Scaling Sim-to-Real Learning for Robot Manipulation

July 2026

Scaling Sim-to-Real Learning for Robot Manipulation

Authors:

Yufei Wang

Abstract:

Building a generalist robot capable of performing diverse tasks in unstructured environments remains a longstanding challenge.
A recent trend in robot learning aims to address this by scaling up demonstration datasets for imitation learning.
However, most large-scale robotics datasets are collected in the real-world, often via manual teleoperation.
This process is labor-intensive, slow, hardware-dependent, and poses safety risks, limiting its scalability.

Physics-based simulation offers a scalable, safe, and efficient alternative for generating large-scale robot datasets, as it scales with computation rather than manpower. However, the full potential of this approach is limited by the following challenges: (1) large manual effort is required to design simulation assets, scenes, and create training supervisions such as reward functions, (2) the sim-to-real gap in both sensing and dynamics hinders real-world deployment of simulation-trained policies, and (3) existing end-to-end policy representations struggle to generalize on large datasets with tremendous variations.

In this thesis, we focus on learning foundational manipulation policies that generalize across diverse objects, tasks, and environments by scaling up sim2real learning, while addressing these challenges:
1. Policy Representation for Broad Generalization.
I develop structured policy representations that enable broad generalization from large, diverse simulation datasets, demonstrating that such representations enable zero-shot transfer to the real world with superior generalization performance on complex tasks involving deformable objects and physical human-robot interaction.
2. Automatic Generation of Large-scale Simulation Datasets.
To make sim-to-real learning practical at scale,
I introduced a novel paradigm, termed Generative Simulation, which leverages generative foundation models to automate the creation of large simulation datasets, including tasks, assets, scenes, and training supervisions (e.g., rewards), with minimal human effort.
3. Efficient Adaptation of Sim2Real Policies. No simulation is perfect. I develop algorithms that efficiently adapt simulation-trained policies using limited real-world data -- for instance, by leveraging additional modalities available only in the real world -- to enhance real-world performance and safety, e.g., in the challenging assistive task of robot-assisted dressing.
Together, these efforts form a cohesive agenda centered on scalable, generalizable, and adaptive robot learning.

Notes:

@phdthesis{Wang-2026-88316,
author = {Yufei Wang},
title = {Scaling Sim-to-Real Learning for Robot Manipulation},
year = {2026},
month = {July},
school = {Carnegie Mellon University},
address = {Pittsburgh, PA},
number = {CMU-RI-TR-26-55},
keywords = {Robot Manipulation, Sim-to-Real Learning},
}
Copyright notice: This material is presented to ensure timely dissemination of scholarly and technical work. Copyright and all rights therein are retained by authors or by other copyright holders. All persons copying this information are expected to adhere to the terms and constraints invoked by each author's copyright. These works may not be reposted without the explicit permission of the copyright holder.