July
2026
Towards Scalable Robot Learning: From Teleoperation to Web-scale Data
Authors:
Abstract:
Humanoid robots operating in human environments must manipulate articulated objects under contact and kinematic constraints that human demonstrations do not satisfy. That mismatch makes the human–humanoid embodiment gap the central bottleneck for learning from human data: robot demonstrations are expensive and sparse, while human demonstrations inhabit a different state-action space and often violate robot kinematic constraints. This thesis studies how to convert human behavior into supervision that remains executable for the target robot body.
The first part develops HAT (Humanoid Policy ∼ Human Policy) for cross-embodiment supervision in humanoid manipulation. It places humans and humanoids in a unified state-action representation, enabling a transformer policy to co-train on human and robot demonstrations and retarget its predictions at deployment. To support this formulation, we introduce PH2D, a task-oriented egocentric human demonstration dataset that expands data scale without discarding embodiment structure.
The second part presents EmbodyHOI, which addresses a harder embodiment-gap setting in dexterous hand-object interaction. It starts from a flow-matching diffusion model trained in human hand-object space, then applies a differentiable guidance function during sampling to steer trajectories toward a target humanoid embodiment, jointly optimizing
wrist reachability and base placement before downstream control. Together, these chapters show that scalable robot manipulation requires data transformations that preserve task structure while respecting the robot body.
The first part develops HAT (Humanoid Policy ∼ Human Policy) for cross-embodiment supervision in humanoid manipulation. It places humans and humanoids in a unified state-action representation, enabling a transformer policy to co-train on human and robot demonstrations and retarget its predictions at deployment. To support this formulation, we introduce PH2D, a task-oriented egocentric human demonstration dataset that expands data scale without discarding embodiment structure.
The second part presents EmbodyHOI, which addresses a harder embodiment-gap setting in dexterous hand-object interaction. It starts from a flow-matching diffusion model trained in human hand-object space, then applies a differentiable guidance function during sampling to steer trajectories toward a target humanoid embodiment, jointly optimizing
wrist reachability and base placement before downstream control. Together, these chapters show that scalable robot manipulation requires data transformations that preserve task structure while respecting the robot body.
copied = false, 2000);
">
@mastersthesis{Chawla-2026-88331,
author = {Chaitanya Chawla},
title = {Towards Scalable Robot Learning: From Teleoperation to Web-scale Data},
year = {2026},
month = {July},
school = {Carnegie Mellon University},
address = {Pittsburgh, PA},
number = {CMU-RI-TR-26-72},
keywords = {Learning from human demonstrations, cross-embodiment learning, dexterous manipulation},
}
author = {Chaitanya Chawla},
title = {Towards Scalable Robot Learning: From Teleoperation to Web-scale Data},
year = {2026},
month = {July},
school = {Carnegie Mellon University},
address = {Pittsburgh, PA},
number = {CMU-RI-TR-26-72},
keywords = {Learning from human demonstrations, cross-embodiment learning, dexterous manipulation},
}