July
2026
What needs to be learned in robot learning? A case study: learning battery insertion from a diagram
Authors:
Abstract:
Manual diagrams are a rich and common knowledge source humans use to learn new skills, but their use for robot learning is still underexplored. A challenge with instruction diagrams is that they communicate task progression in a sparse, qualitative visual format, relying on the learner's prior physical understanding to fill in unmentioned execution details. This thesis investigates the fundamental question of "what needs to be learned in robot learning" by exploring the interplay between explicit information extracted from instructions and implicit physical knowledge discovered through practice or prior knowledge, using cylindrical battery insertion task as a case study.
First, we present a pipeline that compiles static 2D instruction diagrams into metrically accurate 3D simulation environments. We use Vision-Language Models (VLMs) to extract qualitative scene topology and contact modes, followed by a geometric optimization solver that certifies and refines metric object dimensions and spatial subgoals. Second, using the reconstructed task keyframes, we manually designed a control strategy on physical hardware. This hardware deployment exposes the limitations of purely explicit instructions and reveals crucial implicit details such as "tricks" - open-loop primitives exploiting the task's physical properties that yield robustness gains over naive closed-loop methods. Finally, we investigate whether these implicit physical behaviors can be discovered autonomously by reinforcement and imitation learning.
In summary, this thesis demonstrates a paradigm for making use of human data (such as explicit demonstration videos found on YouTube, instructional text, and diagrams) for robot skill learning, showing a promising approach to make use of the vast knowledge base of humans.
First, we present a pipeline that compiles static 2D instruction diagrams into metrically accurate 3D simulation environments. We use Vision-Language Models (VLMs) to extract qualitative scene topology and contact modes, followed by a geometric optimization solver that certifies and refines metric object dimensions and spatial subgoals. Second, using the reconstructed task keyframes, we manually designed a control strategy on physical hardware. This hardware deployment exposes the limitations of purely explicit instructions and reveals crucial implicit details such as "tricks" - open-loop primitives exploiting the task's physical properties that yield robustness gains over naive closed-loop methods. Finally, we investigate whether these implicit physical behaviors can be discovered autonomously by reinforcement and imitation learning.
In summary, this thesis demonstrates a paradigm for making use of human data (such as explicit demonstration videos found on YouTube, instructional text, and diagrams) for robot skill learning, showing a promising approach to make use of the vast knowledge base of humans.
Notes:
copied = false, 2000);
">
@mastersthesis{Doan-2026-88333,
author = {Duc Doan},
title = {What needs to be learned in robot learning? A case study: learning battery insertion from a diagram},
year = {2026},
month = {July},
school = {Carnegie Mellon University},
address = {Pittsburgh, PA},
number = {CMU-RI-TR-26-89},
}
author = {Duc Doan},
title = {What needs to be learned in robot learning? A case study: learning battery insertion from a diagram},
year = {2026},
month = {July},
school = {Carnegie Mellon University},
address = {Pittsburgh, PA},
number = {CMU-RI-TR-26-89},
}