Despite rapid advances in language and vision models, current robots still lag far behind human physical capabilities due to the relative scarcity of real-world data compared to online text and images. How can we leverage abundant language data to advance robotic capabilities? Language provides semantic structure that facilitates the understanding of diverse data, improving sample efficiency in scarce data regimes. It also provides a natural communicative medium when interacting with and learning from humans.
To leverage the first benefit of language, we first take inspiration from how humans teach each other in video tutorials, through simultaneous video and language streams, to more efficiently teach robots new skills. We then show that language can bridge wide visual sim2real gaps, enabling robots to learn tasks with just a few real-world demonstrations by leveraging knowledge from imperfect simulation data. To leverage the second benefit of language, we explore how dialog can enable robots to solve complex manipulation tasks by communicating and collaborating with a wide distribution of human collaborators in the real-world. We develop a robotic framework that requests and proactively offers help through mixed-initiative, free-form dialog, enabling the robot to adapt to changing human preferences and strategically utilize each agent’s physical capabilities. Finally, to accelerate how robots learn to operate new household devices, we combine both benefits of language into a framework that leverages semantics from rich textual corpora, such as user manuals, to establish skill priors that are efficiently refined through active human dialog.
Overall, our thesis provides a foundation for leveraging the key benefits of language to improve the generalization capabilities of robotic manipulation policies to new tasks, domains, devices, and human collaborators.
PhD Thesis, Department of Computer Science, UT Austin.