Robotics Seminar: Learning from Visual Experience: Toward Generalizable Robot Manipulation
Event Details
On 10/2, Prof Matt Walter will visit from TTIC and speak in the UW -- Madison Robotics Seminar. Talk information:
Title: Learning from Visual Experience: Toward Generalizable Robot Manipulation
Abstract: The dominant paradigm in robot learning seeks to achieve generalizable behavior by training policies on large collections of demonstrations that span different robots, environments, objects, sensor configurations, and tasks. This approach raises complementary questions about how to learn from that experience: what supervision is needed, what information should be represented explicitly, and how should generalization be evaluated? In this talk, I will describe a body of work that addresses these questions through learning from visual experience and evaluating the resulting capabilities.
I will begin with work that seeks to take advantage of the large quantities of human video available as a source of supervision for robot learning. These videos show how people perform tasks, but do not provide the actions required for conventional imitation learning. I will first present an approach that learns a task-agnostic, progress-based reward function from action-free video demonstrations. Pretrained on egocentric human videos, this reward model generalizes to robot manipulation tasks without task-specific fine-tuning, supporting goal-conditioned policy learning. I will then describe work that jointly learns a predictive world model and continuous latent action representations from action-free videos. Key to this approach is an adversarial regularization scheme that encourages the latent representations to capture actions consistently across visual contexts, supporting both imitation learning from observation and goal-directed planning.
Next, I will present work that examines the role of camera geometry in learning policies that generalize across viewpoints. Our analysis reveals that policies learn shortcuts, relying on incidental background cues to infer camera pose, and consequently fail when these cues inevitably change during deployment. I will describe an approach that conditions policies on camera geometry through per-pixel representations of viewing rays, providing an explicit relationship between image observations and the robot’s coordinate frame. This approach improves generalization to unseen viewpoints across several policy architectures without requiring depth observations.
Finally, I will examine the extent to which performance on common manipulation benchmarks provides evidence of general robot capability. I will describe diagnostics that identify reliance on task shortcuts, uncertainty in reported improvements, overfitting to evaluation setups, and dependence on the source of training demonstrations, and discuss their implications for evaluating progress in robot learning.
Bio: Matthew R. Walter is an associate professor at the Toyota Technological Institute at Chicago (TTIC). His interests revolve around the realization of intelligent, perceptually aware robots that are able to act robustly and effectively in unstructured environments, particularly with and alongside people. His research focuses on machine learning-based solutions that allow robots to learn to understand and interact with the people, places, and objects in their surroundings. Matthew has investigated these areas in the context of various robotic platforms, including autonomous underwater vehicles, self-driving cars, voice-commandable wheelchairs, mobile manipulators, and autonomous cars for (rubber) ducks. Matthew obtained his Ph.D. from the Massachusetts Institute of Technology and the Woods Hole Oceanographic Institution, where his thesis focused on improving the efficiency of inference for simultaneous localization and mapping.
We value inclusion and access for all participants and are pleased to provide reasonable accommodations for this event. Please email jphanna@cs.wisc.edu to make a disability-related accommodation request. Reasonable effort will be made to support your request.