Robot Self-Discovery: Vision-Based System Teaches Machines to Understand Their Bodies

A New Approach to Robotic Control

In an office at MIT’s Computer Science and Artificial Intelligence Laboratory (CSAIL), a soft robotic hand carefully curls its fingers to grasp a small object. What makes this scene unique is not the mechanical design or embedded sensors—it’s the absence of any such components. Instead, the entire system relies on a single camera that observes the robot’s movements and uses that visual data to control it.

This groundbreaking capability comes from a new system developed by CSAIL scientists, which offers a fresh perspective on robotic control. Rather than relying on hand-designed models or complex sensor arrays, the system allows robots to learn how their bodies respond to control commands, solely through vision. The approach, called Neural Jacobian Fields (NJF), gives robots a form of bodily self-awareness.

Robot Self-Discovery: Vision-Based System Teaches Machines

The research was published in Nature and highlights a shift in how robots are programmed. According to Sizhe Lester Li, an MIT Ph.D. student in electrical engineering and computer science, and lead researcher on the project, “This work points to a shift from programming robots to teaching robots.”

Learning Through Observation

Today, many robotics tasks require extensive engineering and coding. However, the future envisioned by the researchers involves showing a robot what to do and letting it learn how to achieve the goal autonomously. The motivation behind this innovation stems from a simple but powerful reframing: the main barrier to affordable, flexible robotics isn’t hardware—it’s control capability, which can be achieved in multiple ways.

Traditional robots are built to be rigid and sensor-rich, making it easier to construct a digital twin, a precise mathematical replica used for control. But when a robot is soft, deformable, or irregularly shaped, those assumptions fall apart. NJF flips the script by giving robots the ability to learn their own internal model from observation.

Expanding Design Possibilities

This decoupling of modeling and hardware design could significantly expand the design space for robotics. In soft and bio-inspired robots, designers often embed sensors or reinforce parts of the structure just to make modeling feasible. NJF removes that constraint, allowing designers to explore unconventional morphologies without worrying about whether they’ll be able to model or control them later.

“Think about how you learn to control your fingers: you wiggle, you observe, you adapt,” says Li. “That’s what our system does. It experiments with random actions and figures out which controls move which parts of the robot.”

The system has proven robust across a range of robot types. The team tested NJF on a pneumatic soft robotic hand capable of pinching and grasping, a rigid Allegro hand, a 3D-printed robotic arm, and even a rotating platform with no embedded sensors. In every case, the system learned both the robot’s shape and how it responded to control signals, just from vision and random motion.

Future Applications

The researchers see potential far beyond the lab. Robots equipped with NJF could one day perform agricultural tasks with centimeter-level localization accuracy, operate on construction sites without elaborate sensor arrays, or navigate dynamic environments where traditional methods break down.

At the core of NJF is a neural network that captures two intertwined aspects of a robot’s embodiment: its three-dimensional geometry and its sensitivity to control inputs. The system builds on neural radiance fields (NeRF), a technique that reconstructs 3D scenes from images by mapping spatial coordinates to color and density values.

NJF extends this approach by learning not only the robot’s shape but also a Jacobian field, a function that predicts how any point on the robot’s body moves in response to motor commands. To train the model, the robot performs random motions while multiple cameras record the outcomes. No human supervision or prior knowledge of the robot’s structure is required—the system simply infers the relationship between control signals and motion by watching.

Once training is complete, the robot only needs a single monocular camera for real-time closed-loop control, running at about 12 Hertz. This allows it to continuously observe itself, plan, and act responsively. That speed makes NJF more viable than many physics-based simulators for soft robots, which are often too computationally intensive for real-time use.

Advancing Robotics Capabilities

In early simulations, even simple 2D fingers and sliders were able to learn this mapping using just a few examples. By modeling how specific points deform or shift in response to action, NJF builds a dense map of controllability. That internal model allows it to generalize motion across the robot’s body, even when the data are noisy or incomplete.

“What’s really interesting is that the system figures out on its own which motors control which parts of the robot,” says Li. “This isn’t programmed—it emerges naturally through learning, much like a person discovering the buttons on a new device.”

The future of robotics is moving toward soft, bio-inspired designs that can adapt to the real world more fluidly. While these robots are harder to model, NJF aims to lower the barrier, making robotics more affordable, adaptable, and accessible.

Vision as a Reliable Sensor

“Vision alone can provide the cues needed for localization and control—eliminating the need for GPS, external tracking systems, or complex onboard sensors,” says co-author Daniela Rus, MIT professor of electrical engineering and computer science and director of CSAIL. “This opens the door to robust, adaptive behavior in unstructured environments, from drones navigating indoors or underground without maps to mobile manipulators working in cluttered homes or warehouses, and even legged robots traversing uneven terrain.”

While training NJF currently requires multiple cameras and must be redone for each robot, the researchers are imagining a more accessible version. In the future, hobbyists could record a robot’s random movements with their phone and use that footage to create a control model, with no prior knowledge or special equipment required.

The system doesn’t yet generalize across different robots, and it lacks force or tactile sensing, limiting its effectiveness on contact-rich tasks. But the team is exploring new ways to address these limitations, including improving generalization, handling occlusions, and extending the model’s ability to reason over longer spatial and temporal horizons.

“Just as humans develop an intuitive understanding of how their bodies move and respond to commands, NJF gives robots that kind of embodied self-awareness through vision alone,” says Li. “This understanding is a foundation for flexible manipulation and control in real-world environments.”

The research reflects a broader trend in robotics: moving away from manually programming detailed models toward teaching robots through observation and interaction.

Leave a Comment