Dyna Robotics has unveiled DYNA-2, its latest foundation model for embodied intelligence, introducing a new approach to robot learning based largely on human first-person video rather than massive amounts of robot-specific training data.
The company says DYNA-2 was pretrained on more than 1 million hours of first-person human video, equivalent to roughly 170 years of human experience. According to Dyna Robotics, the model can achieve task success rates of up to 90% in certain evaluations, highlighting the potential of human video as a scalable source of training data for general-purpose robots.
DYNA-2 Is Designed as a “World Action Model”
Dyna Robotics describes DYNA-2 as a World Action Model, rather than simply another vision-language model.
The core idea is to teach robots how the physical world responds to human actions. By watching people interact with objects and their surroundings, the model learns relationships between actions and their physical consequences.
For example, a robot needs more than visual recognition to understand that pushing an object can make it move, grasping a container can change its position, or rotating a cap can eventually open it.
Dyna Robotics believes this type of physical understanding could help address some of the limitations of conventional vision-language models, particularly in spatial reasoning, movement prediction and physical interaction.
More Than 1 Million Hours of Human Video
One of the most notable aspects of DYNA-2 is the scale and type of its pre-training data.
Dyna Robotics says the model was pretrained entirely on human video data, with more than 1 million hours of first-person footage. The company compares this amount of information to approximately 170 years of human experience.
The strategy addresses one of the biggest challenges facing general-purpose robotics: collecting enough high-quality robot interaction data.
Unlike digital AI systems, physical robots cannot generate enormous amounts of training data simply by running simulations or processing existing internet content. Real-world robot data requires hardware, environments, human supervision and considerable time.
Jason Ma, co-founder of Dyna Robotics, said general-purpose robotics has long been constrained by this data bottleneck. Scaling physical teleoperation data sufficiently to achieve general intelligence is difficult, while human activity provides a much larger potential source of information.
Teaching Robots How the Physical World Changes
DYNA-2 is designed to predict what happens in the physical environment before executing an action.
The model reportedly learns from two major training objectives: predicting the content of the next video frame and predicting the action that should be taken.
This allows the system to learn several types of information simultaneously, including:
- Spatial relationships between objects
- Human and object movement patterns
- Changes in the environment over time
- How objects respond to physical contact
- The relationship between an action and its resulting outcome
The broader objective is to move beyond simply recognizing what is visible and toward understanding what will happen next.
For robots operating in unpredictable environments, that distinction could be particularly important.
Human Video Could Provide a New Scaling Path for Robotics
Dyna Robotics says its experiments demonstrate a relationship between the amount of human video used during pre-training and robot performance.
The company evaluated DYNA-2 across 15 benchmark tasks and reported that robot performance improved as the amount of human video data increased.
Dyna Robotics refers to this as a “scaling law from human to robot.”
The concept suggests that increasing the amount of human experience available during pre-training could produce relatively predictable improvements in robot capabilities.
If this relationship continues to hold at larger scales, human video could become an important alternative to the traditional approach of collecting enormous quantities of robot-specific interaction data.
Just 13 Minutes of Robot Data for a Bottle-Cap Task
Another experiment highlighted by Dyna Robotics involved two five-finger robotic hands.
The company says that only 13 minutes of robot-specific data were needed for the system to learn how to unscrew a bottle cap.
This is significant because collecting robot-specific data can be considerably more expensive and difficult than collecting human video.
Rather than training the entire capability from scratch using robot demonstrations, the model can first learn general physical concepts from human activity and then use a relatively small amount of robot data to adapt those capabilities to a particular robotic platform.
This approach could potentially reduce the amount of expensive physical data required during downstream training.
Manufacturing Performance Increased From 20% to 90%
Dyna Robotics also reported encouraging results from manufacturing-related tasks.
Using the same post-training dataset, the company says the success rate increased from approximately 20% to 80% and eventually 90% as the scale of pre-training data increased.
The results suggest that improvements in general pre-training can translate into better performance on specialized physical tasks, even when the downstream dataset remains unchanged.
If independently validated, this would support Dyna Robotics’ argument that large-scale human video pre-training can provide a meaningful foundation for robot learning.
DYNA-2 Reportedly Improves on Dyna-1 in Deployment
The company also provided results from an actual customer deployment.
According to Dyna Robotics, DYNA-2 achieved an 87% quality pass rate, compared with 46% for the previous-generation Dyna-1.
The company further reported that video collaborative training improved performance by 133% on tasks requiring robots to perform different actions depending on user instructions.
These results point toward a broader goal for Dyna Robotics: developing robots that can respond flexibly to instructions instead of performing only narrowly predefined actions.
The Challenge of Building General-Purpose Robots
The robotics industry has increasingly focused on foundation models as companies seek to create robots capable of performing a wide range of tasks.
However, physical intelligence remains fundamentally different from language or image understanding. A robot must account for gravity, friction, object weight, contact forces, spatial constraints and the consequences of its own movements.
Human beings acquire much of this knowledge naturally through years of observation and interaction with the physical world. Dyna Robotics is attempting to use first-person human video as a way of transferring some of that experience into an AI model.
DYNA-2’s approach therefore represents a broader shift toward treating human experience itself as a scalable training resource for embodied AI.
What DYNA-2 Could Mean for the Future of Robotics
Dyna Robotics’ latest model does not eliminate the need for real-world robot data. Physical robots still need to learn how their specific hardware behaves and how to translate general knowledge into precise motor actions.
However, the company’s results suggest that large-scale human video pre-training could substantially reduce the amount of robot-specific data required for some tasks.
If the reported scaling relationship continues as models and datasets grow, future robot foundation models could potentially learn increasingly sophisticated physical skills from vast libraries of human activities before being adapted to specific robotic platforms.
For now, DYNA-2 remains an emerging approach, and the company’s reported benchmark and deployment results will need to be evaluated independently. Nevertheless, its emphasis on more than 1 million hours of human video, predictive world modeling and limited robot-specific fine-tuning reflects one of the industry’s most interesting attempts to overcome the data bottleneck in embodied AI.
Rather than teaching robots exclusively through robots, Dyna Robotics is betting that watching humans interact with the world can provide a powerful foundation for teaching machines how to act within it.









