Abstract
Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require data that couple observations with physical states and actions. These signals can be provided, to varying degrees, by multiple data sources. In this work, we organize the embodied data ecosystem as a "pyramid" spanning five complementary sources: real-robot data, UMI-style data, egocentric and exocentric data, simulation data, and general vision-language data. We organize the pyramid around the tension between scalability and robot alignment, and further characterize each source in terms of data quality, diversity, reusability, and physical fidelity. We then analyze recent embodied foundation models through the lens of their data recipes, examining how different sources are selected, aligned, and mixed during pretraining. For embodied brain models, vision-language-action models, and world-action models alike, we relate data composition to capabilities in perception, reasoning, planning, action generation, and world prediction. We close by discussing six open challenges: building large-scale tactile datasets, collecting failure and recovery data, developing scalable data-collection pipelines, aligning actions across embodiments, leveraging egocentric data for dexterous manipulation, and designing principled data recipes for robot learning. We hope this work paves the foundation for the design of next-generation embodied systems.
Community
Multimodal foundation models learned to see and reason from Internet-scale image-text data. Embodied AI, however, requires a fundamentally different ingredient: interaction data that couples perception, physical states, and actions.
But todayβs embodied data ecosystem is highly fragmented. Real-robot trajectories, UMI demonstrations, egocentric videos, simulation, and general vision-language data each provide different forms of supervision, yet their roles and trade-offs remain poorly understood.
In this work, we introduce the Embodied Data Pyramid, a unified taxonomy that organizes embodied data into five complementary layers:
- Real-Robot Data β highest robot alignment and physical fidelity π€
- UMI Data β scalable robot-free demonstrations with action supervision π¦Ύ
- Egocentric & Exocentric Data β rich human interaction priors from the real world π
- Simulation Data β scalable robot-oriented interaction with privileged supervision πΉοΈ
- General Data β web-scale perception, reasoning, and semantic knowledge π
Beyond organizing the data landscape, we analyze how recent Embodied Brain Models, Vision-Language-Action Models (VLAs), and World-Action Models (WAMs) combine these heterogeneous data sources, revealing the emerging trends in embodied pretraining recipes.
Finally, we discuss several open challenges, including tactile data, failure and recovery trajectories, scalable data collection, cross-embodiment action alignment, egocentric priors for dexterous manipulation, and principled data recipes for robot learning.
We hope this survey provides a useful roadmap for understanding, collecting, and utilizing data for the next generation of embodied foundation models.
Feedback and discussions are very welcome!
Get this paper in your agent:
hf papers read 2607.24744 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper