Papers
arxiv:2607.24744

Data Pyramid for Embodied Manipulation

Published on Jul 27
Β· Submitted by
Lingdong Kong
on Jul 28
Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,

Abstract

Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require data that couple observations with physical states and actions. These signals can be provided, to varying degrees, by multiple data sources. In this work, we organize the embodied data ecosystem as a "pyramid" spanning five complementary sources: real-robot data, UMI-style data, egocentric and exocentric data, simulation data, and general vision-language data. We organize the pyramid around the tension between scalability and robot alignment, and further characterize each source in terms of data quality, diversity, reusability, and physical fidelity. We then analyze recent embodied foundation models through the lens of their data recipes, examining how different sources are selected, aligned, and mixed during pretraining. For embodied brain models, vision-language-action models, and world-action models alike, we relate data composition to capabilities in perception, reasoning, planning, action generation, and world prediction. We close by discussing six open challenges: building large-scale tactile datasets, collecting failure and recovery data, developing scalable data-collection pipelines, aligning actions across embodiments, leveraging egocentric data for dexterous manipulation, and designing principled data recipes for robot learning. We hope this work paves the foundation for the design of next-generation embodied systems.

Community

Paper submitter

Multimodal foundation models learned to see and reason from Internet-scale image-text data. Embodied AI, however, requires a fundamentally different ingredient: interaction data that couples perception, physical states, and actions.

But today’s embodied data ecosystem is highly fragmented. Real-robot trajectories, UMI demonstrations, egocentric videos, simulation, and general vision-language data each provide different forms of supervision, yet their roles and trade-offs remain poorly understood.

In this work, we introduce the Embodied Data Pyramid, a unified taxonomy that organizes embodied data into five complementary layers:

  • Real-Robot Data β€” highest robot alignment and physical fidelity πŸ€–
  • UMI Data β€” scalable robot-free demonstrations with action supervision 🦾
  • Egocentric & Exocentric Data β€” rich human interaction priors from the real world πŸ‘“
  • Simulation Data β€” scalable robot-oriented interaction with privileged supervision πŸ•ΉοΈ
  • General Data β€” web-scale perception, reasoning, and semantic knowledge 🌐

Beyond organizing the data landscape, we analyze how recent Embodied Brain Models, Vision-Language-Action Models (VLAs), and World-Action Models (WAMs) combine these heterogeneous data sources, revealing the emerging trends in embodied pretraining recipes.

Finally, we discuss several open challenges, including tactile data, failure and recovery trajectories, scalable data collection, cross-embodiment action alignment, egocentric priors for dexterous manipulation, and principled data recipes for robot learning.

We hope this survey provides a useful roadmap for understanding, collecting, and utilizing data for the next generation of embodied foundation models.

Feedback and discussions are very welcome!

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.24744
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2607.24744 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2607.24744 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2607.24744 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.