Yingjie Lei
ChaceLei2004
·
AI & ML interests
World Models, Computer Vision
Recent Activity
upvoted a paper about 1 month ago
Geometric Action Model for Robot Policy Learning liked a dataset about 1 month ago
facebook/jepa-wms liked a model about 1 month ago
facebook/jepa-wmsOrganizations
None yet
Infra
3D Vision
-
Seed3D 1.0: From Images to High-Fidelity Simulation-Ready 3D Assets
Paper • 2510.19944 • Published • 22 -
Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations
Paper • 2510.23607 • Published • 181 -
Masked Depth Modeling for Spatial Perception
Paper • 2601.17895 • Published • 30
Embodied AI
-
World-in-World: World Models in a Closed-Loop World
Paper • 2510.18135 • Published • 78 -
GigaBrain-0: A World Model-Powered Vision-Language-Action Model
Paper • 2510.19430 • Published • 54 -
World Simulation with Video Foundation Models for Physical AI
Paper • 2511.00062 • Published • 46 -
TwinBrainVLA: Unleashing the Potential of Generalist VLMs for Embodied Tasks via Asymmetric Mixture-of-Transformers
Paper • 2601.14133 • Published • 61
Video Generation
-
VISTA: A Test-Time Self-Improving Video Generation Agent
Paper • 2510.15831 • Published • 24 -
HoloCine: Holistic Generation of Cinematic Multi-Shot Long Video Narratives
Paper • 2510.20822 • Published • 41 -
Video-As-Prompt: Unified Semantic Control for Video Generation
Paper • 2510.20888 • Published • 50 -
The Script is All You Need: An Agentic Framework for Long-Horizon Dialogue-to-Cinematic Video Generation
Paper • 2601.17737 • Published • 56
Technical Reports
RL
Infra
Adaptation
3D Vision
-
Seed3D 1.0: From Images to High-Fidelity Simulation-Ready 3D Assets
Paper • 2510.19944 • Published • 22 -
Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations
Paper • 2510.23607 • Published • 181 -
Masked Depth Modeling for Spatial Perception
Paper • 2601.17895 • Published • 30
Video Understanding
Embodied AI
-
World-in-World: World Models in a Closed-Loop World
Paper • 2510.18135 • Published • 78 -
GigaBrain-0: A World Model-Powered Vision-Language-Action Model
Paper • 2510.19430 • Published • 54 -
World Simulation with Video Foundation Models for Physical AI
Paper • 2511.00062 • Published • 46 -
TwinBrainVLA: Unleashing the Potential of Generalist VLMs for Embodied Tasks via Asymmetric Mixture-of-Transformers
Paper • 2601.14133 • Published • 61
World Models
Video Generation
-
VISTA: A Test-Time Self-Improving Video Generation Agent
Paper • 2510.15831 • Published • 24 -
HoloCine: Holistic Generation of Cinematic Multi-Shot Long Video Narratives
Paper • 2510.20822 • Published • 41 -
Video-As-Prompt: Unified Semantic Control for Video Generation
Paper • 2510.20888 • Published • 50 -
The Script is All You Need: An Agentic Framework for Long-Horizon Dialogue-to-Cinematic Video Generation
Paper • 2601.17737 • Published • 56