TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Paper • 2607.17423 • Published 6 days ago • 165
BadWAM: When World-Action Models Dream Right but Act Wrong Paper • 2607.15207 • Published 9 days ago • 53
Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos Paper • 2607.11523 • Published 12 days ago • 13
AGE: Adaptive-masking for Graph Embedding in Graph Retrieval-Augmented Generation Paper • 2607.00052 • Published 25 days ago • 7
InterleaveThinker: Reinforcing Agentic Interleaved Generation Paper • 2606.13679 • Published Jun 11 • 83
Agent Skills Should Go Beyond Text: The Case for Visual Skills Paper • 2606.01414 • Published May 31 • 10
Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players Paper • 2605.28816 • Published May 27 • 433