TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM Paper • 2607.27205 • Published 6 days ago • 134
Multimodal Speaker Verification as a Threat to Speaker Anonymization Paper • 2607.19636 • Published 13 days ago • 2
DSWorld: A Data Science World Model for Efficient Autonomous Agents Paper • 2607.15901 • Published 18 days ago • 12
LATO.2: Factorized 3D Mesh Generation with Vertex and Topology Flow Paper • 2607.10623 • Published 23 days ago • 14
AI translation of literary texts is "fine", but readers still prefer human translations Paper • 2606.26040 • Published Jun 24 • 9
SkillHone: A Harness for Continual Agent Skill Evolution Through Persistent Decision History Paper • 2606.08671 • Published Jun 23 • 46
InterleaveThinker: Reinforcing Agentic Interleaved Generation Paper • 2606.13679 • Published Jun 11 • 84
Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players Paper • 2605.28816 • Published May 27 • 433
The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm Paper • 2604.20665 • Published May 21 • 7