StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models Paper • 2608.26067 • Published 3 days ago • 17
FlowR2A: Learning Reward-to-Action Distribution for Multimodal Driving Planning Paper • 2606.24231 • Published Jun 23 • 1
VidToMe: Video Token Merging for Zero-Shot Video Editing Paper • 2312.10656 • Published Dec 17, 2023 • 11