VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Paper • 2607.14935 • Published 11 days ago • 169
view article Article Multimodal Embedding & Reranker Models with Sentence Transformers tomaarsen • Apr 9 • 68
video-SALMONN 2 Collection video-SALMONN 2 is a powerful audio-visual large language model (LLM) that generates high-quality audio-visual video captions. • 11 items • Updated Mar 21 • 2
view article Article VLX-Flow: Continuous Video Understanding for Real-Time Multimodal Interaction omlab • 30 days ago • 15
view article Article Profiling in PyTorch (Part 3): Attention is all you profile +2 ariG23498, sergiopaniego, sayakpaul, ror • 17 days ago • 39