OpenMOSS-Team/MOSS-Transcribe-Diarize Audio-Text-to-Text • 0.9B • Updated about 18 hours ago • 117k • 325
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Paper • 2607.14935 • Published 9 days ago • 168
view article Article Multimodal Embedding & Reranker Models with Sentence Transformers tomaarsen • Apr 9 • 67
video-SALMONN 2 Collection video-SALMONN 2 is a powerful audio-visual large language model (LLM) that generates high-quality audio-visual video captions. • 11 items • Updated Mar 21 • 2
view article Article VLX-Flow: Continuous Video Understanding for Real-Time Multimodal Interaction omlab • 28 days ago • 15
view article Article Profiling in PyTorch (Part 3): Attention is all you profile +2 ariG23498, sergiopaniego, sayakpaul, ror • 15 days ago • 38