HOI-R1 Exploring the Potential of Multimodal Large Language Models for Human-Object Interaction Detection. thxplz/HOI-R1_Qwen2.5-VL-3B-Instruct Image-Text-to-Text • 4B • Updated Dec 27, 2025 • 1.18k • 1 thxplz/HOI-R1_Qwen3-VL-4B-Instruct Image-Text-to-Text • 4B • Updated May 22 • 9 • 1
HOI-R1 Exploring the Potential of Multimodal Large Language Models for Human-Object Interaction Detection. thxplz/HOI-R1_Qwen2.5-VL-3B-Instruct Image-Text-to-Text • 4B • Updated Dec 27, 2025 • 1.18k • 1 thxplz/HOI-R1_Qwen3-VL-4B-Instruct Image-Text-to-Text • 4B • Updated May 22 • 9 • 1