HydraHead: From Head-Level Functional Heterogeneity to Specialized Attention Hybridization Paper • 2606.20097 • Published Jun 18 • 19
HydraHead: From Head-Level Functional Heterogeneity to Specialized Attention Hybridization Paper • 2606.20097 • Published Jun 18 • 19
HydraHead: From Head-Level Functional Heterogeneity to Specialized Attention Hybridization Paper • 2606.20097 • Published Jun 18 • 19
Nemotron-Pre-Training-Datasets Collection Large scale pre-training datasets used in the Nemotron family of models. • 15 items • Updated 17 days ago • 187
microsoft/Phi-4-multimodal-instruct Automatic Speech Recognition • 6B • Updated Dec 10, 2025 • 293k • 1.61k