view article Article The Best Open Source and Open-Weight LLM Models to Run Locally in 2026 daya-shankar • May 13 • 19
view article Article We changed one line and the benchmark score moved 0.21 AUROC FINAL-Bench • 7 days ago • 11
view article Article Measuring benchmark optimization in speech recognition +5 tlebryk02, bezzam, aliceebaird, dayllon, jpc, jens-hume-ai, tzirakis • 8 days ago • 57
view article Article Beyond the Aggregate Score: Per-Country Domain Shift in the GWHD Wheat Head Detection Model Zoo dronefreak • 17 days ago • 7
view article Article Build an AI Evaluation from a Hugging Face Dataset Without Writing Python phranzia • 24 days ago • 3
view article Article Build an AI Evaluation from a Hugging Face Dataset Without Writing Python phranzia • 24 days ago • 3
view article Article LettucePrevent - Real-Time Prevention of Factual Hallucinations in RAG lebe1 • about 1 month ago • 9
view article Article GPU Management: Why Idle GPUs Are the New Grounded Aircraft Dharma-AI • 29 days ago • 95
view article Article Is it agentic enough? Benchmarking open models on your own tooling +1 lysandre, SaylorTwift, pcuenq • Jun 18 • 23
view article Article ScarfBench: Benchmarking AI Agents for Enterprise Java Framework Migration ibm-research • Jun 30 • 26
view article Article Welcome Inkling by Thinking Machines +3 burtenshaw, merve, pcuenq, ariG23498, andito • Jul 15 • 165
view article Article Run and Compare AI Evaluations with a CLI for Developers and Coding Agents phranzia • Jul 27 • 2
view article Article Run and Compare AI Evaluations with a CLI for Developers and Coding Agents phranzia • Jul 27 • 2