view article Article Beyond the Aggregate Score: Per-Country Domain Shift in the GWHD Wheat Head Detection Model Zoo dronefreak • 12 days ago • 7
view article Article Build an AI Evaluation from a Hugging Face Dataset Without Writing Python phranzia • 19 days ago • 3
view article Article Build an AI Evaluation from a Hugging Face Dataset Without Writing Python phranzia • 19 days ago • 3
view article Article LettucePrevent - Real-Time Prevention of Factual Hallucinations in RAG lebe1 • 25 days ago • 9
view article Article GPU Management: Why Idle GPUs Are the New Grounded Aircraft Dharma-AI • 24 days ago • 95
view article Article Is it agentic enough? Benchmarking open models on your own tooling +1 lysandre, SaylorTwift, pcuenq • Jun 18 • 23
view article Article ScarfBench: Benchmarking AI Agents for Enterprise Java Framework Migration ibm-research • Jun 30 • 26
view article Article Welcome Inkling by Thinking Machines +3 burtenshaw, merve, pcuenq, ariG23498, andito • Jul 15 • 164
view article Article Run and Compare AI Evaluations with a CLI for Developers and Coding Agents phranzia • 27 days ago • 2
view article Article Run and Compare AI Evaluations with a CLI for Developers and Coding Agents phranzia • 27 days ago • 2