Running Repro - Demystifying LLM-as-a-Judge: Analytically Tractable Model for Inference-Time Scaling 🎯 Track and share experiment logs with an AI coding agent
Running Repro - On Structured State-Space Duality 🎯 Collaborate with an AI agent to view and edit your logbook
Running Reproduction: Optimal Unconstrained Self-Distillation in Ridge Regression: Strict Improvements, Precise Asymptotics, and One-Shot Tuning 🎯 Explore a research logbook and sync with coding agents
Running Reproduction: Flat Minima and Generalization: Insights from Stochastic Convex Optimization 📐 Explore research logbook and sync findings with an AI agent
Running Repro — CapBencher: Give Your LLM Benchmark a Built-in Alarm for Test-Set Overfitting (formerly “How Can I Publish My LLM Benchmark Without Giving the True Answers Away?”) 🎯 Explore and sync LLM benchmark logs with a collaborative agent
Running Repro - A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness 🎯 Collaborate with an AI agent to manage a safety logbook
Running Repro - Minimum Distance Summaries for Robust Neural Posterior Estimation 🎯 Collaborate on experiment logs with an AI coding agent
Running To Grok Grokking — independent ridge reproduction 📐 View and sync a logbook with your coding agent