๐ฑ POCKET โ a 35-billion-parameter model that runs on your iPhone, and on your PC with no GPU
We're releasing POCKET, VIDRAFT's flagship Darwin-36B-Opus compressed for on-device use. No fork, no CUDA, no cloud โ it runs on stock llama.cpp. It's a sparse Mixture-of-Experts model (256 experts, only 8 active per token), so the file can be large while the work per token stays small. That's what lets a 35B model run on a phone, and generate fast on a CPU with no graphics card.
Measured (POCKET-35B IQ1_M vs Bonsai-27B Q1_0): โข CPU generate (Xeon, 16 threads): 27.0 vs 10.1 tok/s โ 2.69ร faster โข GPU generate (H100): 197 vs 89 tok/s โ 2.22ร faster โข GPU prompt processing (H100): 753 vs 1816 โ 0.41ร (Bonsai wins this one โ MoE prefill wakes every expert, so sparsity stops helping there. We say so.) โข Quality (HellaSwag, 400 q): 61.0% vs 60.0% โ a tie (confidence intervals overlap)
On a real consumer laptop โ MacBook M3 Pro (18 GB) โ POCKET wins every axis, prompt processing included: โข Metal generate: 25.4 vs 12.8 โ 1.99ร โข CPU generate: 13.8 vs 4.4 โ 3.13ร โข Metal prompt: 240.7 vs 73.4 โ 3.28ร
One more quiet fact: the same-size, quality-oriented rival Ternary-Bonsai-27B (7.2 GB) fails to load in upstream llama.cpp at all โ it needs the PrismML fork. POCKET runs on the tools you already have: LM Studio, Ollama, PocketPal, MLX.
๐ง Does your LLM know when it's about to be wrong?
Most leaderboards measure accuracy. We measure metacognition โ whether a model catches its own errors. Benchmark + leaderboard + adapters, all open. ๐
The surprise: even a K-AI #1 model (JGOS-31B-Citizen) is the strongest on multiple-choice traps (trap_rate 0.005 โ ~2 misses in 400) yet blind to its own free-form mistakes (self-confidence AUROC = 0.5, pure random). A tiny base-frozen adapter recovers that signal.
Two independent axes (never compared across a row): โ trap_rate โ does it fall for tempting trap options? (lower = stronger) โก adapter gain ฮ โ how much a lightweight adapter catches errors the model itself misses. (higher = more adapter value)
What's open: ๐ 300+100 trap problems (each with a hidden trap + TICOS type) ๐ 24-model leaderboard ๐งฉ 11 per-model adapters โ adapters, NOT fine-tunes (base stays frozen; the adapter just reads the hidden state โ P(wrong))
Submit any HF model โ auto-scored daily at 09:00 KST and added to the board.
NEW RELEASE: Esper 4 is here for Qwen 3.6 27b, along with our new datasets!
- NEW DATASET: Titanium 4 maximizes DevOps and architecture helpfulness, powered by high-difficulty agentic-focused DevOps and architecture data generated with DeepSeek-V4-Pro! - NEW DATASET: Mitakihara 2 brings AI coding and expertise data for AI development, research, deployment, interpretability, operation and experimentation! - Improved coding performance: challenging agentic coding queries from Tachibana 4 allow Esper 4 to tackle harder coding tasks across a variety of languages!
We've been working hard on Esper 4 - it's so exciting to finally bring it to everyone! We hope it helps you build.
We'll be expanding Esper 4 to more models as funding allows - donate for more, faster, better models and datasets: sequelbox/SupportOpenSource
The revolution is coming - we're here to fight for AI you can use and build on your own computer, not a giant corporation charging you for access at their discretion. We've seen what OpenAI, Anthropic, and the ultra-rich taking charge of the AI future looks like, and it's already very clear you won't like living in it. Choose a different future while you still can.