BTL-4 Compact

The whole 35B model in a single 9.96 GB file. 2.30 bits per weight, and it retains 94.1% of the full-precision model's measured behaviour.

BTL-4 is a mixture of experts with roughly 2.1B active parameters per token, so it costs a large model's memory and a small model's compute. Compact is the edition that runs on hardware you already own โ€” one file, one command, a running agent. No base download, no reconstruction.

Loads in llama.cpp, Ollama and LM Studio.

Full-precision weights: badtheorylabs/BTL-4

build size bits/weight behavioural retention
BTL-4-IQ2_XXS.gguf 9.96 GB 2.30 94.1%

Retention is measured, not estimated: 118 items on which the full-precision bf16 model is correct, replayed against this build. It reproduces 111 of them. Per category: 95.0% short-form factual, 100% grounded extraction, 87.2% false-premise rejection. The gate resolves to about ยฑ3.4 points, so treat differences smaller than that as noise.

Run it

llama-cli -m BTL-4-IQ2_XXS.gguf -p "Refactor this function to be pure." -c 8192
llama-server -m BTL-4-IQ2_XXS.gguf --port 8080 -c 8192

Requires a llama.cpp with qwen3_5_moe support (src/models/qwen35moe.cpp).

Architecture

total parameters 35.1B (34.7B excluding the vision tower)
active per token ~2.1B
layers 40 โ€” 30 linear-attention, 10 full-attention
experts 256 per layer, 8 routed per token
context 262,144 native
KV cache ~20 KB/token

Only 10 of 40 layers keep a growing KV cache, and those use 2 KV heads. The whole 262K window costs about 5.2 GB of cache, so long-context work fits on consumer hardware.

Notes on this build

The MTP layer is disabled. The source model declares mtp_num_hidden_layers: 1 and the converter writes block_count = 41 while emitting tensors for only 40 blocks, so a stock loader fails on blk.40.attn_norm.weight. This build sets block_count = 40 and nextn_predict_layers = 0. The multi-token-prediction head is a speculative decoding accessory; the model runs without it.

The vision tower is not included. This is a text-only build.

Quantisation

The 120 expert tensors are IQ2_XXS (2.0625 bpw); everything else follows the Q4_K_M mixture. An importance matrix was computed over 120 chunks of a 3 MB corpus of source code, technical documentation and question prompts โ€” a deliberate match for what this model is for, rather than generic web text.

The router (ffn_gate_inp) and every normalisation tensor stay at f32. Routing decides which experts a token reaches, so error there changes which knowledge gets used rather than degrading it smoothly, and at ~21M parameters it is free to protect.

Where the 2.30 bpw goes: the experts are 93% of all parameters and contribute 1.92 bpw; the remaining 0.38 comes from the 4-bit and 6-bit non-expert matrices plus the f32 router and norms.

Two findings from simulation work on this model shaped the recipe. Range selection dominates everything else at low bit widths โ€” replacing min/max group ranging with a per-group MSE clip search moved retention from 77.1% to 95.8% at an identical byte budget. And protecting the output head, the usual recommendation, is worth nothing: head and embedding at 4-bit retained 118 of 118. IQ2_XXS with an imatrix performs its own importance-weighted range search, which is why it is the build shipped here.

Licence

Apache-2.0, inherited from the base model.

ยฉ 2026 Bad Theory Labs

Downloads last month
-
GGUF
Model size
35B params
Architecture
qwen35moe
Hardware compatibility
Log In to add your hardware

2-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for badtheorylabs/BTL-4-Compact

Quantized
(5)
this model