AlephLM-0 — an anchored expert trunk, distilled against a dense control

This is a live experiment repository, not a finished model release. Runs land here as they finish training, checkpoints push every 30 minutes mid-run, and every arm ships — including any that end up refuted. If you are reading this while the run table below says IN PROGRESS, you are watching the experiment happen.

The question

Mixture-of-experts models normally route with a learned softmax over expert logits — a comparative choice among experts. This program tests a different router: a closed-form signed address over unit anchor directions,

u_k = cos(x, a_k) / τ          w_k = sinh(u_k) / Σ_j cosh(u_j)

where each expert's contribution is w_k · σ(g_k) · E_k(x) per token. The weights are signed — an expert can be recruited negatively (an inhibitory anchor) — and the read is reconstructive rather than competitive: no argmax, no top-k, no load-balancing loss. The anchors, gates, and experts are trained by nothing but the task gradient.

E1 (this repo): does a trunk built this way match or beat a parameter-matched dense trunk under an identical objective, at 32M-row scale? Six runs answer it:

run encoder routing seeds
a1_anchored trunk-expert ff512 + 3 dispatched experts ff512/block signed aleph address, learned anchors s0, s1
a2_dense standard dense ff2048 — (the control) s0, s1
a3_random same as a1 anchors frozen at random init s0, s1

a1 vs a2 is the headline; a1 vs a3 isolates whether learned addressing matters or any fixed partition of the capacity would do.

Architecture

12 layers, d=512, 8 heads, pre-norm, 8192 learned positions, 768-d projected output, CLS readout (settled empirically — see S0e below).

  • Per block, the dense FFN (ff2048) is replaced by 1 always-on trunk expert (ff512) + 3 dispatched experts (ff512 each) — 2048 hidden units total, exact capacity parity with the control.
  • Dispatched-expert output layers are zero-initialized and gates start at σ(−3) ≈ 0.047: at initialization the dispatch contributes exactly zero (bit-exact, asserted at construction), so the anchored trunk is born as its own dense-trunk null hypothesis and the routing must earn its way in. One known consequence: the routing gradient is zero for exactly one step (∂L/∂w = σ(g)·E(x) and E ≡ 0 at init), the same dynamic as LoRA's A-matrix under B=0.
  • Parameter cost of the machinery: +36,900 over dense (+0.063%) — 12 codebooks of 3×512, 36 gates, and the extra expert biases. 58,345,764 vs 58,308,864.

Training recipe (identical for every arm)

Consensus distillation, inherited verbatim from captionbert-8192-v2: the target for each caption is the L2-normalized centroid of five BERT-family teachers, each mapped into the reference member's frame (bert-base) by a whitened Procrustes fit — the precomputed targets cover 33M captions from CC12M.

  • loss = InfoNCE(T=0.07, in-batch negatives) + MSE (F.mse_loss, per-element mean — the batch of 2048 is the negative set, so batch size is part of the objective and is never changed)
  • pure Adam (no weight decay), lr 6e-4, linear warmup 2000 → cosine to 1e-6, grad clip 1.0, AMP fp16, 4 epochs over 64 train chunks (31.9M rows), 2 holdout chunks for eval
  • length-bucketed dynamic padding (ceiling 256 tokens), gradient checkpointing
  • trained on a single RTX 5090 (32GB); worst-case batch measured 30.1 GB reserved

Stage-0 instruments (complete)

S0a — is the rank ceiling the teachers' agreement, or bert's own geometry? (s0a/s0a_erank.json) The consensus target occupies an effective rank of 28.1/768. Raw bert-base rows on the same corpus: 40.7/768 — and 40.3 on out-of-domain STS-B text, so the low rank is the encoder's geometry, not the corpus. Verdict at the matched (L2-normalized) gauge: ratio 1.45× → intermediate — the consensus construction costs ~30% of the member's rank, but the member itself only has ~40 directions to give. Any consensus built in a bert frame is capped near 40 regardless of teacher roster.

S0e — pooling settle (runs/alephlm0-s0e-*). Three identical dense trunks, one seed shared exactly (same init, same batch plan), differing only in readout, 500k rows × 2 epochs:

readout cos→target mimicry R@1
mean over mask .6037 .7745
CLS token .6147 .8180
learned-query attention .6033 .7680

CLS wins both gauges, outside the preregistered tie band (.003 cos / .01 R@1) — notable because the target is a mean-pooled object, and the attention readout (initialized to be exactly mean pooling) declined to move away from mean. Stage 1 therefore trains with the CLS readout.

Run status

run status
runs/alephlm0-s0e-{mean,cls,attn} ✅ complete
s0a/ erank instrument ✅ complete
runs/alephlm0-a2_dense-s0 ✅ complete — mimicry R@1 .9975, cos→target .8418, erank 99.2/768; 8-task capability .6026 (eval/), inside the captionbert-v2/-B band: the dense control is triple-replicated
runs/alephlm0-a3_random-s0 ✅ complete — mimicry .9980, cos→target .8392, erank 98.8; capability .6033 (band center: frozen-random routing matches dense at capacity parity); dispatch-OFF .5772 — the routed experts carry −.026 of task function, degrading gracefully (eval/)
runs/alephlm0-a1_anchored-s0 🔄 IN PROGRESS — probe passed 28.0 GB reserved / 4.9 GB margin (the reworked bank runs leaner than dense)
runs/alephlm0-{a1,a2,a3}-s1 queued

Each run directory carries checkpoints/ (state + rolling model snapshots + final_model.pt + metrics.json), config/ (the exact resolved configuration), and tensorboard/. Anchored runs additionally log per-block routing vitals at every eval: mean dispatched amplitude |w·σ(g)|, anchor drift from initialization, gate openings, and address-usage diversity — the curves that show the routing waking from its zero-initialized silence.

Lineage

  • Teachers: bert-base-uncased, ModernBERT-base, roberta-base, albert-base-v2, distilbert-base-uncased (mean-pooled, 512-token truncation)
  • Dense-recipe provenance: captionbert-8192-v2 (.6077 8-task STS mean, beating its best teacher at 13% of the combined teacher parameters) and its replication captionbert-8192-v2-B
  • The signed-address form and its training laws come from a long-running research program on geometric routing (AMOE); the amplitude-conservation result that motivates per-token signed dispatch was established on adapter collectives before being carried inward here.

Maintained as a live research log. Numbers in this card are measured, not projected; anything not yet measured is marked as such.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AbstractPhil/alephlm-0

Finetuned
(2375)
this model

Datasets used to train AbstractPhil/alephlm-0