Any-Depth Alignment (ADA) β€” Linear Probes (ADA-LP)

Pre-trained ADA-LP probes for Any-Depth Alignment: Unlocking the Innate Safety Alignment of LLMs to Any Depth (ICLR 2026). Code: Any-Depth-Alignment Β· Data: javyduck/any-depth-alignment.

Each probe is a scikit-learn LogisticRegression trained on the hidden state of an injected Safety Token (the assistant header) at one layer. At inference, ADA-LP re-injects the Safety Tokens mid-generation, reads that single hidden state, and applies the probe to decide whether to halt β€” turning the base model into its own guardrail with constant overhead and no weight updates.

What's inside

Probes for 12 models, covering every layer Γ— Safety-Token Γ— hook-position in the paper's ablations (~3.6k .joblib files). Path layout (mirrors the training code):

ckpts/{model_slug}/{safety_slug}/mask_token_none/{hook_slug}/gradual_cache/seed_42/logistic/layer_{L}.joblib

The canonical per-model probe (Safety Token, layer) matches the model registry in the code repo, e.g. google/gemma-2-9b-it β†’ layer 23, meta-llama/Llama-3.1-8B-Instruct β†’ layer 15, deepseek-ai/DeepSeek-R1-Distill-Qwen-7B β†’ layer 13, openai/gpt-oss-120b β†’ layer 33.

Usage

import joblib
probe = joblib.load("ckpts/google_gemma-2-9b-it/.../logistic/layer_23.joblib")
prob_harmful = probe.predict_proba(safety_token_hidden_state)[:, 1]

In the code repo this is fully automated β€” python -m ada.probe.evaluate --model google/gemma-2-9b-it --dataset advbench loads the right probe from the registry.

Citation

@inproceedings{zhang2026anydepth,
  title     = {Any-Depth Alignment: Unlocking the Innate Safety Alignment of LLMs to Any Depth},
  author    = {Zhang, Jiawei and Estornell, Andrew and Li, Bo and Baek, David D. and Xu, Xiaojun},
  booktitle = {ICLR},
  year      = {2026}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support