Model Alpha — Qwen3-8B fine-tuned for Solidity vulnerability detection

Overview

Model Alpha is a fine-tuned version of Qwen3-8B specialized in identifying high-severity vulnerabilities in Solidity smart contracts. It is served as an OpenAI-compatible endpoint and powers a public audit pipeline at https://openai.vast.ai/model-alpha.

The weights released here are the same AWQ pack-quantized int4 weights running in production, byte-identical to the deployed endpoint. The format is compressed-tensors (pack-quantized, W4A16, group_size=128).

Training data

Trained on approximately 3,291 labelled examples (2,961 train / 330 test) drawn from public smart contract audits across the following categories:

  • DeFi protocols — lending markets, DEXs (AMM + order-book), staking, vaults, yield aggregators
  • NFT and gaming — marketplaces, royalty engines, on-chain randomness, claim flows
  • Cross-chain bridges — message passing, signature verification, replay protection
  • Oracles and price feeds — TWAP, Chainlink integration, fallback handling
  • Account abstraction and wallets — session keys, paymasters, signature replay
  • Stablecoins and synthetic assets — mint/burn, collateral, liquidation
  • Governance and DAOs — voting, timelock, proposal execution

Examples were constructed by combining function-level Solidity source with structured call-graph context and labelled by audit methodology (vulnerability type, severity, dollar impact where applicable).

Intended use

  • Smart contract security auditing (defensive / white-hat)
  • Vulnerability triage during code review
  • Research on LLM-based code analysis for adversarial code

Not intended for: identifying exploits for offensive use, generating malicious contracts, or auditing contracts outside EVM-compatible chains.

Evaluation

Wake Arena v8 (94 high-severity vulnerabilities, Ackee Blockchain)

Setting Result
Primary function only, 3 samples, 2/3 majority vote 69 / 94 = 73.4%
All sibling functions, 3 samples, majority (ceiling) 88 / 94 = 93.6%
Pashov skill (published external measurement, @0xTomass) 49 / 94 = 52.1%
Opus 4.7 MAX (same external measurement) 48 / 94 = 51.1%

Methodology documented in evaluation_results/wake_arena/. Same-model Alpha scores range 21.3% → 93.6% depending on prompt format and sample count.

evmbench (117 high-severity vulnerabilities, OpenAI Frontier Evals)

Metric Alpha Claude Fable GLM-5.1
Detected 6 / 117 16 / 117 15 / 117
Total $ award on detected $20,691 $2,422 $646

Alpha's single highest-value catch — Sequence H-02 ($20,367) — was a partial-signature replay vulnerability that neither Fable nor GLM-5.1 detected under the same pipeline. That single find is worth 8.4× Fable's entire 16-catch portfolio. Source: evmbench_full_alpha_20260610_021252.json.

Specifications

  • Base model: unsloth/Qwen3-8B-bnb-4bit
  • Architecture: Qwen3ForCausalLM (36 layers, hidden_size 4096)
  • Context length: 16,384 tokens (production), up to 40,960 (architecture max)
  • Quantization: AWQ W4A16, pack-quantized, group_size=128, via compressed-tensors
  • Disk size: ~5.7 GB
  • Serving: vLLM with --quantization compressed-tensors

Usage

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "l-lyubenov/model-alpha"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto",
)

prompt = """Audit this Solidity function for high-severity vulnerabilities:

function withdraw(uint256 amount) external {
    uint256 balance = balances[msg.sender];
    require(amount <= balance, "insufficient");
    (bool ok, ) = msg.sender.call{value: amount}("");
    require(ok, "transfer failed");
    balances[msg.sender] = balance - amount;
}
"""
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=512, do_sample=False)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Or via vLLM:

vllm serve l-lyubenov/model-alpha \
  --quantization compressed-tensors \
  --max-model-len 16384 \
  --gpu-memory-utilization 0.90

Limitations

  • Trained primarily on EVM-compatible Solidity; performance on Vyper, Move, Cairo is unknown
  • May produce false positives on safe patterns (assembly blocks, low-level calls, deliberate reentrancy)
  • High flag rate with moderate precision — recommended as a triage tool, not a final arbiter
  • Single-turn chat only; not optimized for multi-turn code review dialogue

License

Public release — weights are available to anyone with a Hugging Face account. Use is restricted to defensive security purposes (auditing, vulnerability research on authorized systems, academic study, tooling that identifies or remediates vulnerabilities). Commercial use, redistribution as a hosted service, and offensive use are not granted by this release. See LICENSE for the full terms.

Citation

If you use Model Alpha in published work, please cite the public benchmark results:

@software{model_alpha_2026,
  title  = {Model Alpha: Qwen3-8B fine-tuned for Solidity vulnerability detection},
  author = {Lyubenov, Lyuboslav},
  year   = {2026},
  url    = {https://huggingface.co/l-lyubenov/model-alpha}
}
Downloads last month
98
Safetensors
Model size
2B params
Tensor type
BF16
·
I32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for l-lyubenov/model-alpha

Finetuned
Qwen/Qwen3-8B
Quantized
(4)
this model