PyroDash-4B-SFT / README.md
huoyunhf's picture
Upload folder using huggingface_hub
cdd7bbe verified
|
Raw
History Blame Contribute Delete
5.64 kB
---
license: apache-2.0
license_link: https://huggingface.co/Qwen/Qwen3.5-4B/blob/main/LICENSE
library_name: transformers
pipeline_tag: image-text-to-text
base_model: Qwen/Qwen3.5-4B
tags:
- pyrodash
- collaborative-decoding
- llm-offload
- qwen3.5
- sft
language:
- en
- zh
---
# PyroDash-4B-SFT
---
This repository hosts **PyroDash-4B-SFT** — the **offload cold-start (Stage 2)** checkpoint of [PyroDash](https://github.com/PyroMind-Dynamics/pyroDash), fine-tuned from [Qwen/Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B).
Companion models: [GRPO λ=0.05](https://huggingface.co/pyromind/PyroDash-4B-GRPO-Lambda-0.05) · [GRPO λ=0.6](https://huggingface.co/pyromind/PyroDash-4B-GRPO-Lambda-0.6)
---
<table>
<tr>
<td width="50%" valign="top">
<img src="https://raw.githubusercontent.com/PyroMind-Dynamics/pyroDash/main/docs/assets/inference.png" alt="Inference Architecture" width="100%"/>
</td>
<td width="50%" valign="middle">
We propose **PyroDash**, a token-level dynamic reasoning paradigm for collaborative inference between small and large language models. PyroDash enables the small model to autonomously emit the control token `<|llm_offload|>` during autoregressive streaming decoding; the collaboration engine then dynamically offloads the local reasoning chain to a large model based on this control signal. This approach requires neither an additional router model nor retraining of the large model, and is naturally compatible with closed-source LLM services.
</td>
</tr>
</table>
During training, PyroDash follows a three-stage progressive optimization pipeline: (1) train the control-token embedding layer so the small model acquires basic offloading expressiveness; (2) **cold-start the offload capability** (this checkpoint) to establish a collaboration pattern between the small and large models; and (3) apply GRPO reinforcement learning that jointly optimizes the dynamic offloading policy with a task-accuracy reward and a large-model call-cost penalty, achieving an adaptive balance between reasoning quality and compute cost.
<p align="center">
<img src="https://raw.githubusercontent.com/PyroMind-Dynamics/pyroDash/main/docs/assets/training.png" alt="Three-stage progressive training pipeline" width="80%"/>
</p>
## Model Details
| Item | Value |
| --- | --- |
| Base model | [Qwen/Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B) |
| Stage | (2) Offload cold-start SFT (LoRA → merged) |
| Control token | <code>&lt;\|llm_offload\|&gt;</code> |
| SFT dataset | [EasyHard-24K](https://huggingface.co/datasets/pyromind/easyhard-24k) |
| Expert LLM (eval) | GLM-5.2-FP8 |
| Precision | bfloat16 |
## Quick Start
### 1. Setup
```bash
git clone https://github.com/PyroMind-Dynamics/pyroDash.git
cd pyroDash
pip install -r requirements.txt
```
### 2. Run evaluation (`evaluation/math_eval.sh`)
Edit placeholders in [`evaluation/math_eval.sh`](https://github.com/PyroMind-Dynamics/pyroDash/blob/main/evaluation/math_eval.sh), then:
```bash
bash evaluation/math_eval.sh
```
The script (1) starts a local **vLLM** server for the small model on port `8001`, (2) runs `math_eval.py`, and (3) stops vLLM on exit.
#### Parameters
| Variable / flag | Meaning | Example |
|-----------------|---------|---------|
| `MODEL` | Local merged model path (vLLM serve + tokenizer) | `/path/to/your/merged_model` |
| `--glm-base-url` | OpenAI-compatible API for the large/relay model | `http://your-glm-host:8000/v1` |
| `--glm-api-key` | API key for that endpoint | `your-glm-api-key` |
| `--glm-model` | Served model name on the GLM side | `your-glm-model` |
| `--output-dir` | Per-dataset JSON output directory | `./results_500` |
| `--datasets` | Benchmarks (space-separated) | `gsm8k minerva olympiad aime2024 aime2025` |
Tokenizer must include the special token `<|llm_offload|>`.
## Results
<p align="center">
<img src="https://raw.githubusercontent.com/PyroMind-Dynamics/pyroDash/main/docs/assets/fig_cost_accuracy_pareto.png" alt="Cost–Accuracy Pareto" width="70%"/>
</p>
| Method | Avg. Acc. (%) | LLM Token Ratio (%) | Avg. LLM Calls | Cost ($) |
| --- | ---: | ---: | ---: | ---: |
| Qwen3.5-4B | 28.36 | 0.00 | 0.000 | 2.26 |
| **Qwen3.5-4B (+SFT) ← this** | **46.25** | **0.00** | **0.000** | **1.32** |
| RouteLLM (~75% GLM-5.2-FP8) | 52.74 | 77.37 | 0.808 | 44.62 |
| GlimpRouter (τ=0.9) | 54.20 | 75.11 | 1.20 | 31.61 |
| PyroDash (λ=0.1) | 55.29 | 8.19 | 0.058 | 4.71 |
| PyroDash (λ=0.6) | 54.55 | 1.90 | 0.012 | 1.78 |
| PyroDash (λ=0.05) | 64.04 | 95.34 | 0.975 | 39.29 |
| GLM-5.2-FP8 | 57.68 | 100.00 | 1.000 | 49.36 |
## Resources
| Resource | Link |
| --- | --- |
| Project website | [PyroMind-Dynamics.github.io/pyroDash](https://PyroMind-Dynamics.github.io/pyroDash/) |
| Code | [github.com/PyroMind-Dynamics/pyroDash](https://github.com/PyroMind-Dynamics/pyroDash) |
| SFT dataset (EasyHard-24K) | [huggingface.co/datasets/pyromind/easyhard-24k](https://huggingface.co/datasets/pyromind/easyhard-24k) |
| Hugging Face org | [huggingface.co/pyromind](https://huggingface.co/pyromind) |
## Citation
```bibtex
@misc{pyrodash2026,
title = {PyroDash: Cost-Efficient Token-Level Small-Large Model Collaborative Inference},
author = {{PyroMind Dynamics}},
year = {2026},
note = {Preprint}
}
@misc{pyromind2026easyhard24k,
title = {{EasyHard-24K} v0.02},
author = {{PyroMind Dynamics}},
year = {2026},
howpublished = {\url{https://huggingface.co/datasets/pyromind/easyhard-24k}}
}
```
## License
Apache 2.0 (derived from [Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B)).