--- license: apache-2.0 license_link: https://huggingface.co/Qwen/Qwen3.5-4B/blob/main/LICENSE library_name: transformers pipeline_tag: image-text-to-text base_model: Qwen/Qwen3.5-4B tags: - pyrodash - collaborative-decoding - llm-offload - qwen3.5 - sft language: - en - zh --- # PyroDash-4B-SFT --- This repository hosts **PyroDash-4B-SFT** — the **offload cold-start (Stage 2)** checkpoint of [PyroDash](https://github.com/PyroMind-Dynamics/pyroDash), fine-tuned from [Qwen/Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B). Companion models: [GRPO λ=0.05](https://huggingface.co/pyromind/PyroDash-4B-GRPO-Lambda-0.05) · [GRPO λ=0.6](https://huggingface.co/pyromind/PyroDash-4B-GRPO-Lambda-0.6) ---
Inference Architecture We propose **PyroDash**, a token-level dynamic reasoning paradigm for collaborative inference between small and large language models. PyroDash enables the small model to autonomously emit the control token `<|llm_offload|>` during autoregressive streaming decoding; the collaboration engine then dynamically offloads the local reasoning chain to a large model based on this control signal. This approach requires neither an additional router model nor retraining of the large model, and is naturally compatible with closed-source LLM services.
During training, PyroDash follows a three-stage progressive optimization pipeline: (1) train the control-token embedding layer so the small model acquires basic offloading expressiveness; (2) **cold-start the offload capability** (this checkpoint) to establish a collaboration pattern between the small and large models; and (3) apply GRPO reinforcement learning that jointly optimizes the dynamic offloading policy with a task-accuracy reward and a large-model call-cost penalty, achieving an adaptive balance between reasoning quality and compute cost.

Three-stage progressive training pipeline

## Model Details | Item | Value | | --- | --- | | Base model | [Qwen/Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B) | | Stage | (2) Offload cold-start SFT (LoRA → merged) | | Control token | <\|llm_offload\|> | | SFT dataset | [EasyHard-24K](https://huggingface.co/datasets/pyromind/easyhard-24k) | | Expert LLM (eval) | GLM-5.2-FP8 | | Precision | bfloat16 | ## Quick Start ### 1. Setup ```bash git clone https://github.com/PyroMind-Dynamics/pyroDash.git cd pyroDash pip install -r requirements.txt ``` ### 2. Run evaluation (`evaluation/math_eval.sh`) Edit placeholders in [`evaluation/math_eval.sh`](https://github.com/PyroMind-Dynamics/pyroDash/blob/main/evaluation/math_eval.sh), then: ```bash bash evaluation/math_eval.sh ``` The script (1) starts a local **vLLM** server for the small model on port `8001`, (2) runs `math_eval.py`, and (3) stops vLLM on exit. #### Parameters | Variable / flag | Meaning | Example | |-----------------|---------|---------| | `MODEL` | Local merged model path (vLLM serve + tokenizer) | `/path/to/your/merged_model` | | `--glm-base-url` | OpenAI-compatible API for the large/relay model | `http://your-glm-host:8000/v1` | | `--glm-api-key` | API key for that endpoint | `your-glm-api-key` | | `--glm-model` | Served model name on the GLM side | `your-glm-model` | | `--output-dir` | Per-dataset JSON output directory | `./results_500` | | `--datasets` | Benchmarks (space-separated) | `gsm8k minerva olympiad aime2024 aime2025` | Tokenizer must include the special token `<|llm_offload|>`. ## Results

Cost–Accuracy Pareto

| Method | Avg. Acc. (%) | LLM Token Ratio (%) | Avg. LLM Calls | Cost ($) | | --- | ---: | ---: | ---: | ---: | | Qwen3.5-4B | 28.36 | 0.00 | 0.000 | 2.26 | | **Qwen3.5-4B (+SFT) ← this** | **46.25** | **0.00** | **0.000** | **1.32** | | RouteLLM (~75% GLM-5.2-FP8) | 52.74 | 77.37 | 0.808 | 44.62 | | GlimpRouter (τ=0.9) | 54.20 | 75.11 | 1.20 | 31.61 | | PyroDash (λ=0.1) | 55.29 | 8.19 | 0.058 | 4.71 | | PyroDash (λ=0.6) | 54.55 | 1.90 | 0.012 | 1.78 | | PyroDash (λ=0.05) | 64.04 | 95.34 | 0.975 | 39.29 | | GLM-5.2-FP8 | 57.68 | 100.00 | 1.000 | 49.36 | ## Resources | Resource | Link | | --- | --- | | Project website | [PyroMind-Dynamics.github.io/pyroDash](https://PyroMind-Dynamics.github.io/pyroDash/) | | Code | [github.com/PyroMind-Dynamics/pyroDash](https://github.com/PyroMind-Dynamics/pyroDash) | | SFT dataset (EasyHard-24K) | [huggingface.co/datasets/pyromind/easyhard-24k](https://huggingface.co/datasets/pyromind/easyhard-24k) | | Hugging Face org | [huggingface.co/pyromind](https://huggingface.co/pyromind) | ## Citation ```bibtex @misc{pyrodash2026, title = {PyroDash: Cost-Efficient Token-Level Small-Large Model Collaborative Inference}, author = {{PyroMind Dynamics}}, year = {2026}, note = {Preprint} } @misc{pyromind2026easyhard24k, title = {{EasyHard-24K} v0.02}, author = {{PyroMind Dynamics}}, year = {2026}, howpublished = {\url{https://huggingface.co/datasets/pyromind/easyhard-24k}} } ``` ## License Apache 2.0 (derived from [Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B)).