Speculative-decoding subfolder test fixture

A tiny, randomly-initialised Llama pair for exercising subfolder loading of a speculative draft model. It exists so inference engines can test the code path in CI without downloading a real checkpoint.

These weights are random. The model produces gibberish. That is the point.

.              <- target model, 2 layers, ~4M params
โ””โ”€โ”€ DFlash/    <- draft model,  1 layer,  ~4M params

Both share the same tokenizer and vocabulary, so the draft is a valid speculator for the target.

Why a subfolder needs engine support

A Hugging Face repo id may contain at most one /, so org/repo/DFlash is not a valid id and cannot be passed as a single string. A draft bundled inside its target's repository therefore has to be addressed as a repo id plus a separate sub-path, which most engines have no parameter for.

This matters in practice: shipping a draft inside its target repo keeps the pair versioned together and downloadable in one fetch, rather than as two repos that can drift apart.

Usage

from huggingface_hub import snapshot_download
snapshot_download("poolside/spec-decoding-subfolder-fixture")     # target + draft
snapshot_download("poolside/spec-decoding-subfolder-fixture",
                  allow_patterns=["DFlash/*"])                    # draft only

Engines adding support should accept the draft as (model="poolside/spec-decoding-subfolder-fixture", subfolder="DFlash").

Provenance

Built with transformers from LlamaConfig (hidden 64, intermediate 128, 4 heads, tied embeddings), torch.manual_seed(0). Tokenizer copied from hf-internal-testing/llama-tokenizer. No trained parameters, no proprietary data.

Downloads last month
-
Safetensors
Model size
2.13M params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support