开源项目

dflash

dflash

轻量级块扩散模型,专为投机性解码设计,实现高效并行草稿生成以加速LLM推理。支持Gemma、Qwen、Llama等多种主流模型,可与vLLM、SGLang等推理引擎无缝集成,并提供了MLX(Apple Silicon)支持。开源训练配方即将发布,便于用户自定义加速。研究向,非生产级建议。

README

DFlash:面向闪速推测解码的块扩散

论文 | 博客 | 模型

DFlash 是一种轻量级的块扩散模型,专为推测解码设计。它能够实现高效且高质量的并行草稿生成。

DFlash 架构

https://github.com/user-attachments/assets/5b29cabb-eb95-44c9-8ffe-367c0758de8c

支持的模型

模型 DFlash 草稿
gemma-4-26B-A4B-it z-lab/gemma-4-26B-A4B-it-DFlash
gemma-4-31B-it z-lab/gemma-4-31B-it-DFlash
Qwen3.6-27B z-lab/Qwen3.6-27B-DFlash
Qwen3.6-35B-A3B z-lab/Qwen3.6-35B-A3B-DFlash
MiniMax-M2.5 (Preview) z-lab/MiniMax-M2.5-DFlash
Kimi-K2.5 z-lab/Kimi-K2.5-DFlash
Qwen3.5-4B z-lab/Qwen3.5-4B-DFlash
Qwen3.5-9B z-lab/Qwen3.5-9B-DFlash
Qwen3.5-27B z-lab/Qwen3.5-27B-DFlash
Qwen3.5-35B-A3B z-lab/Qwen3.5-35B-A3B-DFlash
Qwen3.5-122B-A10B z-lab/Qwen3.5-122B-A10B-DFlash
Qwen3-Coder-Next z-lab/Qwen3-Coder-Next-DFlash
Qwen3-Coder-30B-A3B z-lab/Qwen3-Coder-30B-A3B-DFlash
gpt-oss-20b z-lab/gpt-oss-20b-DFlash
gpt-oss-120b z-lab/gpt-oss-120b-DFlash
Qwen3-4B(非思考模式) z-lab/Qwen3-4B-DFlash-b16
Qwen3-8B(非思考模式) z-lab/Qwen3-8B-DFlash-b16
Llama-3.1-8B-Instruct z-lab/LLaMA3.1-8B-Instruct-DFlash-UltraChat
DeepSeek-V4-Flash 即将推出
DeepSeek-V4-Pro 即将推出
MiniMax-M2.7 即将推出
GLM-5.1 即将推出

欢迎随时提交 GitHub issue 请求支持更多模型。我们也将很快开源训练配方,以便您可以训练自己的 DFlash 草稿模型来加速任何 LLM。

📦 安装

为每个后端使用单独的虚拟环境以避免冲突。

后端 安装命令
Transformers uv pip install -e ".[transformers]"
SGLang uv pip install -e ".[sglang]"
vLLM 见下方
MLX(Apple Silicon) pip install -e ".[mlx]"

vLLM: vLLM v0.20.1+ 包含核心 DFlash 支持。对大部分模型使用标准安装:

uv pip install -e ".[vllm]"

Gemma4 DFlash 当前需要我们的临时 vLLM Gemma4 构建版本。推荐使用 Docker:

docker pull ghcr.io/z-lab/vllm-openai:gemma4-dflash-cu130

Gemma4 的源码回退方案:

uv pip install -U --torch-backend=auto \
  "vllm @ git+https://github.com/vllm-project/vllm.git@refs/pull/41703/head"

较新的非 Gemma4 SWA 草稿模型使用 SWA 支持分支:

uv pip install -U --torch-backend=auto \
  "vllm @ git+https://github.com/vllm-project/vllm.git@refs/pull/40898/head"

🚀 快速上手

vLLM

使用 Docker 运行 Gemma4:

docker run --rm -it \
  --gpus all \
  --ipc=host \
  --shm-size=16g \
  -p 8000:8000 \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  ghcr.io/z-lab/vllm-openai:gemma4-dflash-cu130 \
  google/gemma-4-26B-A4B-it \
  --host 0.0.0.0 \
  --port 8000 \
  --speculative-config '{"method": "dflash", "model": "z-lab/gemma-4-26B-A4B-it-DFlash", "num_speculative_tokens": 15, "attention_backend": "flash_attn"}' \
  --attention-backend triton_attn \
  --max-num-batched-tokens 32768 \
  --trust-remote-code

非 Gemma4 模型:

vllm serve Qwen/Qwen3.5-27B \
  --speculative-config '{"method": "dflash", "model": "z-lab/Qwen3.5-27B-DFlash", "num_speculative_tokens": 15}' \
  --attention-backend flash_attn \
  --max-num-batched-tokens 32768

SGLang

export SGLANG_ALLOW_OVERWRITE_LONGER_CONTEXT_LEN=1

# 可选:启用调度重叠(实验性,可能不稳定)
# export SGLANG_ENABLE_SPEC_V2=1
# export SGLANG_ENABLE_DFLASH_SPEC_V2=1
# export SGLANG_ENABLE_OVERLAP_PLAN_STREAM=1

python -m sglang.launch_server \
    --model-path Qwen/Qwen3.5-35B-A3B \
    --speculative-algorithm DFLASH \
    --speculative-draft-model-path z-lab/Qwen3.5-35B-A3B-DFlash \
    --speculative-num-draft-tokens 16 \
    --tp-size 1 \
    --attention-backend trtllm_mha \
    --speculative-draft-attention-backend fa4 \
    --mem-fraction-static 0.75 \
    --mamba-scheduler-strategy extra_buffer \
    --trust-remote-code

Transformers

仅 Qwen3 和 LLaMA-3.1 模型支持 Transformers 后端。

from transformers import AutoModel, AutoModelForCausalLM, AutoTokenizer

draft = AutoModel.from_pretrained("z-lab/Qwen3-8B-DFlash-b16", trust_remote_code=True, dtype="auto", device_map="cuda:0").eval()
target = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-8B", dtype="auto", device_map="cuda:0").eval()
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen3-8B")

messages = [{"role": "user", "content": "How many positive whole-number divisors does 196 have?"}]
input_ids = tokenizer.apply_chat_template(messages, return_tensors="pt", add_generation_prompt=True, enable_thinking=False).to(draft.device)

output = draft.spec_generate(input_ids=input_ids, max_new_tokens=2048, temperature=0.0, target=target, stop_token_ids=[tokenizer.eos_token_id])
print(tokenizer.decode(output[0], skip_special_tokens=False))

MLX(Apple Silicon)

社区已经在 MLX 上实现了许多优秀的 DFlash 实现;这里我们提供一个简单且高效的版本,已在 Apple M5 Pro 上使用 Qwen3、Qwen3.5 和 Gemma-4 模型测试通过。

from dflash.model_mlx import load, load_draft, stream_generate

model, tokenizer = load("Qwen/Qwen3.5-4B")
draft = load_draft("z-lab/Qwen3.5-4B-DFlash")

messages = [{"role": "user", "content": "How many positive whole-number divisors does 196 have?"}]
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True, enable_thinking=True)
tps = 0.0
for r in stream_generate(model, draft, tokenizer, prompt, block_size=16, max_tokens=2048, temperature=0.6):
    print(r.text, end="", flush=True)
    tps = r.generation_tps
print(f"\nThroughput: {tps:.2f} tok/s")

📊 评估

所有基准测试共享相同的数据集(gsm8k、math500、humaneval、mbpp、mt-bench)。数据集在首次运行时自动下载并缓存为 JSONL 格式,存放于 cache/ 目录。

vLLM:

python -m dflash.benchmark --backend vllm \
    --base-url http://127.0.0.1:8000 --model Qwen/Qwen3.5-27B \
    --dataset gsm8k --num-prompts 128 --concurrency 1 --enable-thinking

SGLang:

python -m dflash.benchmark --backend sglang \
    --base-url http://127.0.0.1:30000 --model Qwen/Qwen3.5-35B-A3B \
    --dataset gsm8k --num-prompts 128 --concurrency 1 --enable-thinking

Transformers(仅 Qwen3 和 LLaMA):

torchrun --nproc_per_node=8 -m dflash.benchmark --backend transformers \
    --model Qwen/Qwen3-8B --draft-model z-lab/Qwen3-8B-DFlash-b16 \
    --dataset gsm8k --max-samples 128

MLX:

python -m dflash.benchmark --backend mlx \
    --model mlx-community/gemma-4-31b-it-4bit --draft-model z-lab/gemma-4-31B-it-DFlash \
    --dataset gsm8k --max-samples 128 --enable-thinking

致谢

衷心感谢 @dcw02、@gongy 以及 @modal-labs 团队在将 DFlash 适配到 SGLang 过程中提供的快速且高质量的支持。同样衷心感谢 NVIDIA 的 @benchislett 在将 DFlash 适配到 vLLM 并使其惠及更广泛的服务社区方面所做的工作。

引用

如果您觉得 DFlash 有用,请引用我们的工作。如需分享对 DFlash 的反馈或请求新模型支持,请填写此表单:DFlash 反馈。

@article{chen2026dflash,
  title   = {{DFlash: Block Diffusion for Flash Speculative Decoding}},
  author  = {Chen, Jian and Liang, Yesheng and Liu, Zhijian},
  journal = {arXiv preprint arXiv:2602.06036},
  year    = {2026}
}
开源项目z-lab2026-05-07原文

相关内容