百科.dev
全部条目AI 编程趋势榜开源项目技术资讯提交条目
登录
< 返回工具列表
D

dflash

> 编程语言
开源

DFlash: Flash 猜测解码的区块扩散

5.6K stars0 点赞0 次浏览
访问官网GitHub

工具介绍

DFlash: Flash 猜测解码的区块扩散

# DFlash: Block Diffusion for Flash Speculative Decoding **DFlash** is a lightweight **block diffusion** model designed for speculative decoding. It enables efficient and high-quality parallel drafting. DFlash 2 [**Blog**](https://inco.ai/blog/dflash2/) | [**Models**](https://huggingface.co/collections/z-lab/dflash-2)

https://github.com/user-attachments/assets/f786e7c5-c2bc-47d4-8a32-1f730a689e1b DFlash [**Paper**](https://arxiv.org/abs/2602.06036) | [**Blog**](https://z-lab.ai/projects/dflash/) | [**Models**](https://huggingface.co/collections/z-lab/dflash) https://github.com/user-attachments/assets/5b29cabb-eb95-44c9-8ffe-367c0758de8c ## Supported Models ### DFlash 2 Available checkpoints: [Muse-Glimmer-30B](https://huggingface.co/z-lab/Muse-Glimmer-30B-DFlash2) and [Qwen3.8-27B](https://huggingface.co/z-lab/Qwen3.8-27B-DFlash2). See the [DFlash 2 collection](https://huggingface.co/collections/z-lab/dflash-2) for updates. ### DFlash Public checkpoints are available in the [DFlash collection](https://huggingface.co/collections/z-lab/dflash): - **Qwen:** Qwen3.6 (27B, 35B-A3B), Qwen3.5 (4B, 9B, 27B, 35B-A3B, 122B-A10B, 397B-A17B), Qwen3 (4B/8B non-thinking, Coder-Next, Coder-30B-A3B) - **Gemma:** Gemma 4 (12B, 31B, 26B-A4B) - **MiniMax:** M2.5, M2.7 - **Kimi:** K2.5, K2.6, K2.7-Code - **Others:** GPT-OSS (20B, 120B), Llama-3.1-8B, GLM 5.1, Alpamayo 1.5/R1 10B Use the Transformers or MLX backends below for their explicitly listed model families. Other checkpoints can be benchmarked through an OpenAI-compatible SGLang or vLLM server. ## Installation Install the base package for an OpenAI-compatible server, or include local inference dependencies. The local install uses MLX on Apple Silicon and Transformers on Linux. ```bash pip install dflash pip install "dflash[local]" # local inference ``` For serving benchmarks, install a supported version of [SGLang](https://github.com/sgl-project/sglang/pull/35371), [vLLM](https://github.com/vllm-project/vllm/pull/52816), [oMLX](https://github.com/z-lab/omlx-fork/releases/download/0.6.2-dflash2/oMLX-0.6.2-zlab-dflash2-arm64-signed.dmg), or [llama.cpp](https://github.com/ggml-org/llama.cpp/pull/27342) separately, launch its OpenAI-compatible server with DFlash, and pass its `--base-url` below. ## Quick Start ### Transformers The Transformers backend supports DFlash 2 for Muse-Glimmer-30B, and DFlash for Qwen3 and LLaMA-3.1-8B. Muse uses `reasoning_strength`: `low`, `medium`, `high` (default), or `xhigh`. ```bash dflash generate transformers \ --model meta-models/Muse-Glimmer-30B \ --draft z-lab/Muse-Glimmer-30B-DFlash2 \ --reasoning high --temperature 1 --top-p 0.95 --top-k 64 \ "How many positive whole-number divisors does 196 have?" ``` ### MLX (Apple Silicon) The MLX backend supports DFlash 2 for Qwen3.8-27B, and DFlash for Qwen3, Qwen3.5, Qwen3.6, and Gemma 4. Qwen3.8 uses `reasoning_effort`: `low`, `medium`, or `xhigh` (default). For quantized targets or drafts, use `block_size <= 5`: MLX's current quantized matmul kernel becomes less efficient at larger verify widths. The example below runs both the target and draft with 4-bit weights. ```bash dflash generate mlx \ --model mlx-community/Qwen3.8-27B-4bit \ --draft z-lab/Qwen3.8-27B-DFlash2 \ --draft-bits 4 --block-size 5 --reasoning xhigh \ "How many positive whole-number divisors does 196 have?" ``` ### OpenAI-compatible server Launch the latest SGLang or vLLM server separately, then run: ```bash dflash generate openai \ --base-url http://127.0.0.1:8000 --model Qwen/Qwen3.8-27B \ "How many positive whole-number divisors does 196 have?" ``` ## Evaluation All benchmarks share the same datasets (gsm8k, math500, humaneval, mbpp, mt-bench), downloaded and cached by Hugging Face Datasets. **OpenAI-compatible server** (SGLang or vLLM): ```bash dflash benchmark openai \ --base-url http://127.0.0.1:8000 --model Qwen/Qwen3.8-27B \ --dataset gsm8k --num-prompts 128 --concurrency 1 --reasoning xhigh \ --temperature 1 --top-p 0.95 --top-k 20 ``` **Transformers** (Muse-Glimmer-30B DFlash 2): ```bash dflash benchmark transformers \ --model meta-models/Muse-Glimmer-30B --draft z-lab/Muse-Glimmer-30B-DFlash2 \ --dataset gsm8k --max-samples 128 --reasoning high ``` **MLX** (Qwen3.8-27B 4-bit DFlash 2): ```bash dflash benchmark mlx \ --model mlx-community/Qwen3.8-27B-4bit --draft z-lab/Qwen3.8-27B-DFlash2 \ --dataset gsm8k --max-samples 128 --reasoning xhigh --block-size 5 --draft-bits 4 ``` ## Acknowledgement Huge thanks to [@dcw02](https://github.com/dcw02), [@gongy](https://github.com/gongy), and the team at [@modal-labs](https://github.com/modal-labs) for their fast, high-quality support in bringing DFlash to SGLang. And huge thanks as well to [@benchislett](https://github.com/benchislett) at NVIDIA for his work in bringing DFlash to vLLM and helping make it available to the broader serving community. ## Citation If you find DFlash useful, please cite our work. To share feedback on DFlash or request new model support, please fill out this form: [DFlash Feedback](https://forms.gle/4YNwfqb4nJdqn6hq9). ```bibtex @article{chen2026dflash, title = {{DFlash: Block Diffusion for Flash Speculative Decoding}}, author = {Chen, Jian and Liang, Yesheng and Liu, Zhijian}, journal = {arXiv preprint arXiv:2602.06036}, year = {2026} } @misc{inco2026dflash2, title = {DFlash 2: Keep Drafting Parallel}, author = {{Inco AI}}, year = {2026}, month = {August}, url = {https://inco.ai/blog/dflash2/} } ```

Issues· 0 开放

查看全部 Issues在 GitHub 打开

暂无开放 Issues,或尚未同步最近议题。

> 标签

Python

暂无评论,来聊聊你的看法吧

> 工具信息

发布日期2026年8月1日
最后更新2026年9月17日
分类编程语言
定价开源

> 相关工具

T
TypeScript
JavaScript 的超集,为前端与全栈提供静态类型
P
Python
通用编程语言,广泛用于 Web、数据与 AI
G
Go
Google 推出的简洁高效系统语言