Back to repositories
z-labDFlash

DFlash: open source worth watching

A lightweight block-diffusion model for speculative decoding, designed for efficient parallel drafting.

Why we are featuring it

Speeding up inference by improving the draft.

DFlash tackles a concrete problem: producing useful parallel proposals to reduce the cost of generation. Its backends and benchmarks make it valuable for anyone studying model infrastructure.

What is inside

  • Supports Qwen, Gemma, Llama, Kimi, MiniMax, and other models.
  • Transformers, SGLang, vLLM, and MLX backends.
  • Draft models published on Hugging Face.
  • Benchmarks for GSM8K, Math500, HumanEval, MBPP, and MT-Bench.
DFlash | Trending Repo | DevOP