Why we are featuring it
Speeding up inference by improving the draft.
DFlash tackles a concrete problem: producing useful parallel proposals to reduce the cost of generation. Its backends and benchmarks make it valuable for anyone studying model infrastructure.
What is inside
- Supports Qwen, Gemma, Llama, Kimi, MiniMax, and other models.
- Transformers, SGLang, vLLM, and MLX backends.
- Draft models published on Hugging Face.
- Benchmarks for GSM8K, Math500, HumanEval, MBPP, and MT-Bench.