DeepSpec
View on GitHubDeepSpec: a full-stack codebase for training and evaluating speculative decoding algorithms
DeepSpec is a full-stack Python codebase for training and evaluating speculative-decoding draft models (DSpark, DFlash, Eagle3) against targets like Qwen3 and Gemma. It covers data prep, 8-GPU training, and benchmark evaluation, plus released checkpoints.
Use Cases
Train draft models for speculative decodingAccelerate LLM inference throughput and latencyEvaluate speculative-decoding acceptance ratesReproduce Eagle3, DFlash and DSpark checkpointsBuild target-answer caches from promptsBenchmark draft models on GSM8K, MATH500, AIME, HumanEval, MBPP, LiveCodeBench, MT-BenchFine-tune draft models for domain-specific targetsMulti-GPU distributed draft model training
Built With
- Language
- Python
- Frameworks
- PyTorch · Hugging Face Transformers · Triton · TensorBoard · Safetensors · Hugging Face Datasets · SpecForge
Tags
speculative-decoding · llm-inference · inference-optimization · draft-model · model-training · eagle3 · dflash · dspark · pytorch · triton · distributed-training · llm-evaluation · throughput · latency · transformers · benchmarking