What Is Speculative Decoding? (ANSWERED)
Model question on speculative decoding — draft model, target verification, lossless speedup, and production deployment notes.
TL;DR — Quick Answer
Speculative decoding uses a small draft model to propose several tokens quickly; the large target model verifies them in parallel in one forward pass. Accepted tokens advance faster than sequential decoding — lossless vs target-only sampling when verification matches exact decoding rules. Speedup depends on acceptance rate (draft quality and alignment). Used in inference servers (vLLM, TensorRT-LLM) for latency-sensitive serving — not a replacement for model routing or quantization.
The Interview Question
What is speculative decoding and how does it speed up LLM inference without changing outputs?
Deep Explanation
Sign in to unlock full answer
Get deep explanations, PDF export & all LLMs questions
- 12 more sections of deep explanation
- Real-world examples
- Common mistakes
- Interviewer expectations
- Follow-up questions
Speculative DecodingInferenceLatencyDraft ModelGoogleMetaNVIDIA