Lecture 22: Hacker's Guide to Speculative Decoding in VLLM

Опубликовано: 08 Февраль 2026
на канале: GPU MODE
11,512
217

Abstract: We will discuss how vLLM combines continuous batching with speculative decoding with a focus on enabling external contributors. Topics include proposer/scorer/verifier framework, proposal methods, lookahead scheduling, dynamic speculative decoding, and future contribution ideas.

Speaker: Cade Daniel

Slides: https://docs.google.com/presentation/...