Speculative Decoding, Beyond Text
A complete mental model for replacing several expensive serial decode steps with cheap guesses and one parallel verification pass—without changing the target model’s output law when the exact algorithm is used. The same idea reaches language, vision-language, speech recognition, audio codecs, discrete image/video generators, music, motion, proteins, and action models. What changes is the token, conditioning, and state contract.
- An autoregressive target still decides every committed token; a cheaper process only proposes a block.
- The target scores that block in one causal forward pass because training-style token scoring is parallel across positions.
- Greedy decoding accepts the longest prefix matching the target argmax, then emits the target correction or a bonus token.
- Exact sampling accepts proposal y with
min(1, p(y)/q(y)); rejection samples from normalized[p−q]₊. - The simple form needs aligned output units, ordering, conditioning, processors, and transactional state—not merely a “small model.”
- Expected committed tokens are
1 + α + … + α^γ; speed comes only if that saved target work exceeds draft, verification, and rollback cost. - Lossless means the target distribution is preserved. “Typical,” tolerant, or approximate acceptance may be useful, but is a different promise.
- Across modalities, you may speculate only across positions whose conditioning is already known. New microphone samples or environment feedback form a hard boundary.
- Measure end-to-end. Speculation accelerates the autoregressive section, not vision/audio encoders, prefill, codec decoders, or network queues.
01The mental model: propose, verify, commit
Any autoregressive generator can be written as
c is everything already known: a text prompt, image embeddings, encoded speech, speaker
identity, a protein motif, or a robot’s current observation. zₜ need not be a word token. It can
be a timestamp, acoustic-codec ID, image code, motion primitive, or quantized action.
Normal decoding asks the large target model for one next-token distribution at a time. Each result is needed
before the next step can begin, so the target is called serially. Speculative decoding uses a cheap distribution
q to propose γ future tokens, then asks target p to score those positions in one
causal block. Causality is preserved: position i can see only the committed prefix and earlier proposals.
What is parallel?
The target’s scoring of a known candidate block. The draft itself is often still autoregressive, just much cheaper.
Who has authority?
The target. In an exact method, a draft token is never committed merely because the draft likes it.
Where is the win?
Fewer expensive target invocations and better hardware utilization per invocation—not fewer logical dependencies.
Parallel sampling heads, blockwise decoding, lookahead, retrieval proposals, early exit, and tree decoding can all supply candidates, but they are proposal mechanisms. Beam search changes the search objective. Quantization changes arithmetic. Prompt caching shortens prefill. Speculative decoding is the verification-and-commit protocol; it can combine with several of those techniques.
02Two exact algorithms: greedy equality and sampling correction
Exact greedy decoding
The draft proposes y₁…yγ. The target supplies its argmax for each proposal position plus one
extra position. Accept draft tokens left-to-right while yᵢ = argmax pᵢ. At the first mismatch,
discard that token and everything after it, then emit the target argmax at that position. If every proposal
matches, accept all and emit the target’s extra “bonus” token.
This reproduces ordinary greedy target decoding, subject to the same logit processors, constraints, position semantics, and deterministic numerical path. A different kernel shape can cause a low-precision argmax tie to flip, so “same algorithm” is not automatically bit-identical on every backend.
Exact stochastic decoding
Let qᵢ and pᵢ be the draft and target distributions after temperature, vocabulary
masking, top-k/top-p, repetition penalties, grammar constraints, and every other processor. Draw proposal
yᵢ ~ qᵢ, then accept it with
If a proposal is rejected, draw the correction from
If all proposals are accepted, sample the bonus token from the target’s next distribution. This rejection
correction is why the final sequence has the target model’s distribution even though candidates came from
q. Leviathan et al. formalized this exact procedure and reported 2–3× improvements on their tested
Transformer workloads—not a universal promise.
Averaged over proposals y ~ q, the acceptance probability at one position is
Σₓ min(p(x), q(x)) = 1 − TV(p,q). Acceptance therefore measures distributional overlap, not simply
whether the draft is “smart.” This is why draft-to-target alignment is the useful training objective.
the one-token proof of exactness
Token x is proposed and accepted with probability
q(x)·min(1,p(x)/q(x)) = min(p(x),q(x)). A rejection happens with total probability
Σ[p−q]₊, equal to the residual normalizer. Multiplying that rejection probability by the
normalized residual adds [p(x)−q(x)]₊. Therefore the total probability of outputting
x is min(p(x),q(x)) + [p(x)−q(x)]₊ = p(x).
The ratio is evaluated only for a sampled token, so its q(y) is positive. If p=q,
rejection has probability zero and the zero residual never needs to be sampled.
p(y)=0.12 and q(y)=0.30, what is the
acceptance probability? Why can’t rejection simply resample from p?
answer and intuition
Acceptance is 0.12/0.30 = 0.4. The draft proposed y too often,
so most such proposals must be removed. Sampling rejected cases directly from p would add target
probability a second time. The residual [p−q]₊ restores probability mass only where the target was
underrepresented by accepted proposals.
| method | acceptance rule | promise | use it when |
|---|---|---|---|
| exact greedy | draft token equals target argmax | same greedy token path, within numerical determinism | parity and simple debugging matter most |
| exact stochastic | ratio test + residual correction | same target distribution | sampling quality/diversity must be preserved |
| approximate lenient | probability, rank, distance, or semantic tolerance | distribution may change | quality/speed tradeoff is explicit and evaluated |
| search beam/block | task-specific candidate ranking | usually comparable quality, not target sampling law | search quality, not distributional identity, is the objective |
Ordinary and speculative samplers consume random numbers in a different order. They can sample from the same distribution without producing the same sequence for the same seed. Test greedy token equality separately; test sampling with small exact categorical cases and statistical checks.
03One round, slowly
- Checkpoint: both caches describe exactly the committed prefix of length
L. - Draft: draw or greedily choose
γtokens. Save eachqᵢvalue needed for acceptance. - Verify: run the target on the proposal block with a causal mask. Align logits carefully: the distribution for proposal
yᵢis conditioned on the prefix plusy<i. - Accept: scan left-to-right. Rejection ends the round; later proposals are invalid because their conditioning included a rejected token.
- Emit: commit accepted tokens plus exactly one target correction, or all proposals plus one target bonus.
- Reconcile: truncate/advance both caches and every auxiliary state to the new committed length.
Every completed round advances by at least one token. With γ proposals, it advances by 1 through
γ+1 tokens. That extra target token is not decorative: it guarantees progress and captures the
target work already performed beyond a fully accepted block.
A pure greedy verifier
accepted? What should happen to proposals after it?from dataclasses import dataclass
from typing import TypeAlias
Token: TypeAlias = int
@dataclass(frozen=True)
class GreedyRound:
"""The immutable decision from one target verification block."""
accepted: tuple[Token, ...]
emitted: tuple[Token, ...]
rejected_at: int | None
def verify_greedy(
draft: tuple[Token, ...],
target_argmax: tuple[Token, ...],
) -> GreedyRound:
"""Accept the matching prefix and emit one target-owned token."""
if len(target_argmax) != len(draft) + 1:
raise ValueError("target must score every draft position plus one bonus")
for index, token in enumerate(draft):
if token != target_argmax[index]:
accepted = draft[:index]
correction = target_argmax[index]
return GreedyRound(accepted, accepted + (correction,), index)
return GreedyRound(draft, draft + (target_argmax[-1],), None)
The function contains no cache operations. That separation is intentional: acceptance math should be testable with tiny scripted distributions before being entangled with model state.
04The draft-model contract
A draft need not be an independently trained smaller Transformer. It is any cheap mechanism that can propose the next units and, for exact sampling, expose their proposal probabilities. Separate correctness requirements from performance preferences.
Hard requirements for the simple exact algorithm
A subtle but liberating fact: the draft is allowed to be inaccurate or to ignore part of the conditioning; exact verification will correct it. That hurts acceptance, not correctness. The target distribution and the accounting around the actual proposal process are what must be exact.
| contract | why it matters | what breaks |
|---|---|---|
| same output units and IDs | p(y) and q(y) must refer to the same event | token IDs from different tokenizers/codebooks are incomparable |
| same autoregressive order | both must factorize the sequence along the same axis | an image raster draft cannot verify a target using a different group order |
| exact target conditioning | p must be the ordinary baseline distribution at that position | wrong prefix, mask, position, image/audio state, or processor changes the model being sampled |
| actual proposal law | sampling must draw from the recorded normalized qᵢ; the ratio must use that same law | using raw logits after sampling from processed logits invalidates the correction proof |
| target block scoring | the expensive model must score candidates causally in one call | replaying one target step per token erases the intended speedup |
| transactional state | rejected futures must be removed everywhere | KV, positions, grammar state, or buffers silently drift |
| proposal probabilities | the ratio test needs qᵢ(yᵢ) and the residual needs the full compatible qᵢ | exact sampling is impossible; greedy can still work |
| frozen round context | every verified position must use conditioning already available for the block | new audio or environment feedback changes the conditional problem mid-round |
Performance requirements
- Cheap enough: judge measured draft latency, memory traffic, synchronization, and transfers—not parameter count alone.
- Aligned enough: acceptance on the real domain, languages, modalities, temperatures, and prompt lengths matters more than perplexity in isolation.
- Conditioned enough: an audio-blind ASR or vision-blind VLM draft can remain lossless under exact verification, but usually wastes grounded proposals.
- Small working set: a draft that evicts target weights or KV from useful memory may make the target verification slower.
- Operationally simple: shared tokenizer, processors, device, and cache format reduce bookkeeping and bugs.
Universal assisted decoding maps between tokenizer domains by converting candidate text and re-encoding it. That expands support, but alignment, boundary repair, probability accounting, and overhead become more complex. Start with identical IDs unless tokenizer mismatch is itself the research question.
05Six ways to obtain proposals
| family | mechanism | strength | main cost/risk |
|---|---|---|---|
| smaller sibling | separate model with matching vocabulary and task | simple mental model; independently deployable | extra weights, cache, and memory traffic |
| distilled draft | train q on target behavior, ideally on-policy | optimizes alignment rather than generic quality | training/data pipeline; drift as target changes |
| self-speculation | early-exit target layers propose; full layers verify | shares weights, tokenizer, and much state | target must be trained or calibrated for useful early exits |
| multi-token heads | extra heads predict several future offsets; verify a tree/block | one backbone pass yields many candidates | branch packing, tree attention, training and cache logic |
| feature-level draft | predict future hidden features, then recover candidate tokens | may align better than token-only small models | target-specific auxiliary modules and integration |
| model-free | prompt lookup, n-grams, token maps, retrieval, repetition | near-zero parameters; strong on copied/structured spans | domain-sensitive acceptance; exact sampling probabilities may be awkward |
Medusa is the canonical multi-head example; LayerSkip demonstrates early-exit self-speculation. DistillSpec shows why training the draft to match the target’s actual visited distribution can improve acceptance. These are not mutually exclusive: a server can select prompt lookup for copied spans, a small draft elsewhere, and disable speculation when recent acceptance collapses.
answer
You cannot decide from acceptance alone. Insert both into the cost model in the next section, including block-verification cost. The smaller, less accurate draft often wins because drafter quality has value only through target steps saved per unit draft cost.
06Predicting speed before benchmarking
Let γ be proposal length and α the probability each next proposal survives, treated
as constant and independent only for a rough model. The expected number of committed tokens per round is
If one ordinary target step costs Cₜ, each draft step costs C_d, block verification
costs Cᵥ(γ), and bookkeeping costs C_b, a first estimate is
This is deliberately a latency model, not a FLOP model. Verification processes more tokens than a one-token decode step, but in one launch-rich block it may use the GPU far more efficiently. Conversely, long visual prefixes, large KV reads, saturated batches, or an expensive draft can make it slower.
acceptance and cost simulator
Costs are normalized to one ordinary target decode step. Verification cost includes the target’s whole γ-token block. The model is educational; profile your actual shapes.
the simulator will show which candidate prefix survives.
What the simple equation hides
αchanges with proposal position, temperature, domain, and recent context. Report it by offset, not only as one average.Cᵥ(γ)is not constant: attention reads grow, kernel shapes change, and longer blocks may cross efficient tile boundaries.- Draft and target may contend for memory, streams, or devices. CPU drafting can add transfers and synchronization.
- EOS, short answers, and early rejection truncate rounds. Prefill and non-AR stages dilute decode-only gains through Amdahl’s law.
Start conservative, track recent accepted prefix length, and increase γ only while marginal candidates remain cheap and likely to survive. The optimum is a runtime policy, not a sacred constant.
07KV caches and state: treat each round as a transaction
Most broken implementations get the acceptance formula right and the state machine wrong. At round start,
record the committed length L. Draft and target may write speculative KV entries, but none are final
until acceptance is known.
The logit-shift invariant
If target cache ends at the committed prefix and you already retained its next-token logits, those logits score
y₁. Feeding y₁…yγ returns logits that score y₂…yγ and the bonus position.
Implementations that compare proposal i with output row i without accounting for this
shift are off by one.
State that must roll back or recompute
Model state
target KV, draft KV, cache length, RoPE/MRoPE positions, sliding-window indices, recurrent/convolutional state, cross-attention projections.
Decoder state
grammar DFA, no-repeat n-grams, timestamp constraints, forced tokens, beam hypotheses, EOS flags, sampler RNG accounting, stream buffers.
With a static cache, rejected values may remain physically stored only if logical length, masks, and position indices guarantee they can never be read. “We overwrote the pointer later” is not proof. Padded cache attention requires an explicit valid-position mask.
The verification pass has already computed their target KV. Re-running a normal target decode step for each accepted token restores correctness in some skeletons but spends the serial work speculation was meant to save. Commit or compact the block outputs instead.
08LLMs: the cleanest starting point
Recommended first implementation
- Greedy, batch one, same tokenizer, small sibling draft, fixed
γ=3or4. - Share identical prompt tokens but keep target and draft caches separate.
- Implement one target block-verification call and transactional cache truncation.
- Require exact token equality against ordinary greedy generation over odd prompt/output lengths.
- Only then add exact sampling, adaptive length, batching, trees, or different tokenizers.
Where each drafter works
| workload | good first proposal source | watch |
|---|---|---|
| summarization / grounded QA | prompt lookup or n-gram + small model fallback | copied spans accept well; factual pivots may not |
| code | retrieval/prompt lookup, distilled code draft, multi-token heads | whitespace/tokenizer and grammar processors |
| chat | domain-aligned small sibling or self-speculation | temperature and multilingual distribution shift |
| constrained JSON/tools | same draft plus shared grammar state | processor mismatch breaks exactness before model mismatch does |
Framework support is an implementation detail, not an algorithm guarantee. Current Hugging Face assisted generation supports assistant models, prompt lookup, self-speculation, and tokenizer-translation variants, but batching/cache constraints vary by release. vLLM exposes several speculative backends. Check the installed version and benchmark your serving scheduler rather than copying an old configuration.
09Vision-language models
For a VLM that autoregressively emits text, the verification math is unchanged. The difficult part is conditioning. The target distribution depends on visual features, layout/position encodings, and special image tokens. A text-only draft may predict fluent glue words well and fail exactly where visual grounding matters.
Three architectures
shared vision state
Encode the image once; project or expose compatible visual features to both decoders. Fastest when architecture permits it.
small multimodal draft
A compact VLM receives the same image and prompt. Better grounding, but may duplicate encoder work and memory.
text-only draft
Cheap and easy. Safe under exact verification, but low acceptance around objects, OCR, spatial relations, and rare entities.
VLM checklist
- Freeze and reuse the visual prefix within a round; match MRoPE/position IDs and image-token placement.
- Do not claim that speculative decode accelerates the image encoder or multimodal prefill.
- Report acceptance separately for visually grounded tokens and ordinary syntax if possible.
- Long visual KV prefixes can make every block verification expensive; model visual-token compression as a separate lever.
- Evaluate text/token parity for exact decoding and task quality (OCR, VQA, grounding) for approximate schemes.
Early multimodal studies report that speculative decoding can help VLMs, but performance is workload- and architecture-sensitive. Treat newer multimodal speculators and benchmarks as emerging evidence, not as a reason to assume an LLM draft will transfer unchanged.
10LM-based ASR: audio is not optional conditioning
Encoder-decoder ASR such as Whisper-like systems has a non-autoregressive audio encoder followed by an autoregressive text decoder. Encode the current audio once. For useful acceptance, both draft and target decoders should be conditioned on semantically equivalent audio state; an audio-blind draft can still be corrected exactly, but will tend to fail on acoustically determined tokens. Their projected cross-attention KV can remain separate if widths differ.
| ASR concern | required treatment |
|---|---|
| vocabulary | align text, language, task, timestamp, no-speech, and EOS tokens—not only ordinary words |
| forced prefix | apply language/task prompts and suppression rules identically before comparing distributions |
| timestamps | share timestamp-range constraints and monotonic state; exact text with wrong timing is not parity |
| streaming | new audio changes c; discard or revalidate proposals created under the older acoustic context |
| beam search | start with greedy; speculative beams require per-hypothesis caches and candidate accounting |
| measurement | report real-time factor, first-token latency, chunk latency, encoder/decode split, WER, and timestamp quality |
A smaller audio-conditioned decoder is the direct route. Model-free token-map/n-gram drafting can be useful in structured low-perplexity domains; recent SpecASR work explores adaptive proposal lengths and draft recycling. These are promising research directions, but reproduce them on your audio, language mix, and streaming policy.
11TTS, audio, images, video, motion, proteins, and actions
The acceptance protocol is modality-agnostic; the factorization is not. First write down exactly what one autoregressive step means.
| generator | one token / axis | draft must share | what speculation does not speed up |
|---|---|---|---|
| AR TTS / speech LM | semantic or acoustic codec IDs; sometimes several codebooks per frame | text, speaker, language, acoustic prefix, codebook order, EOS | text/audio encoders and non-AR waveform decoder |
| music/audio LM | codec IDs over time and codebooks | conditioning tags, timing grid, interleaving pattern | codec decoder and post-processing |
| discrete image/video AR | VQ code IDs in raster, grouped, or space-time order | codebook, scan/group order, class/text condition | VQ encoder/decoder and prompt encoder |
| continuous AR image | continuous vectors/patches | compatible probability densities and ordering | everything outside AR sampling |
| motion / action tokens | quantized poses or action chunks | agent state, goal, observation history, safety constraints | environment/sensors and downstream controller |
| protein / molecule LM | residue, atom, bond, or structural token | alphabet, constraints, generation order, conditioning | structure relaxation and external scoring |
Multi-codebook audio needs an axis decision
A codec model may generate time-major frames, codebook-major streams, or one semantic stream followed by parallel residual codebooks. You may speculate across time only where the target is autoregressive across time; you may speculate within a frame only if that is the target’s declared factorization. Flattening IDs in a convenient order changes the conditional model. For a lossless implementation, validate exact codec-ID parity; waveform similarity and listening tests are necessary only after choosing an approximate acceptance rule.
Discrete and continuous outputs are different mathematics
The ratio/residual equations above are easiest for finite discrete vocabularies. Continuous AR models require density-aware acceptance and a way to sample the positive residual density; nearest-neighbor tolerance is an approximate method unless proven otherwise. Recent continuous-image speculation work develops specialized machinery—do not paste the categorical formula onto vectors.
The external-feedback boundary
You can speculate only while future conditioning is fixed. An offline ASR model can draft across a known audio segment. A streaming ASR model cannot safely treat yet-unheard samples as fixed. A robot may draft an open-loop action chunk internally, but it cannot commit actions beyond a new environment observation merely because the target verified them under the old state. Verify before execution and replan at feedback boundaries.
12Serving, batching, and when speculation loses
Speculative decoding was first attractive at batch one, where one-token target calls underuse hardware and weight/KV traffic dominates. In a busy continuous-batching server, ordinary decoding may already fill the GPU. Variable accepted lengths also create ragged work and scheduling pressure.
Disable or reconsider it when
- prefill, image/audio encoding, codec decoding, or network queueing dominates end-to-end latency;
- the target is already small or the batch already saturates compute;
- outputs are so short that setup and extra model loading dominate;
- acceptance collapses under high temperature, domain/language shift, or grounded multimodal tokens;
- the draft consumes enough memory bandwidth or VRAM to slow/evict the target;
- verification uses a poor kernel path for its block shape, or rejected KV is replayed serially;
- the serving engine cannot pack variable-length proposal blocks efficiently.
Useful runtime policies
acceptance controller
Track recent accepted prefix by request/domain; shrink γ or turn speculation off after repeated early rejection.
proposal router
Use prompt lookup for copied spans, an audio/vision-aware model for grounded regions, and a small general draft elsewhere.
load-aware scheduler
Favor speculation at low concurrency/latency mode; reassess at throughput-oriented large batches.
memory-aware placement
Measure same-GPU contention versus CPU/second-GPU transfer. Never infer it from parameter count.
13A safe implementation path for decode-lab
Build correctness in layers. Keep ordinary eager generation as the oracle and switchable fallback throughout.
- Pure verifier: unit-test greedy prefix acceptance and stochastic residual math with tiny hand-computed distributions.
- No-cache reference: score the full committed prefix plus proposals and prove logit alignment.
- Target block path: add causal multi-token verification; compare every relevant logit with the no-cache reference.
- Transactional caches: checkpoint lengths, retain accepted block KV, truncate rejected suffixes, then synchronize the draft.
- Greedy parity: exact IDs over many prompts, odd lengths, early EOS, γ larger than remaining output, and forced constraints.
- Exact sampling: add processed
p/q, ratio tests, residual correction, and statistical validation. - Performance: profile draft, target verify, cache reconciliation, synchronization, and end-to-end latency independently.
- Only then: adaptive γ, trees/heads, multiple devices, batching, VLM/ASR/TTS conditioning, or approximate acceptance.
Audit of the current teaching skeleton
src/qwen3_lm/inference.py::InferenceEngine.speculative_generate is currently a conceptual sketch,
not a production-ready speculative path. Before benchmarking it, account for these exact failure modes:
- After prefill, drafting starts by feeding the prompt’s final token again instead of using the next-token logits returned by prefill, so the draft context begins with a duplicated token.
- The target verification call scores
mergedwithout the target cache/prefix state, so it does not necessarily evaluate candidates under the committed context. - Its output row
ipredicts the token after proposali, yet the loop compares that row with proposali; this is the logit-shift error from section 7. - Accepted target tokens are then replayed through
decode_stepone at a time, paying the serial target work again. - The draft cache has already advanced through every guess; advancing it again for accepted tokens, without rollback on rejection, desynchronizes its logical history.
- The routine implements a greedy comparison only; it does not implement exact speculative sampling.
The repair is to give verification the correct committed target state, retain target KV from the accepted block rather than replaying it, roll both engines back/forward to one committed length, and test greedy parity before adding the stochastic protocol. This lesson documents that design; it intentionally does not refactor runtime code in the same pass.
14The benchmark and validation contract
A speedup without a matched decoding contract is not evidence. Freeze this table before running the GPU.
| dimension | record |
|---|---|
| models | exact target/draft checkpoints, revisions, dtypes, quantization, tokenizer/codebook |
| decode policy | greedy or sampling; temperature/top-k/top-p; penalties, grammar, EOS; γ and adaptation |
| workload | prompt/source lengths, output lengths, domain/language/modality, batch/concurrency |
| latency | cold/warm, prefill, draft, verify, reconciliation, time-to-first-token, inter-token, p50/p95, end-to-end |
| acceptance | accepted-prefix histogram by proposal offset, tokens per target invocation, rejection causes |
| resources | peak VRAM, target/draft KV, memory placement, power if relevant |
| correctness | greedy exact IDs; sampling distribution tests; cache/logit parity; constraint/timestamp state |
| modality quality | VQA/OCR, WER/timestamps, codec IDs/audio quality, image metrics, task reward—as applicable |
Report both decode-only and end-to-end speedup. Use warmups, synchronize device timing correctly, include the non-speculative target baseline in the same process, and publish regressions too. CPU structural tests can validate state and math; they do not prove CUDA kernel behavior, actual model quality, or GPU latency.
15One-page decision sheet
Is it applicable?
- Is generation autoregressive?
- Are several future conditions already known?
- Can the target score a causal candidate block?
- Is serial target decode material end-to-end?
Is it exact?
- Same events/order/conditioning?
- Same processed distributions?
- Greedy equality or ratio + residual?
- All state transactional?
Will it be fast?
- Measure draft cost.
- Measure block verify cost.
- Measure acceptance by offset.
- Check memory and batch contention.
How should I start?
- Greedy, batch one.
- Same tokenizer/codebook.
- γ = 3–4.
- Exact parity before sampling or fusion.
A candidate is cheap; a commitment is expensive. Draft freely inside a frozen conditional world, let the authoritative model verify in parallel, commit only the valid prefix, and make state changes follow that decision. That is speculative decoding across every modality.
Primary sources and further study
- Leviathan, Kalman & Matias (ICML 2023), Fast Inference from Transformers via Speculative Decoding — exact stochastic algorithm and original systems analysis.
- Google Research, Looking back at speculative decoding — retrospective, intuition, and extensions beyond text.
- Zhou et al. (ICLR 2024), DistillSpec — draft/target alignment through task- and divergence-aware distillation.
- Cai et al. (ICML 2024), Medusa — multiple decoding heads and tree attention.
- Elhoushi et al. (ACL 2024), LayerSkip — early-exit self-speculative decoding.
- Sun et al., Block Verification Accelerates Speculative Decoding — verification need not be limited to a left-to-right rejection scan.
- Gagrani et al. (CVPRW 2024), On Speculative Decoding for Multimodal LLMs.
- SpecASR (2025 preprint) and Token Map Drafting for ASR (2025 preprint) — emerging audio-conditioned and model-free ASR approaches.
- So et al. (ICCV 2025), Grouped Speculative Decoding for Autoregressive Image Generation.
- Continuous Speculative Decoding for Autoregressive Image Generation — specialized treatment for continuous outputs.
- VADUSA (2024 preprint) — speculative AR speech synthesis with an explicit quality/speed tolerance mechanism.
- Hugging Face Transformers assisted decoding guide and vLLM speculative decoding docs — current framework interfaces; verify against the installed version.