lab

Speculative Decoding, Beyond Text

A complete mental model for replacing several expensive serial decode steps with cheap guesses and one parallel verification pass—without changing the target model’s output law when the exact algorithm is used. The same idea reaches language, vision-language, speech recognition, audio codecs, discrete image/video generators, music, motion, proteins, and action models. What changes is the token, conditioning, and state contract.

the whole lesson in nine lines
  1. An autoregressive target still decides every committed token; a cheaper process only proposes a block.
  2. The target scores that block in one causal forward pass because training-style token scoring is parallel across positions.
  3. Greedy decoding accepts the longest prefix matching the target argmax, then emits the target correction or a bonus token.
  4. Exact sampling accepts proposal y with min(1, p(y)/q(y)); rejection samples from normalized [p−q]₊.
  5. The simple form needs aligned output units, ordering, conditioning, processors, and transactional state—not merely a “small model.”
  6. Expected committed tokens are 1 + α + … + α^γ; speed comes only if that saved target work exceeds draft, verification, and rollback cost.
  7. Lossless means the target distribution is preserved. “Typical,” tolerant, or approximate acceptance may be useful, but is a different promise.
  8. Across modalities, you may speculate only across positions whose conditioning is already known. New microphone samples or environment feedback form a hard boundary.
  9. Measure end-to-end. Speculation accelerates the autoregressive section, not vision/audio encoders, prefill, codec decoders, or network queues.

01The mental model: propose, verify, commit

Any autoregressive generator can be written as

p(z₁:T | c) = ∏ₜ p(zₜ | c, z<t)

c is everything already known: a text prompt, image embeddings, encoded speech, speaker identity, a protein motif, or a robot’s current observation. zₜ need not be a word token. It can be a timestamp, acoustic-codec ID, image code, motion primitive, or quantized action.

Normal decoding asks the large target model for one next-token distribution at a time. Each result is needed before the next step can begin, so the target is called serially. Speculative decoding uses a cheap distribution q to propose γ future tokens, then asks target p to score those positions in one causal block. Causality is preserved: position i can see only the committed prefix and earlier proposals.

committed prefix→ draft γ guesses→ one target verification→ accept longest prefix→ correction / bonus

What is parallel?

The target’s scoring of a known candidate block. The draft itself is often still autoregressive, just much cheaper.

Who has authority?

The target. In an exact method, a draft token is never committed merely because the draft likes it.

Where is the win?

Fewer expensive target invocations and better hardware utilization per invocation—not fewer logical dependencies.

not the same as these

Parallel sampling heads, blockwise decoding, lookahead, retrieval proposals, early exit, and tree decoding can all supply candidates, but they are proposal mechanisms. Beam search changes the search objective. Quantization changes arithmetic. Prompt caching shortens prefill. Speculative decoding is the verification-and-commit protocol; it can combine with several of those techniques.

02Two exact algorithms: greedy equality and sampling correction

Exact greedy decoding

The draft proposes y₁…yγ. The target supplies its argmax for each proposal position plus one extra position. Accept draft tokens left-to-right while yᵢ = argmax pᵢ. At the first mismatch, discard that token and everything after it, then emit the target argmax at that position. If every proposal matches, accept all and emit the target’s extra “bonus” token.

This reproduces ordinary greedy target decoding, subject to the same logit processors, constraints, position semantics, and deterministic numerical path. A different kernel shape can cause a low-precision argmax tie to flip, so “same algorithm” is not automatically bit-identical on every backend.

Exact stochastic decoding

Let qᵢ and pᵢ be the draft and target distributions after temperature, vocabulary masking, top-k/top-p, repetition penalties, grammar constraints, and every other processor. Draw proposal yᵢ ~ qᵢ, then accept it with

aᵢ = min(1, pᵢ(yᵢ) / qᵢ(yᵢ)).

If a proposal is rejected, draw the correction from

rᵢ(x) = [pᵢ(x) − qᵢ(x)]₊ / Σⱼ[pᵢ(j) − qᵢ(j)]₊.

If all proposals are accepted, sample the bonus token from the target’s next distribution. This rejection correction is why the final sequence has the target model’s distribution even though candidates came from q. Leviathan et al. formalized this exact procedure and reported 2–3× improvements on their tested Transformer workloads—not a universal promise.

Averaged over proposals y ~ q, the acceptance probability at one position is Σₓ min(p(x), q(x)) = 1 − TV(p,q). Acceptance therefore measures distributional overlap, not simply whether the draft is “smart.” This is why draft-to-target alignment is the useful training objective.

the one-token proof of exactness

Token x is proposed and accepted with probability q(x)·min(1,p(x)/q(x)) = min(p(x),q(x)). A rejection happens with total probability Σ[p−q]₊, equal to the residual normalizer. Multiplying that rejection probability by the normalized residual adds [p(x)−q(x)]₊. Therefore the total probability of outputting x is min(p(x),q(x)) + [p(x)−q(x)]₊ = p(x).

The ratio is evaluated only for a sampled token, so its q(y) is positive. If p=q, rejection has probability zero and the zero residual never needs to be sampled.

predict before opening: if p(y)=0.12 and q(y)=0.30, what is the acceptance probability? Why can’t rejection simply resample from p?
answer and intuition

Acceptance is 0.12/0.30 = 0.4. The draft proposed y too often, so most such proposals must be removed. Sampling rejected cases directly from p would add target probability a second time. The residual [p−q]₊ restores probability mass only where the target was underrepresented by accepted proposals.

methodacceptance rulepromiseuse it when
exact greedydraft token equals target argmaxsame greedy token path, within numerical determinismparity and simple debugging matter most
exact stochasticratio test + residual correctionsame target distributionsampling quality/diversity must be preserved
approximate lenientprobability, rank, distance, or semantic tolerancedistribution may changequality/speed tradeoff is explicit and evaluated
search beam/blocktask-specific candidate rankingusually comparable quality, not target sampling lawsearch quality, not distributional identity, is the objective
distributional exactness is not seed identity

Ordinary and speculative samplers consume random numbers in a different order. They can sample from the same distribution without producing the same sequence for the same seed. Test greedy token equality separately; test sampling with small exact categorical cases and statistical checks.

03One round, slowly

  1. Checkpoint: both caches describe exactly the committed prefix of length L.
  2. Draft: draw or greedily choose γ tokens. Save each qᵢ value needed for acceptance.
  3. Verify: run the target on the proposal block with a causal mask. Align logits carefully: the distribution for proposal yᵢ is conditioned on the prefix plus y<i.
  4. Accept: scan left-to-right. Rejection ends the round; later proposals are invalid because their conditioning included a rejected token.
  5. Emit: commit accepted tokens plus exactly one target correction, or all proposals plus one target bonus.
  6. Reconcile: truncate/advance both caches and every auxiliary state to the new committed length.
the +1 invariant

Every completed round advances by at least one token. With γ proposals, it advances by 1 through γ+1 tokens. That extra target token is not decorative: it guarantees progress and captures the target work already performed beyond a fully accepted block.

A pure greedy verifier

predict before reading the code: should the mismatching draft token be included in accepted? What should happen to proposals after it?
from dataclasses import dataclass
from typing import TypeAlias

Token: TypeAlias = int


@dataclass(frozen=True)
class GreedyRound:
    """The immutable decision from one target verification block."""

    accepted: tuple[Token, ...]
    emitted: tuple[Token, ...]
    rejected_at: int | None


def verify_greedy(
    draft: tuple[Token, ...],
    target_argmax: tuple[Token, ...],
) -> GreedyRound:
    """Accept the matching prefix and emit one target-owned token."""
    if len(target_argmax) != len(draft) + 1:
        raise ValueError("target must score every draft position plus one bonus")

    for index, token in enumerate(draft):
        if token != target_argmax[index]:
            accepted = draft[:index]
            correction = target_argmax[index]
            return GreedyRound(accepted, accepted + (correction,), index)

    return GreedyRound(draft, draft + (target_argmax[-1],), None)

The function contains no cache operations. That separation is intentional: acceptance math should be testable with tiny scripted distributions before being entangled with model state.

04The draft-model contract

A draft need not be an independently trained smaller Transformer. It is any cheap mechanism that can propose the next units and, for exact sampling, expose their proposal probabilities. Separate correctness requirements from performance preferences.

Hard requirements for the simple exact algorithm

A subtle but liberating fact: the draft is allowed to be inaccurate or to ignore part of the conditioning; exact verification will correct it. That hurts acceptance, not correctness. The target distribution and the accounting around the actual proposal process are what must be exact.

contractwhy it matterswhat breaks
same output units and IDsp(y) and q(y) must refer to the same eventtoken IDs from different tokenizers/codebooks are incomparable
same autoregressive orderboth must factorize the sequence along the same axisan image raster draft cannot verify a target using a different group order
exact target conditioningp must be the ordinary baseline distribution at that positionwrong prefix, mask, position, image/audio state, or processor changes the model being sampled
actual proposal lawsampling must draw from the recorded normalized qᵢ; the ratio must use that same lawusing raw logits after sampling from processed logits invalidates the correction proof
target block scoringthe expensive model must score candidates causally in one callreplaying one target step per token erases the intended speedup
transactional staterejected futures must be removed everywhereKV, positions, grammar state, or buffers silently drift
proposal probabilitiesthe ratio test needs qᵢ(yᵢ) and the residual needs the full compatible qᵢexact sampling is impossible; greedy can still work
frozen round contextevery verified position must use conditioning already available for the blocknew audio or environment feedback changes the conditional problem mid-round

Performance requirements

  • Cheap enough: judge measured draft latency, memory traffic, synchronization, and transfers—not parameter count alone.
  • Aligned enough: acceptance on the real domain, languages, modalities, temperatures, and prompt lengths matters more than perplexity in isolation.
  • Conditioned enough: an audio-blind ASR or vision-blind VLM draft can remain lossless under exact verification, but usually wastes grounded proposals.
  • Small working set: a draft that evicts target weights or KV from useful memory may make the target verification slower.
  • Operationally simple: shared tokenizer, processors, device, and cache format reduce bookkeeping and bugs.
different tokenizers are possible, not free

Universal assisted decoding maps between tokenizer domains by converting candidate text and re-encoding it. That expands support, but alignment, boundary repair, probability accounting, and overhead become more complex. Start with identical IDs unless tokenizer mismatch is itself the research question.

05Six ways to obtain proposals

familymechanismstrengthmain cost/risk
smaller siblingseparate model with matching vocabulary and tasksimple mental model; independently deployableextra weights, cache, and memory traffic
distilled drafttrain q on target behavior, ideally on-policyoptimizes alignment rather than generic qualitytraining/data pipeline; drift as target changes
self-speculationearly-exit target layers propose; full layers verifyshares weights, tokenizer, and much statetarget must be trained or calibrated for useful early exits
multi-token headsextra heads predict several future offsets; verify a tree/blockone backbone pass yields many candidatesbranch packing, tree attention, training and cache logic
feature-level draftpredict future hidden features, then recover candidate tokensmay align better than token-only small modelstarget-specific auxiliary modules and integration
model-freeprompt lookup, n-grams, token maps, retrieval, repetitionnear-zero parameters; strong on copied/structured spansdomain-sensitive acceptance; exact sampling probabilities may be awkward

Medusa is the canonical multi-head example; LayerSkip demonstrates early-exit self-speculation. DistillSpec shows why training the draft to match the target’s actual visited distribution can improve acceptance. These are not mutually exclusive: a server can select prompt lookup for copied spans, a small draft elsewhere, and disable speculation when recent acceptance collapses.

predict: a 1B draft accepts 90% of tokens but costs 45% of a target step; a tiny draft accepts 72% and costs 6%. Which is faster?
answer

You cannot decide from acceptance alone. Insert both into the cost model in the next section, including block-verification cost. The smaller, less accurate draft often wins because drafter quality has value only through target steps saved per unit draft cost.

06Predicting speed before benchmarking

Let γ be proposal length and α the probability each next proposal survives, treated as constant and independent only for a rough model. The expected number of committed tokens per round is

E[N] = 1 + α + α² + … + α^γ = (1 − α^(γ+1)) / (1 − α).

If one ordinary target step costs Cₜ, each draft step costs C_d, block verification costs Cᵥ(γ), and bookkeeping costs C_b, a first estimate is

speedup ≈ E[N] · Cₜ / (γ·C_d + Cᵥ(γ) + C_b).

This is deliberately a latency model, not a FLOP model. Verification processes more tokens than a one-token decode step, but in one launch-rich block it may use the GPU far more efficiently. Conversely, long visual prefixes, large KV reads, saturated batches, or an expensive draft can make it slower.

acceptance and cost simulator

Costs are normalized to one ordinary target decode step. Verification cost includes the target’s whole γ-token block. The model is educational; profile your actual shapes.

3.36expected tokens / round
1.47normalized round cost
2.29×estimated speedup

the simulator will show which candidate prefix survives.

What the simple equation hides

  • α changes with proposal position, temperature, domain, and recent context. Report it by offset, not only as one average.
  • Cᵥ(γ) is not constant: attention reads grow, kernel shapes change, and longer blocks may cross efficient tile boundaries.
  • Draft and target may contend for memory, streams, or devices. CPU drafting can add transfers and synchronization.
  • EOS, short answers, and early rejection truncate rounds. Prefill and non-AR stages dilute decode-only gains through Amdahl’s law.
adaptive γ

Start conservative, track recent accepted prefix length, and increase γ only while marginal candidates remain cheap and likely to survive. The optimum is a runtime policy, not a sacred constant.

07KV caches and state: treat each round as a transaction

Most broken implementations get the acceptance formula right and the state machine wrong. At round start, record the committed length L. Draft and target may write speculative KV entries, but none are final until acceptance is known.

checkpoint L→speculative writes →verify r accepted→ truncate to L+r→append correction / bonus

The logit-shift invariant

If target cache ends at the committed prefix and you already retained its next-token logits, those logits score y₁. Feeding y₁…yγ returns logits that score y₂…yγ and the bonus position. Implementations that compare proposal i with output row i without accounting for this shift are off by one.

State that must roll back or recompute

Model state

target KV, draft KV, cache length, RoPE/MRoPE positions, sliding-window indices, recurrent/convolutional state, cross-attention projections.

Decoder state

grammar DFA, no-repeat n-grams, timestamp constraints, forced tokens, beam hypotheses, EOS flags, sampler RNG accounting, stream buffers.

With a static cache, rejected values may remain physically stored only if logical length, masks, and position indices guarantee they can never be read. “We overwrote the pointer later” is not proof. Padded cache attention requires an explicit valid-position mask.

do not replay accepted target tokens one by one

The verification pass has already computed their target KV. Re-running a normal target decode step for each accepted token restores correctness in some skeletons but spends the serial work speculation was meant to save. Commit or compact the block outputs instead.

08LLMs: the cleanest starting point

Recommended first implementation

  1. Greedy, batch one, same tokenizer, small sibling draft, fixed γ=3 or 4.
  2. Share identical prompt tokens but keep target and draft caches separate.
  3. Implement one target block-verification call and transactional cache truncation.
  4. Require exact token equality against ordinary greedy generation over odd prompt/output lengths.
  5. Only then add exact sampling, adaptive length, batching, trees, or different tokenizers.

Where each drafter works

workloadgood first proposal sourcewatch
summarization / grounded QAprompt lookup or n-gram + small model fallbackcopied spans accept well; factual pivots may not
coderetrieval/prompt lookup, distilled code draft, multi-token headswhitespace/tokenizer and grammar processors
chatdomain-aligned small sibling or self-speculationtemperature and multilingual distribution shift
constrained JSON/toolssame draft plus shared grammar stateprocessor mismatch breaks exactness before model mismatch does

Framework support is an implementation detail, not an algorithm guarantee. Current Hugging Face assisted generation supports assistant models, prompt lookup, self-speculation, and tokenizer-translation variants, but batching/cache constraints vary by release. vLLM exposes several speculative backends. Check the installed version and benchmark your serving scheduler rather than copying an old configuration.

09Vision-language models

For a VLM that autoregressively emits text, the verification math is unchanged. The difficult part is conditioning. The target distribution depends on visual features, layout/position encodings, and special image tokens. A text-only draft may predict fluent glue words well and fail exactly where visual grounding matters.

Three architectures

shared vision state

Encode the image once; project or expose compatible visual features to both decoders. Fastest when architecture permits it.

small multimodal draft

A compact VLM receives the same image and prompt. Better grounding, but may duplicate encoder work and memory.

text-only draft

Cheap and easy. Safe under exact verification, but low acceptance around objects, OCR, spatial relations, and rare entities.

VLM checklist

  • Freeze and reuse the visual prefix within a round; match MRoPE/position IDs and image-token placement.
  • Do not claim that speculative decode accelerates the image encoder or multimodal prefill.
  • Report acceptance separately for visually grounded tokens and ordinary syntax if possible.
  • Long visual KV prefixes can make every block verification expensive; model visual-token compression as a separate lever.
  • Evaluate text/token parity for exact decoding and task quality (OCR, VQA, grounding) for approximate schemes.

Early multimodal studies report that speculative decoding can help VLMs, but performance is workload- and architecture-sensitive. Treat newer multimodal speculators and benchmarks as emerging evidence, not as a reason to assume an LLM draft will transfer unchanged.

10LM-based ASR: audio is not optional conditioning

Encoder-decoder ASR such as Whisper-like systems has a non-autoregressive audio encoder followed by an autoregressive text decoder. Encode the current audio once. For useful acceptance, both draft and target decoders should be conditioned on semantically equivalent audio state; an audio-blind draft can still be corrected exactly, but will tend to fail on acoustically determined tokens. Their projected cross-attention KV can remain separate if widths differ.

ASR concernrequired treatment
vocabularyalign text, language, task, timestamp, no-speech, and EOS tokens—not only ordinary words
forced prefixapply language/task prompts and suppression rules identically before comparing distributions
timestampsshare timestamp-range constraints and monotonic state; exact text with wrong timing is not parity
streamingnew audio changes c; discard or revalidate proposals created under the older acoustic context
beam searchstart with greedy; speculative beams require per-hypothesis caches and candidate accounting
measurementreport real-time factor, first-token latency, chunk latency, encoder/decode split, WER, and timestamp quality
draft choices for ASR

A smaller audio-conditioned decoder is the direct route. Model-free token-map/n-gram drafting can be useful in structured low-perplexity domains; recent SpecASR work explores adaptive proposal lengths and draft recycling. These are promising research directions, but reproduce them on your audio, language mix, and streaming policy.

11TTS, audio, images, video, motion, proteins, and actions

The acceptance protocol is modality-agnostic; the factorization is not. First write down exactly what one autoregressive step means.

generatorone token / axisdraft must sharewhat speculation does not speed up
AR TTS / speech LMsemantic or acoustic codec IDs; sometimes several codebooks per frametext, speaker, language, acoustic prefix, codebook order, EOStext/audio encoders and non-AR waveform decoder
music/audio LMcodec IDs over time and codebooksconditioning tags, timing grid, interleaving patterncodec decoder and post-processing
discrete image/video ARVQ code IDs in raster, grouped, or space-time ordercodebook, scan/group order, class/text conditionVQ encoder/decoder and prompt encoder
continuous AR imagecontinuous vectors/patchescompatible probability densities and orderingeverything outside AR sampling
motion / action tokensquantized poses or action chunksagent state, goal, observation history, safety constraintsenvironment/sensors and downstream controller
protein / molecule LMresidue, atom, bond, or structural tokenalphabet, constraints, generation order, conditioningstructure relaxation and external scoring

Multi-codebook audio needs an axis decision

A codec model may generate time-major frames, codebook-major streams, or one semantic stream followed by parallel residual codebooks. You may speculate across time only where the target is autoregressive across time; you may speculate within a frame only if that is the target’s declared factorization. Flattening IDs in a convenient order changes the conditional model. For a lossless implementation, validate exact codec-ID parity; waveform similarity and listening tests are necessary only after choosing an approximate acceptance rule.

Discrete and continuous outputs are different mathematics

The ratio/residual equations above are easiest for finite discrete vocabularies. Continuous AR models require density-aware acceptance and a way to sample the positive residual density; nearest-neighbor tolerance is an approximate method unless proven otherwise. Recent continuous-image speculation work develops specialized machinery—do not paste the categorical formula onto vectors.

The external-feedback boundary

the most important beyond-text rule

You can speculate only while future conditioning is fixed. An offline ASR model can draft across a known audio segment. A streaming ASR model cannot safely treat yet-unheard samples as fixed. A robot may draft an open-loop action chunk internally, but it cannot commit actions beyond a new environment observation merely because the target verified them under the old state. Verify before execution and replan at feedback boundaries.

12Serving, batching, and when speculation loses

Speculative decoding was first attractive at batch one, where one-token target calls underuse hardware and weight/KV traffic dominates. In a busy continuous-batching server, ordinary decoding may already fill the GPU. Variable accepted lengths also create ragged work and scheduling pressure.

Disable or reconsider it when

  • prefill, image/audio encoding, codec decoding, or network queueing dominates end-to-end latency;
  • the target is already small or the batch already saturates compute;
  • outputs are so short that setup and extra model loading dominate;
  • acceptance collapses under high temperature, domain/language shift, or grounded multimodal tokens;
  • the draft consumes enough memory bandwidth or VRAM to slow/evict the target;
  • verification uses a poor kernel path for its block shape, or rejected KV is replayed serially;
  • the serving engine cannot pack variable-length proposal blocks efficiently.

Useful runtime policies

acceptance controller

Track recent accepted prefix by request/domain; shrink γ or turn speculation off after repeated early rejection.

proposal router

Use prompt lookup for copied spans, an audio/vision-aware model for grounded regions, and a small general draft elsewhere.

load-aware scheduler

Favor speculation at low concurrency/latency mode; reassess at throughput-oriented large batches.

memory-aware placement

Measure same-GPU contention versus CPU/second-GPU transfer. Never infer it from parameter count.

13A safe implementation path for decode-lab

Build correctness in layers. Keep ordinary eager generation as the oracle and switchable fallback throughout.

  1. Pure verifier: unit-test greedy prefix acceptance and stochastic residual math with tiny hand-computed distributions.
  2. No-cache reference: score the full committed prefix plus proposals and prove logit alignment.
  3. Target block path: add causal multi-token verification; compare every relevant logit with the no-cache reference.
  4. Transactional caches: checkpoint lengths, retain accepted block KV, truncate rejected suffixes, then synchronize the draft.
  5. Greedy parity: exact IDs over many prompts, odd lengths, early EOS, γ larger than remaining output, and forced constraints.
  6. Exact sampling: add processed p/q, ratio tests, residual correction, and statistical validation.
  7. Performance: profile draft, target verify, cache reconciliation, synchronization, and end-to-end latency independently.
  8. Only then: adaptive γ, trees/heads, multiple devices, batching, VLM/ASR/TTS conditioning, or approximate acceptance.

Audit of the current teaching skeleton

src/qwen3_lm/inference.py::InferenceEngine.speculative_generate is currently a conceptual sketch, not a production-ready speculative path. Before benchmarking it, account for these exact failure modes:

  • After prefill, drafting starts by feeding the prompt’s final token again instead of using the next-token logits returned by prefill, so the draft context begins with a duplicated token.
  • The target verification call scores merged without the target cache/prefix state, so it does not necessarily evaluate candidates under the committed context.
  • Its output row i predicts the token after proposal i, yet the loop compares that row with proposal i; this is the logit-shift error from section 7.
  • Accepted target tokens are then replayed through decode_step one at a time, paying the serial target work again.
  • The draft cache has already advanced through every guess; advancing it again for accepted tokens, without rollback on rejection, desynchronizes its logical history.
  • The routine implements a greedy comparison only; it does not implement exact speculative sampling.
fix design, before code

The repair is to give verification the correct committed target state, retain target KV from the accepted block rather than replaying it, roll both engines back/forward to one committed length, and test greedy parity before adding the stochastic protocol. This lesson documents that design; it intentionally does not refactor runtime code in the same pass.

14The benchmark and validation contract

A speedup without a matched decoding contract is not evidence. Freeze this table before running the GPU.

dimensionrecord
modelsexact target/draft checkpoints, revisions, dtypes, quantization, tokenizer/codebook
decode policygreedy or sampling; temperature/top-k/top-p; penalties, grammar, EOS; γ and adaptation
workloadprompt/source lengths, output lengths, domain/language/modality, batch/concurrency
latencycold/warm, prefill, draft, verify, reconciliation, time-to-first-token, inter-token, p50/p95, end-to-end
acceptanceaccepted-prefix histogram by proposal offset, tokens per target invocation, rejection causes
resourcespeak VRAM, target/draft KV, memory placement, power if relevant
correctnessgreedy exact IDs; sampling distribution tests; cache/logit parity; constraint/timestamp state
modality qualityVQA/OCR, WER/timestamps, codec IDs/audio quality, image metrics, task reward—as applicable

Report both decode-only and end-to-end speedup. Use warmups, synchronize device timing correctly, include the non-speculative target baseline in the same process, and publish regressions too. CPU structural tests can validate state and math; they do not prove CUDA kernel behavior, actual model quality, or GPU latency.

15One-page decision sheet

Is it applicable?

  1. Is generation autoregressive?
  2. Are several future conditions already known?
  3. Can the target score a causal candidate block?
  4. Is serial target decode material end-to-end?

Is it exact?

  1. Same events/order/conditioning?
  2. Same processed distributions?
  3. Greedy equality or ratio + residual?
  4. All state transactional?

Will it be fast?

  1. Measure draft cost.
  2. Measure block verify cost.
  3. Measure acceptance by offset.
  4. Check memory and batch contention.

How should I start?

  1. Greedy, batch one.
  2. Same tokenizer/codebook.
  3. γ = 3–4.
  4. Exact parity before sampling or fusion.
the transferable principle

A candidate is cheap; a commitment is expensive. Draft freely inside a frozen conditional world, let the authoritative model verify in parallel, commit only the valid prefix, and make state changes follow that decision. That is speculative decoding across every modality.

Primary sources and further study

  1. Leviathan, Kalman & Matias (ICML 2023), Fast Inference from Transformers via Speculative Decoding — exact stochastic algorithm and original systems analysis.
  2. Google Research, Looking back at speculative decoding — retrospective, intuition, and extensions beyond text.
  3. Zhou et al. (ICLR 2024), DistillSpec — draft/target alignment through task- and divergence-aware distillation.
  4. Cai et al. (ICML 2024), Medusa — multiple decoding heads and tree attention.
  5. Elhoushi et al. (ACL 2024), LayerSkip — early-exit self-speculative decoding.
  6. Sun et al., Block Verification Accelerates Speculative Decoding — verification need not be limited to a left-to-right rejection scan.
  7. Gagrani et al. (CVPRW 2024), On Speculative Decoding for Multimodal LLMs.
  8. SpecASR (2025 preprint) and Token Map Drafting for ASR (2025 preprint) — emerging audio-conditioned and model-free ASR approaches.
  9. So et al. (ICCV 2025), Grouped Speculative Decoding for Autoregressive Image Generation.
  10. Continuous Speculative Decoding for Autoregressive Image Generation — specialized treatment for continuous outputs.
  11. VADUSA (2024 preprint) — speculative AR speech synthesis with an explicit quality/speed tolerance mechanism.
  12. Hugging Face Transformers assisted decoding guide and vLLM speculative decoding docs — current framework interfaces; verify against the installed version.