lab

DSpark, every bolt and rivet

Semi-autoregressive speculative decoding with confidence-scheduled verification — the paper (arXiv 2607.05147, DeepSeek-AI + PKU, Jul 2026) and the DeepSpec reference implementation, explained from analogy to pseudocode, ending in a minimal plan to train one for a local VLM.

01TL;DR

Speculative decoding speeds up LLM generation by having a small draft model propose a block of tokens while a large target model verifies them in one parallel pass. Speedup = (T_draft + T_verify) / τ where τ is how many tokens get accepted per round. You win by drafting faster, drafting better, or verifying smarter.

DSpark attacks all three, and it has two independent ideas:

  • Semi-autoregressive drafting. Use a fast parallel backbone (DFlash-style, predicts all γ positions in one forward pass) but add a tiny sequential head (a low-rank Markov transition or an RNN) that re-weights each position's logits given the token that was actually sampled before it. Fixes the classic parallel-drafter failure (multi-modal collisions like "of problem") at ~1% latency cost. This is the draft-quality idea.
  • Confidence-scheduled verification. Train a tiny confidence head that predicts each draft position's acceptance probability, calibrate it (sequential temperature scaling), then at serving time choose per-request how many draft tokens to verify, maximizing expected system throughput given the engine's measured steps-per-second curve and current batch. This is the system-efficiency idea — crucial under concurrency, where verifying low-confidence tokens wastes batch capacity.

Results: +16–18% accepted length over DFlash and +27–31% over Eagle3 on Qwen3-4/8/14B; 60–85% faster per-user tokens in DeepSeek-V4 production at matched throughput; the sequential head adds only 0.2–1.3% latency.

Lossless by construction DSpark's semi-autoregressive head is a proper autoregressive factorization of per-token softmaxes, and its scheduler admits tokens causally (the admission decision for position k never sees token k or later). So the output distribution is exactly the target model's — no quality loss. The paper's Appendix A is a worked counterexample of what goes wrong if you let the scheduler peek.

02The problem: decode latency

An autoregressive LM emits one token per forward pass; wall time is proportional to output length. Speculative decoding (Chen et al. 2023; Leviathan et al. 2023) breaks the chain:

each round:
  1. draft model M_d proposes γ candidate tokens x_1..x_γ
  2. target model M_t verifies the block in ONE forward pass
  3. acceptance rule: at each position k, accept x_k w.p. min(1, p_t(x_k)/p_d(x_k))
     going left→right; the FIRST rejection at k discards x_{k+1}..x_γ
  4. the target then samples a "bonus" token at the first rejected position
     (or after the last accepted one) from the target's own distribution

The acceptance rule is rejection sampling, so the resulting distribution is exactly the target's — lossless. The average per-token latency is

L = (T_draft + T_verify) / τ          (1)

with τ = accepted tokens per round (including the bonus). Three levers: draft faster (lower T_draft), draft better (higher τ), verify smarter (spend T_verify only where it pays).

Drafter families

Autoregressive (EAGLE/Eagle3, MTP)

  • Each draft position conditions on previously sampled tokens → high τ per position.
  • But T_draft ∝ γ: drafting latency grows linearly with block size, so you use short blocks and shallow draft networks.

Parallel (Medusa, DFlash, DART)

  • All γ positions predicted in ONE forward pass → T_draft ≈ constant in γ.
  • But positions are independent → they marginalize over all possible predecessors and suffer multi-modal collision: when two continuations are plausible ("of course" / "no problem"), the drafter emits hybrids ("of problem"). Acceptance decays sharply along the block.

DFlash is the SOTA parallel drafter DSpark builds on. Its key trick is KV injection: during prefill, hidden states from m target layers are concatenated and projected into draft space, then injected into every draft layer's keys/values (see §5).

03Analogy: the writer's desk

Think of a famous novelist — call her the Editor — who writes every word of her story herself. She's brilliant but slow: each word takes her an hour, and she must re-read everything she's written to choose the next one (that's the target model: one token per forward pass).

Now imagine her staff of assistants. They can't improve her prose, but they can make her faster, because every word she writes is still her own — they only ever propose and schedule.

Speculative decoding = the intern

An intern reads the chapter so far and writes the next 7 words on a sticky note. The Editor reads the whole note at once (one parallel pass), nods through the words she would have written anyway ("course" ✓, "of" ✓...), and the moment a word isn't hers she crosses it out and writes her own replacement. The intern's rejected words are thrown away. Result: the Editor's words, faster.

Parallel drafter (DFlash) = the intern fills all 7 blanks at once

The intern writes all 7 words in one burst, each chosen from the chapter without seeing the other blanks. Fast. But when the story could go "of course" or "no problem", the intern writes "of problem". The Editor rejects the hybrid, and everything after it dies. Acceptance decays with every blank.

DSpark semi-AR = intern + whisperer

The intern still drafts all 7 blanks in one burst (keeps the speed). But a whisperer then walks the note left→right, very cheaply: after blank 1 gets "of", the whisperer nudges blank 2 toward "course" and away from "problem" (the Markov head: a tiny low-rank "what word typically follows this word?" table). A fancier whisperer keeps a notepad of the whole prefix (RNN head). The whisperer adds ~1% time because it's just a small lookup per position.

Confidence head = self-assessment stamps

Each draft word gets a stamp: the whisperer's estimate of "how likely is the Editor to accept this word given she accepted all the ones before it". It's trained to match the real acceptance odds, and then calibrated (STS) because interns are overconfident — they're taught to fix their stamp chart against a month of observed accept/reject outcomes.

Prefix scheduler = the assignment editor

A dispatcher decides how many of the intern's 7 words are even worth showing the Editor. Two inputs: the stamps (survival odds, chained as a product — the probability the whole prefix through word k survives is the running product of stamps), and the Editor's current workload. When the Editor's desk is empty, send everything — extra words are nearly free. When she's swamped with other authors' manuscripts (high concurrency / batch), only show the words with real survival odds, so her time isn't burned on doomed suffixes.

The one rule = no peeking

The dispatcher must decide word k's fate before seeing word k or anything after it. If the dispatcher peeked at how confident the intern was about word k+1 before admitting word k, the intern would game it, and the final manuscript would subtly stop being the Editor's own writing (Appendix A shows this quantitatively). Early-stopping in the greedy search is what enforces the rule.

The whole pipeline keeps the promise: the Editor's prose, word-for-word, just delivered faster — 60–85% faster per reader in the production deployment.

04Plain English

Idea 1 — draft with a parallel brain + a sequential tongue

DFlash showed you can generate a whole draft block in one pass if the draft model is fed rich features from the target model's own layers (KV injection). But each position is predicted in isolation, so drafts read like "of problem". DSpark keeps the parallel brain (so T_draft stays flat) and attaches a lightweight sequential tongue: after the parallel pass, positions are finalized left→right, each one adding a small "bias" computed from the token just sampled. The bias is either a low-rank bigram table (Markov head) or a tiny GRU (RNN head). Because the bias is just an additive logit shift, the final per-position distributions are still exact softmaxes — so rejection sampling still works and the whole scheme stays lossless.

A 2-layer DSpark drafter already beats a 5-layer DFlash drafter on accepted length — sequential dependency is a better use of parameters than depth.

Idea 2 — don't verify tokens that are going to die

Verifying a draft token costs target compute. Under light load that's fine; under concurrency, every token you verify occupies batch slots other requests need. So DSpark learns to predict each draft position's survival probability (confidence head), calibrates those predictions (they're overconfident by 3–8% ECE raw, ~1% after STS), and then a scheduler chooses the verification length per request by maximizing expected system throughput: Θ = expected_accepts × SPS(batch_size), where SPS is the engine's measured steps-per-second curve. The greedy search stops at the first throughput drop — that early stop is what keeps the scheme lossless (no peeking at future tokens).

Offline, this shows as a confidence threshold: raise it and acceptance rate climbs from 45.7% → 95.7% on chat. Online, it shows as a load-adaptive verification budget: ~4–6 tokens per request when the engine is idle, shrinking smoothly as concurrency rises.

05Architecture: semi-autoregressive drafting

5.1 The parallel backbone (DFlash)

Code: deepspec/modeling/dspark/qwen3/modeling.py (also gemma4/).

The draft model is a small stack of standard decoder layers (5 for the paper's Qwen3 configs; you can use 2). It shares the target's embedding and LM head (copied, then frozen). Two things make it "DFlash-style":

  • Context feature extraction. Hidden states from m target layers (config target_layer_ids=[1,9,17,25,33] for Qwen3-4B's 36 layers) are concatenated and projected:
    H_ctx = RMSNorm(W_c · concat(H[l1], ..., H[lm]))     # W_c: Linear(m·d → d)
    In the code this is self.fc + self.hidden_norm, applied once in _forward_backbone. The final target layer is excluded on purpose — transformers stores the final normalized hidden at the last index, which is inconsistent with the raw decoder outputs the training cache captures (see assert_no_final_target_layer).
  • KV injection. In every draft attention layer, the projected context features are concatenated with the draft block's own K/V along the sequence axis:
    K = concat([W_k · H_ctx, W_k · H_draft])   # same for V
    # attention within the block is BIDIRECTIONAL (is_causal=False),
    # and every draft position sees all context features.
    This is the code's k_ctx / k_noise / torch.cat(...) in Qwen3DSparkAttention.forward.

The anchor trick. Original DFlash feeds [anchor + γ masks] and predicts only the mask positions (γ logits from γ+1 inputs). DSpark's modification: the anchor itself is the first prediction position — γ input tokens [anchor + (γ−1) masks] produce γ draft logits. One less token of compute, same quality.

# draft block input (block_size = γ = 7 in configs)
draft_input_ids = [anchor_token, MASK, MASK, MASK, MASK, MASK, MASK]
# noise embedding: embed(anchor) at slot 0, embed(MASK) elsewhere
# forward the backbone ONCE over these γ tokens + injected context
# output: hidden h_1..h_γ  →  base logits U_k = lm_head(h_k)

The mask token id is 151669 in the Qwen3 config — any reserved/unused token id works.

5.2 The sequential stage

The parallel pass gives base logits U_k. The sequential head adds a prefix-dependent transition bias B_k(x_0, x_<k, x_k), producing a proper autoregressive factorization of the block distribution:

P(X | x_0) = ∏_k p_k(x_k | x_0, x_<k)
p_k(v | x_0, x_<k) = softmax_v( U_k(v) + B_k(x_0, x_<k, v) )      (4)

At inference you sample left→right: x_1 ~ p_1(·|x_0), then x_2 ~ p_2(·|x_0,x_1), etc. Each p_k is an ordinary softmax, so exact per-token probabilities are available for rejection sampling — this is the crucial property that CRF-style or CTC-style parallel heads lose (globally normalized / marginalized), and why DSpark stays lossless.

Because the sequential loop is inherently O(γ) tiny steps, the block must stay lightweight (T_sequential ≪ T_parallel). Two instantiations are in the paper, three in the code (deepspec/modeling/dspark/markov_head.py).

06Markov head & RNN head

6.1 Vanilla Markov head (paper default, code: VanillaMarkov)

Restrict the bias to depend only on the immediately preceding token — a first-order transition B(x_{k−1}, x_k). In principle that's a full V×V matrix; approximate it with a low-rank factorization B = W1 · W2 with rank r = 256:

# W1: Embedding(V → r)   W2: Linear(r → V, no bias)
B(x_{k-1}, ·) = W2( W1[x_{k-1}] )           (5)
# storage: (V·r + r·V) = for Qwen3 vocab 151936: ~2 × 39M params ≈ 156 MB bf16
# per step: one embedding lookup + one r×V GEMM — cheap

This is the "whisperer": once position 1 sampled "of", the bias at position 2 boosts "course" and suppresses "problem", fixing the parallel drafter's cross-mode collision.

6.2 Gated Markov head (code only — not in the paper's main text)

Same low-rank transition, but the previous-token embedding is gated by the backbone's hidden state at the current position:

gate = σ( W_g · [h_k ; W1[x_{k-1}]] )        # Linear(d + r → r)
bias = W2( gate ⊙ W1[x_{k-1}] )

6.3 RNN head (paper's second instantiation, code: RNNHead)

Removes the memoryless limitation: a recurrent state s_k accumulates the full prefix within the block. A single joint linear projects [s_{k−1}; W1[x_{k−1}]; h_k] ∈ R^{2r+d} into three r-wide chunks (gate, candidate, output) — one linear, then elementwise ops, matching paper eq. (6):

z_k    = [s_{k-1} ; W1[x_{k-1}] ; h_k]                    # 2r + d
g, c, o = chunk( W_joint · z_k, 3 )                       # each ∈ R^r
s_k    = σ(g) ⊙ s_{k-1}  +  (1 − σ(g)) ⊙ tanh(c)          # GRU-style update
B_k    = W2( tanh(o) )                                    # logit bias

The paper found the RNN head gives only marginal gains over Markov, mostly at longer proposal lengths, so Markov is the default (better deployment properties). Note the code's compute_step_bias for the RNN is stateless (zero-init state) — used only for compatibility; the real path is apply_block_logits / sample_block_tokens which unroll the state across the block.

6.4 Sampling (code: VanillaMarkov.sample_block_tokens)

def sample_block_tokens(base_logits, first_prev_token_ids, hidden_states, temperature):
    tokens, corrected = [], []
    prev = first_prev_token_ids                     # the anchor token id
    for k in range(block_size):
        step_logits = base_logits[:, k] + bias(prev, hidden_states[:, k])
        x_k = sample(softmax(step_logits / T))      # T=0 → argmax
        tokens.append(x_k); corrected.append(step_logits)
        prev = x_k                                  # ← the sequential dependency
    return stack(tokens), stack(corrected)

07Confidence head & calibration

7.1 The head

Code: deepspec/eval/dspark/confidence_head.py, AcceptRatePredictor in common.py. A linear projection + sigmoid over [backbone hidden h_k ; Markov embedding of the previous draft token] (when confidence_head_with_markov=True):

c_k = σ( w · [h_k ; W1[x_{k-1}]] )                 (7)   # scalar in (0,1)

c_k models the conditional probability that draft token k survives verification given all tokens before it were accepted. The training label is the analytical per-step acceptance probability — the total variation distance between draft and target distributions:

c*_k = 1 − ½·‖p_d − p_t‖₁                          (8)

Supervised with binary cross-entropy. Because each c_k is conditional, the joint probability that the whole prefix through k survives is the cumulative product:

a_k = ∏_{i ≤ k} c_i          # "prefix survival probability"

7.2 Post-hoc calibration: Sequential Temperature Scaling (STS)

The scheduler needs absolute magnitudes of cumulative acceptance (to compute expected throughput), not just rankings. Neural confidence heads are overconfident (raw ECE 3–8%; Figure 6 in the paper). STS fixes this left→right, one position at a time:

for k = 1..γ:                          # in order
    # keep positions 1..k−1 fixed at their calibrated values
    grid-search temperature T_k minimizing
        ECE of (∏_{i≤k} calibrate_i(c_i))  vs  observed prefix-acceptance
        on a held-out set of rollouts
# calibrate_i(c_i) = σ(logit(c_i) / T_i)   — temperature scaling, order-preserving

Temperature scaling preserves ranking, so the head keeps discriminating well (AUC 0.81–0.90) while ECE drops to ~1% and the cumulative products become trustworthy. In the repo, the eval loop (ConfidenceHeadRecorder) accumulates the raw material — per-position cumprod predictions vs observed accept-prefix labels, binned into ECE/AUC/Brier + reliability diagrams — but the STS grid search itself is production-side, not in DeepSpec.

08Hardware-aware prefix scheduler

Now the system idea. Consider R active requests; request r has per-position confidences c_{r,1..γ} and a scheduled verification length ℓ_r ∈ {0..γ}. The total verification batch (in tokens) sent to the target is B = Σ_r (1 + ℓ_r) (the "1" = the anchor, which is always verified — it also produces the bonus). Expected accepted tokens: τ = Σ_r (1 + Σ_{j≤ℓ_r} a_{r,j}) where a_{r,j} = ∏_{i≤j} c_{r,i}.

Let SPS(B) = engine throughput in steps/second for a forward pass with B tokens — profiled once at startup into a small cost table. Expected system throughput: Θ = τ · SPS(B). Maximize Θ over all length assignments. Because a_{r,j} is non-increasing in j, extending a request by one position has marginal gain exactly a_{r,j} — so the objective has a natural greedy admission structure (Algorithm 1):

# Algorithm 1 — Hardware-Aware Prefix Scheduler (paper, single-node greedy form)
1. for each request r:  a_{r,j} ← ∏_{i≤j} c_{r,i}   for j = 1..γ
2. candidate pool E ← {(r, j) : a_{r,j} > 0}, sorted by a_{r,j} DESCENDING
3. init: ℓ_r ← 0;  B ← R;  τ* ← R;  Θ_best ← R · SPS(R);  ℓ*_r ← 0
4. for (r, j) in E (sorted order):
5.     ℓ_r ← j;  B ← B + 1;  τ* ← τ* + a_{r,j}
6.     Θ ← τ* · SPS(B)
7.     if Θ > Θ_best:  Θ_best ← Θ;  ℓ* ← ℓ
8.     else:  break                       # ← early stop (the losslessness rule)
9. return ℓ*_1..ℓ*_R

Notes:

  • Sorting the global pool by survival probability naturally respects prefix dependencies: a request's tokens are admitted in order because its a-values are monotone decreasing.
  • Unimodality assumption. The early-stop break finds the global max only if Θ is unimodal, i.e., SPS(B) decays smoothly. Real hardware curves are jagged (step-wise degradation) — see §12's production adaptation.
  • Non-anticipation. The break means the scheduler never evaluates admission for position k+1 (whose confidence depends on the realized token x_k) before committing to position k. Without the break, admission of x_1 could depend on the sampled value of x_1 itself — selection bias, and the output distribution drifts off the target's (Appendix A, next section).

8.1 Single-request form (what you'd actually build locally)

For batch = 1 the scheduler degenerates to: choose ℓ maximizing Θ(ℓ) = (1 + Σ_{j≤ℓ} a_j) · SPS(1 + ℓ) over ℓ = 0..γ, breaking at the first drop. You can profile SPS on your own GPU once (measure steps/sec of the target for 1..γ+1 tokens) and hard-code the table. Offline, the repo's eval path instead uses a simple static threshold on the per-position confidence (see §11).

09The Appendix A counterexample

Why is early stopping a correctness requirement, not just a heuristic? The paper proves it with a concrete case. Single request, γ = 2, first-position survival a₁ = 0.8, and a profiled capacity curve:

SPS(1) = 1.0,  SPS(2) = 0.5,  SPS(3) = 0.45

Θ(ℓ=0) = 1 · SPS(1)              = 1.0
Θ(ℓ=1) = (1 + 0.8) · SPS(2)      = 0.9

Now suppose the scheduler does not stop and evaluates ℓ=2 before committing anything. The second survival probability is a₂ = a₁·c₂, and c₂ depends on the realized first token x₁ (the Markov confidence head uses x₁). Two cases:

Case 1: x₁ yields c₂ = 0.9

a₂ = 0.72 → Θ(2) = (1 + 0.8 + 0.72)·0.45 = 1.134, the global max → ℓ = 2, and x₁ is admitted.

Case 2: x₁ yields c₂ = 0

a₂ = 0 → Θ(2) = 0.81, global max is Θ(0) = 1.0 → ℓ = 0, and x₁ is not admitted.

So whether the first draft token gets admitted depends on the value of the first draft token itself. With vocabulary {A, B}, target p_t = (0.7, 0.3) and draft p_d = (0.5, 0.5): the acceptance rule gives x₁ = A always survives (min(1, 0.7/0.5) = 1), and x₁ = B is rejected, the target re-samples from p_t. The retrospective scheduler's output distribution becomes

Pr(Y = A) = 0.5·1 + 0.5·0.7 = 0.85,   Pr(Y = B) = 0.15

— i.e. (0.85, 0.15), different from the target's (0.7, 0.3). The scheme is no longer lossless. The causal greedy scheduler avoids this: since Θ(1) = 0.9 < Θ(0) = 1.0 it breaks immediately and returns ℓ = 0 without ever evaluating c₂, so the admission decision for x₁ used only pre-token information.

10Training

10.1 Data pipeline (scripts/data/, README.md)

  • Prompts: Open-PerfectBlend (1.3M samples: chat 17.6%, math 39.4%, code 38.9%, instruction-following 4.1%). Only prompts are used.
  • Regenerate answers with the target model (SGLang server; Qwen3 recommended sampling: temperature 0.7, top-p 0.8, top-k 20). Drafts are trained on the target's own distribution — that's what makes them good at being verified by it.
  • Target cache (deepspec/data/target_cache_dataset.py): precompute and mmap, per training sequence:
    • input_ids, loss_mask (assistant tokens only)
    • target_hidden_states: concatenated hidden states at target_layer_ids, captured via forward hooks on the target's decoder layers (scripts/data/prepare_target_cache.py) — this is the KV-injection feature
    • target_last_hidden_states: the final-norm hidden (pre-LM-head) so the draft workers can compute target logits locally with the shared frozen lm_head, instead of transferring full-vocab logits across workers (paper §5.1: O(d) communication instead of O(V))
Storage warning The default Qwen3-4B cache is ~38 TB (README). The paper's production training instead packs anchor-bounded blocks (see below) and streams hidden states; for a local project you'll want short sequences, few feature layers, and a modest corpus (see §14).

10.2 Anchor sampling & packing

During training, each sequence contributes num_anchors = 512 random anchor positions (common.py: sample_anchor_positions). An anchor is valid only if both the anchor token and its next token are loss-enabled (i.e., inside the assistant answer). For each anchor, the draft input is [anchor_emb, MASK×(γ−1)] at positions anchor_pos .. anchor_pos+γ−1, labels are the next γ ground-truth tokens. All anchors of one sequence are packed into the same batch row and separated by a custom flex_attention block mask:

# dspark_mask_mod (common.py):
#   draft query in block b may attend to:
#     - context tokens with kv_idx < anchor_pos[b]      (causal context)
#     - draft tokens inside its OWN block               (bidirectional)
#   (blocks are fully isolated from each other)
# position ids: context 0..seq_len−1, then each block gets anchor_pos + [0..γ−1]

This is the paper's "anchor-bounded sequence packing" via token-level attention indices — the draft's compute cost decouples from context length, and padding is avoided.

10.3 Losses

All three terms are position-weighted with w_k = exp(−(k−1)/γ) (code: loss_decay_gamma = 4.0), emphasizing early block positions, and masked by eval_mask — the contiguous enabled prefix per block (a cumprod over per-position validity, so a block is supervised only up to its first disabled position):

L_ce   = − Σ_k w_k · log p_d(x*_k)                          # predict ground-truth token   α=0.1
L_tv   =   Σ_k w_k · ‖p_d − p_t‖₁                            # match target distribution     α=0.9
L_conf = − Σ_k w_k · [c*_k log c_k + (1−c*_k) log(1−c_k)]    # confidence BCE               α=1.0

L = 0.1·L_ce + 0.9·L_tv + 1.0·L_conf                          (12)

Notes:

  • p_t comes from applying the frozen shared lm_head to the cached target_last_hidden_states gathered at label_pos − 1 (the target's prediction for the label position).
  • TV distance is the direct proxy for acceptance rate: per-step acceptance = 1 − ½·TV. Minimizing L_tv maximizes expected τ — this is the "train for verifiability" term.
  • Teacher forcing: the Markov head's previous-token input at train time is the ground-truth previous token (prev_token_ids = [anchor, x*_1..x*_{γ−1}]), not the sampled one.
  • Loss normalization is done with all-reduced denominators across ranks (correct global mean when local batch sizes are tiny).

10.4 Config & hyperparameters (config/dspark/dspark_qwen3_4b.py)

FieldValueMeaning
block_size7γ draft positions per round
num_draft_layers5draft backbone depth (Eagle3 used 1)
target_layer_ids[1,9,17,25,33]KV-injection features from target
mask_token_id151669reserved token for blank slots
num_anchors512sampled blocks per sequence per step
markov_rank / type256 / vanillalow-rank transition head
loss_decay_gamma4.0position weight decay
loss alphas0.1 / 0.9 / 1.0ce / tv / confidence
confidence_head_with_markovTrueconfidence input = [h_k; W1[x_{k-1}]]
lr / warmup / wd6e-4 / 4% / 0AdamW
precision / compilebf16 / torch.compileflex_attention for the block mask
epochs10on Open-PerfectBlend

The draft is trained from scratch except the shared embedding + lm_head, which are copied from the target and frozen (initialize_embeddings_and_head).

11The inference loop

Code: deepspec/eval/base_evaluator.py (loop + verification), eval/dspark/evaluator.py + draft_ops.py (DSpark specifics).

11.1 One decoding round

# state: target KV cache (committed tokens), draft KV cache (context features),
#        target_hidden_states = features of the newly accepted span

propose (draft_ops.build_dspark_proposal):
  1. draft_input_ids = [anchor(=last committed token)] + [MASK]*(γ−1)
  2. h = draft_backbone(draft_input_ids,
                       injected_ctx = project(target_hidden_states),   # KV injection
                       draft KV cache, positions = cache_len..start+γ)
     # cache trick: the draft's DynamicCache holds PROJECTED CONTEXT FEATURES
     # for committed tokens; each round appends the new span's features then
     # crops back to the committed length (past_key_values_draft.crop(start))
  3. U_k = lm_head(h_k)                                  # base logits
  4. sample x_1..x_γ left→right with the Markov/RNN bias  (sequential loop)
     + confidence c_k = σ(w·[h_k; W1[x_{k-1}]]) per position
  5. schedule: prefix length ℓ ← first index with c_k below threshold
     (or the full hardware-aware / Algorithm-1 form in production)
  6. if ℓ == 0: return a degenerate proposal (verify only the anchor)

verify (base_evaluator.verify_draft_tokens):
  1. target forward over [anchor, x_1..x_ℓ] in ONE pass (with target KV cache)
  2. p_t = softmax(target_logits / T)   (T=1 in eval; T≈0 → greedy one-hot)
  3. accept_prob = clamp(p_t(x_k)/p_d(x_k), max=1) for each k
  4. accept_mask = (rand < accept_prob)  →  prefix = cumprod(accept_mask)
     accepted = count of leading 1s      # first 0 truncates the rest
  5. next_token = sample(p_t) at the first rejected position
                  (or at the last position if all accepted = bonus)
  6. commit: [x_1..x_accepted, next_token]  →  advance start, crop target KV

update (evaluator._update):
  target_hidden_states ← features of the newly committed span (accepted+1 rows)
  → next round's injection

The eval harness is single-request (bsz=1), temperature 1.0, and reports acceptance_length = accepted drafts + 1 bonus, plus per-position acceptance rates (denominator = rounds where the position was proposed). The repo's threshold mode is the offline diagnostic; the full Algorithm-1 scheduler with the SPS table is production-side (not in DeepSpec).

11.2 Why the draft cache can just hold context features

Because the block attention is bidirectional and non-causal, and the only context the draft needs is the projected target features of committed tokens, the draft's KV cache is exactly those features — it never needs to re-run over the whole context, and it never caches draft-block K/V across rounds (those are cropped away after each forward). This is what makes the draft cheap at long contexts.

12Results

12.1 Offline: accepted length τ (Table 1, macro-avg over 9 benchmarks)

TargetDrafterMathCodeChatvs DFlashvs Eagle3
Qwen3-4BEagle34.563.872.40——
DFlash4.804.442.95——
DSpark5.575.123.49+16.3%+30.9%
Qwen3-8BEagle34.664.152.58——
DFlash4.774.462.97——
DSpark5.655.283.50+18.4%+26.7%
Qwen3-14BEagle34.523.992.52——
DFlash4.744.452.92——
DSpark5.635.243.47+18.3%+30.0%

(Row values are benchmark means within each domain; full per-benchmark table is in the paper. All drafters retrained in the same framework/data, Eagle3 TTT horizon = 7 matched to block 7, same target feature layers, 5 draft layers for DFlash/DSpark vs 1 for Eagle3. τ includes the bonus token.)

12.2 Position-wise acceptance (Figure 2) — the "why"

  • Position 1: parallel drafters win — DFlash 0.88 vs Eagle3 0.81 (math), 0.72 vs 0.53 (chat), because a single draft forward allows deep networks while AR drafters must stay shallow (T_draft ∝ γ). The first token is the highest-leverage one (a rejection kills the block).
  • Later positions: Eagle3 rises or stays flat (0.53 → 0.74 on chat), DFlash decays (0.87 → 0.78 code, 0.72 → 0.63 chat) — multi-modal collision.
  • DSpark: starts at 0.93 (math) and stays high and stable across the block — inherits the deep parallel capacity, fixes the suffix decay with the sequential head.

12.3 Ablations

  • Depth: DSpark improves with draft layers; a 2-layer DSpark already beats 5-layer DFlash everywhere — sequential modeling is a better parameter spend than depth (Figure 3).
  • Proposal length: DSpark's edge over DFlash grows with γ: at γ=7 it's +16%/+15%/+18% (math/code/chat); at γ=15, +30%/+26%/+22%. RNN head gives only marginal extra gains, mostly at long γ — hence Markov as default (Figure 4).
  • Latency overhead: the sequential loop adds 0.2% (γ=4) to 1.3% (γ=16) of round latency at batch 128 (Figure 4, right).
  • Confidence head quality: threshold sweep shows acceptance rate climbing 45.7% → 95.7% (chat), 76.9% → 92.5% (math), 67.6% → 92.0% (code) as the threshold rises (Figure 5). Raw head: AUC 0.81–0.90 but ECE 3–8%; after STS, ~1% ECE (Figure 6).

12.4 Production (DeepSeek-V4-Flash / V4-Pro preview)

  • Backbone: 3 MoE layers with mHC + sliding-window attention 128; γ = 5; Markov head; confidence head trained end-to-end + STS.
  • vs the MTP-1 production baseline: +60–85% (Flash) / +57–78% (Pro) per-user tokens at matched throughput; +51% / +52% aggregate throughput at moderate SLAs (80 / 35 tok/s/user).
  • At strict SLAs (120 Flash / 50 Pro tok/s/user) the baseline's concurrency collapses; DSpark keeps serving — nominal +661% / +406% throughput, read as "extends the feasible interactivity frontier" rather than a well-utilized comparison (Figure 7).
  • Verification budget adapts to load: 4–6 tokens/request when idle → shrinks smoothly as concurrency rises (Figure 8).

12.5 Production adaptations of the scheduler (paper §5.2)

  • Jagged SPS curves: real capacity is step-wise, breaking the unimodality assumption. Production removes the early-stop break and does a global search — safe because of the next point.
  • Asynchronous scheduling / ZOS: CUDA-graph replay and zero-overhead scheduling need the next batch size before the current step ends. Production computes the truncation length K from confidence outputs two steps prior (dynamic top-K admission). The two-step-old information is the causal barrier: it can't see the current tokens, so losslessness holds even with a global search.
  • Variable-length verification kernels: tokens across requests are flattened and processed as independent elements; intra-sequence structure is carried by a marker tensor in the sparse-attention implementation (only the index-attention and compress kernels changed on V4).

12.6 Limitations (from the paper)

The draft block itself costs a fixed parallel pass; for requests with inherently low acceptance (complex queries), that compute is unrecoverable. Future work: difficulty-aware early exit in the draft model.

13DeepSpec repo map

DeepSpec (MIT) contains Eagle3, DFlash and DSpark in one framework (adapted from SpecForge). The DSpark-relevant files:

FileWhat lives there
config/dspark/dspark_qwen3_4b.pyAll hyperparameters (block 7, 5 layers, layer ids, ranks, alphas, lr...)
modeling/dspark/common.pyAnchor sampling, block mask, eval mask, noise embedding, position ids, confidence head class
modeling/dspark/qwen3/modeling.pyDraft model: attention w/ KV injection, decoder layer, backbone forward, logits, sampling entry points
modeling/dspark/markov_head.pyVanilla / Gated / RNN heads + sample_block_tokens
modeling/dspark/loss.pyCE + TV/L1 + confidence BCE, position weights, all-reduced denominators, τ metric
trainer/dspark_trainer.py, base_trainer.pyStep loop: model forward → loss
eval/dspark/evaluator.py, draft_ops.pyProposal construction (draft cache, sampling, confidence prefix)
eval/base_evaluator.pySpec-decode loop + rejection-sampling verification (algorithm-independent)
eval/dspark/confidence_head.pyConfidence recording: per-position ECE/AUC/Brier, reliability diagrams
data/*, scripts/data/*Target cache (mmap shards), collator, CUDA prefetcher; data prep (download → regen → cache)

Released checkpoints: deepseek-ai/dspark_qwen3_{4b,8b,14b}_block7 and deepseek-ai/dspark_gemma4_12b_block7 (plus eagle3_*_ttt7 and dflash_*_block7).

14Minimal local-VLM implementation plan

Goal: a DSpark drafter for a small VLM (e.g. the ~2–4B class) that trains and serves on a single 6 GB GPU, and can later slot into the kaminari-style engine. The key insight that makes this easy: the drafter is a text-decoder-only model. It never sees pixels — vision enters through the KV-injection features of the target decoder, which already encode image content after prefill. So the DSpark recipe transfers to any VLM whose decoder is a standard transformer.

14.1 Milestones (each independently verifiable)

  1. Verification harness first. Implement the spec-decode loop with a trivial drafter (e.g. prompt-lookup or the target's own greedy continuations). Prove losslessness: greedy target-only vs spec-decode outputs must be bit-identical (T≈0), and rejection sampling must reproduce target statistics. Build the τ / acceptance-rate metrics.
  2. Target cache. Hook m=3–5 decoder layers (never the last), regenerate answers on a modest prompt corpus (few thousand), store input_ids/loss_mask/hidden_features/last_hidden in bf16 mmap. Sizing on the 3050: seq 1024, d≈2048, 5 layers → ~40 MB/sample; a 5k-sample cache is ~200 GB — so cap sequence length (768–1024) and consider 3 layers or fp16, or train on a slice first. Better: cache only the anchor windows you'll actually train on (paper §5.1).
  3. Draft model + training. 2-layer backbone (2L beats 5L DFlash!), vanilla Markov rank 128–256, confidence head on, block 7, num_anchors 32–64 (VRAM), bf16, torch.compile; loss alphas 0.1/0.9/1.0, lr ~6e-4, warmup 4%. Watch the logits tensor: anchors × γ × vocab × 2 bytes can blow 6 GB — chunk the lm_head application if needed.
  4. Eval & serving. Single-request scheduler (argmax over ℓ of (1+Σa)·SPS(1+ℓ) with a locally-profiled SPS table, early-stop break), or plain threshold mode from the repo. Then swap into kaminari: verification is just a target forward over the scheduled prefix.

14.2 Gotchas (learned from the code)

  • Never include the final target layer in target_layer_ids (norm inconsistency between transformers' output_hidden_states and your cache hooks).
  • M-RoPE VLMs: Qwen3-VL-style models use 3D position embeddings. The draft uses 1D positions (create_position_ids). Fine if you keep the draft's positions aligned with the text-stream positions of the target; don't reuse target position tensors directly.
  • Mask token id: pick a truly unused token (151669 works for Qwen3; verify for your tokenizer), and exclude it from generation templates.
  • loss_mask must cover exactly the assistant answer; anchors need the next token enabled too.
  • Eval temperature consistency: train targets are T=1 softmaxes; evaluate at the same T for matched acceptance.
  • Thermal throttling on small GPUs (memory's standing lesson): benchmark fresh processes, interleaved A/B, report min — the 3050 throttles after ~2 consecutive runs.
  • Copy baseline logic verbatim when porting; the paper's "simplified" versions of masking/anchoring are easy to get subtly wrong.

15Interactive scheduler simulator

Drag the per-position conditional confidences (the c_k from the confidence head). The simulator chains them into survival probabilities, evaluates the greedy single-request schedule against a profiled SPS curve, and shows what a retrospective (peeking) scheduler would pick instead.

0.85
0.75
0.62
0.48
0.35

—