lab

The inference-efficiency playbook — post-training and serving for faster models

A working map of what an inference/ML-efficiency engineer actually does: the cost model, the six families of post-training that buy serving speed, the training-free compression toolbox, the systems layer — and the concrete 2026 production numbers from fal's research blog and Baseten's model-performance posts that prove the recipe works end to end. Includes the full catalog of post-training-for-efficiency methods (§02) the non-training engineering map across LLMs, diffusion, audio and VLMs (§05), the open engine room's production numbers from vLLM's engineering blog (§09), and how to run the whole post-training-for-efficiency loop on a 6 GB laptop — including Aug-2026 cloud-rental rates for the runs that need more (§12) — plus a worked application to the audio lane: the fal recipe on Stable Audio 3, and mini-DSpark drafters for speech-token models (§13).

01The mental model everything hangs on

Every technique attacks one of four multiplicative terms:

latency/cost ≈ (tokens generated) × (bytes moved per token) × (time per byte) ÷ (parallelism amortized over batch)
  • Decode is memory-bound. One token at a time; every step re-reads all weights plus the whole KV cache. Arithmetic intensity is ~O(1) per matmul. Anything that shrinks bytes read per token (quantization, sparsity, KV compression, fewer KV heads) is nearly free speedup at small batch.
  • Prefill is compute-bound. O(n²) attention, full matmuls. Attacked with chunked prefill, sparse attention, prefix caching — a different toolbox.
  • Batch changes the regime. At large batch, decode approaches compute-bound: quantization's win shrinks, KV capacity and scheduling dominate. A trick that wins at batch 1 can lose at 64.
  • Token count is a hidden dimension. Generating 2000 reasoning tokens to say "yes" is 2000× the decode cost. Post-training can shrink this directly — most people forget it.
the frontier is compositionality The strongest 2026 systems stack everything at once — low-bit weights + compressed/paged KV + prefix reuse + continuous batching + speculative decoding — and the open research is making the layers cooperate rather than fight (e.g. QuantSpec, arXiv:2502.10424: self-speculative decoding over a hierarchical quantized KV cache, >90% acceptance, ~2.5× end-to-end).

02Post-training routes that buy inference speed

Changing the weights by training so the deployed artifact is cheaper. Six families:

2.1 Distillation into a smaller student

  • Classic soft-target KD (Hinton, 1503.02531) → sequence-level KD (Kim & Rush, 1603.09785) → on-policy distillation that fixes the train/inference mismatch (GKD, 2307.13386; DistiLLM, 2401.10491) and reverse-KD so students stop covering modes they can't (MiniLLM, 2305.08445).
  • Prune-then-adapt: Sheared-LLaMA (2310.06694) — structured 3D pruning + continued training beats from-scratch small models at 25% of the compute; Minitron (2407.14679) — prune width/depth and MoE experts, then a short KD re-alignment (~3% of original training cost).
  • The canonical inference-aware ASR example: distil-whisper (2311.00430) — decoder-only distillation of Whisper (6→2 decoder layers, vocab cut, synthetic alignment retraining): 4.9× faster at <1% WER loss, plus a cascaded text-only correction model.
  • Fal's pragmatics (see §07): online KD, offline KD and plain PEFT landed at the same quality on their task — so ship PEFT, it's the cheapest.

2.2 Training the model to emit fewer tokens

  • Concise-reasoning / overthinking control: token-budget penalties in RL (2403.05587), CCoT curriculum (2412.20993), L1 length-controlled DPO (2503.04697), Chain of Draft (2502.18600) — think in ≤5-word drafts, ~90% token cut at near-parity accuracy.
  • Vocabulary work: MiniLLM-style vocab distillation shrinks the lm_head (a 151k-row embedding is real decode cost); tokenizer retraining cuts token counts 20–30% for CJK/multilingual.
  • Adaptive computation: CALM early-exit adapters (2205.14135) and Meta's LayerSkip (2404.16710) — train draft-at-bottom/verify-at-top inside one model; self-speculation by construction.

2.3 Training the decoding algorithm itself (MTP & draft heads)

The most active 2024–2026 front — detailed in §03. Post-hoc head-training jobs cost GPU-hours, not runs: Medusa (2401.10774), EAGLE-3 (2503.01840), DFlash (2602.06036), DSpark; DeepSeek-V3 (2412.19464) proved MTP heads trained in the main run serve as a free draft module; post-hoc MTP self-distillation (2602.06019) retrofits the trick onto existing checkpoints (>3× decode, <5% accuracy loss).

2.4 Quantization-aware post-training

You cannot PTQ your way below 4 bits cleanly — see §04 for the ladder: learned rounding (LSQ 1810.01875, GMLQ 2102.03586), EfficientQAT (2407.11062: ~25% of usual QAT cost), progressive UPQ FP16→INT4→INT2 (2506.09104), end-loss-guided PTQ GuidedQuant (ICML 2025), and Quantization-Aware Distillation (2601.20088): KL from the bf16 teacher into an NVFP4 student — the current default at the frontier when you can't rerun the original SFT/RL recipe.

2.5 Architecture surgery + fine-tune recovery

  • MHA→GQA uptraining (2305.13245): shrink KV heads with a small fine-tune; KV cache shrinks proportionally.
  • Sliding-window + global interleaves (Gemma recipe) → bounded KV for streaming.
  • MLA low-rank KV latents (DeepSeek-V2, 2405.04434): ~93% KV reduction — wants training in the loop.
  • Layer/block pruning with cheap recovery: LLM-Pruner (2305.15286), ShortGPT (2310.11102 — delete 25% of layers in minutes), SliceGPT (2401.15024 — orthogonalize then slice, no retrain).
  • MoE compression: expert pruning/merging (I-MoE 2405.04534, Minitron expert slimming).

2.6 Inference-aware post-training: optimize the procedure, not the likelihood

  • InfAlign (2412.19792): align against the actual test-time policy (best-of-N), not single-sample quality.
  • Diffusion step-distillation is the same philosophy with a sampler instead of a decoder: consistency models (2303.01469), progressive distillation (2202.00512), DMD/DMD2 (2311.18828 / 2405.14867), adversarial ADD → SDXL-Turbo (2311.17042), MeanFlow (2502.06085). Deployed artifact runs 4 steps, not 50 — and the objective was chosen because of the sampler. See §06 for the 2026 production version.
  • Model merging (DARE-TIES, soups) as cheap capability recovery after an efficiency fine-tune ate generality.

2.7 The full catalog — every post-training method that buys inference speed

Grouped by which term of the §01 cost equation it attacks. A method is listed only if it needs gradient steps on real data — pure training-free compression lives in the §04 toolbox. Every entry links to its representative paper.

familymethodwhat inference gainskey refs
student size soft-logit KD → studentsmaller model ≈ proportional decode + KV cut1503.02531
sequence-level KD (synthetic teacher data)same, without logit access1603.09785
on-policy KD (student samples, teacher scores)kills exposure bias on long generationsGKD 2307.13386 · DistiLLM 2401.10491
reverse-KD generative distillationsmaller student, no mode-averaging mushMiniLLM 2305.08445
speculative KD (teacher as the drafter and target)student + built-in acceptance lift in one job2404.09728
task distillation of a whole pipeline (ASR)distil-whisper: 4.9× at <1% WER loss2311.00430
prune-then-adapt (width/depth/head/experts + KD recovery)cheap small model from a big oneSheared-LLaMA 2310.06694 · Minitron 2407.14679
layer/block deletion + short retraindepth cut w/ recovery; SliceGPT needs no retrainShortGPT 2310.11102 · LLM-Pruner 2305.15286 · SliceGPT 2401.15024
MoE expert pruning / merging / shrinkingsmaller active set + denser expert tablesI-MoE 2405.04534 · Minitron
tokens generated length-controlled RL / DPO (budget or penalty in the reward)linear latency cut on reasoning-heavy traffic2403.05587 · L1 2503.04697 · CCoT 2412.20993
concise-thought formats (Chain-of-Draft)~90% fewer CoT tokens at parity2502.18600
latent / continuous reasoningthinking in fewer, denser stepsCoconut 2412.06769
vocabulary / tokenizer shrink for the deployment language mixfewer decode steps per unit of meaningMiniLLM (vocab) · tokenizer retrain
early-exit training (per-layer confidence)easy tokens skip deep layersCALM 2205.14135 · LayerSkip 2404.16710
draft / verify structure parallel decoding heads (frozen backbone)self-drafting, ~1–2×Medusa 2401.10774 · Hydra 2402.05109
feature-space autoregressive draft headhigher acceptance; ~2× ceilingEAGLE 2401.15077 · -2 · -3 2503.01840
MTP module trained with the main modeldraft built-in; free at servingDeepSeek-V3 2412.19464
post-hoc MTP via online self-distillationretrofit existing ckpts, >3× decode2602.06019
block-diffusion parallel drafter (+ Markov refine heads)~3×; DFlash 4.6 acceptanceDFlash 2602.06036 · DSpark (DeepSeek)
double early-exit adapter (self-speculative)no draft model, losslessKangaroo 2404.18911 · Draft&Verify 2309.08168
warm-started drafter retrain on task outputsacceptance recovery for fine-tuned targetsfal DSpark post
number format learned rounding / quant (QAT)int8→int4 weights+acts deployable-nativeLSQ 1810.01875 · GMLQ
budgeted / progressive QAT recipesQAT at ~25% cost; INT2 instruction modelsEfficientQAT 2407.11062 · UPQ 2506.09104
quantization-aware distillation (teacher → quantized student)recover NVFP4/FP4 quality post-SFT/RLQAD 2601.20088 · GuidedQuant ICML'25
rotation + quant fine-tuning (incoherence processing)makes 2–4 bit tractableQuaRot 2404.00456 · SpinQuant
ternary / 1.58-bit trainingextreme bandwidth cut (BW GEMM)BitNet b1.58 2402.17764
sparsity-aware fine-tune (2:4 semi-structured)real 1.3–1.7× on Ampere+ siliconMaini et al. FGS; NVIDIA blog
attention / KV structure MHA→GQA uptrainingKV heads ÷N w/ small fine-tuneGQA 2305.13245
MLA low-rank KV-latent conversion~93% KV cutDeepSeek-V2 2405.04434
sliding-window + global interleaving fine-tunebounded KV for streamingGemma 2/3 recipe
KV-quantization-aware fine-tuningtolerates 2–4 bit cache at long contextKVQuant-style + in-loop fake-quant
sampler / procedure CFG / guidance distillation into one branchhalves DiT passes per step2207.12598 · fal Ideogram §06
progressive step distillationsteps ÷2 per stage2202.00512
consistency models / LCM / TDM / MeanFlow1–8 step samplers2303.01469 · LCM 2310.04378 · TDM 2503.06674 · MeanFlow 2502.06085
distribution matching (DMD/DMD2; GAN-first variants)few-step, data-free; the shipped recipeDMD 2311.18828 · DMD2 2405.14867
adversarial distillation (ADD)1-step image gen (SDXL-Turbo)2311.17042
inference-aware alignment (train vs best-of-N policy)quality at the deployed procedureInfAlign 2412.19792
multimodal / recovery / economics VLM visual-token compression w/ tuning÷14 image prefill tokensLLaVA-PruMerge 2403.15388
single-step distilled TTS / ultra-small speech LMssub-150 ms first-soundConsistencyTTS · Chatterbox Turbo
model merging for capability recovery (soup/TIES/DARE)restores generality lost to efficiency tunes — no training2204.06504 · TIES 2306.01708 · DARE 2311.03099
LoRA / adapter economics (fine-tune per tenant, share base)marginal model cost ≈ 0 per customerLoRA 2106.09685 · S-LoRA 2311.03285

03Speculative decoding: the 2026 frontier

Exact greedy/sampling acceptance (Leviathan 2211.17192, Chen 2302.01318) preserves the target distribution; the engineering race is all about acceptance length and drafting overhead:

independent draft model
(original specdec)
→ parallel heads on frozen backbone
(Medusa, 2401.10774)
→ autoregressive draft in feature space
(EAGLE-3, 2503.01840 — ~2× ceiling)
→ MTP heads trained with the model
(DeepSeek-V3; Qwen3.6 ships them)
→ block-diffusion parallel draft
(DFlash, 2602.06036 — ~3×)
→ DFlash + Markov refine heads
(DSpark, DeepSeek — 4.6 acceptance)
  • DFlash's argument: EAGLE is autoregressive — drafting 8 tokens costs 8 draft passes and errors accumulate, capping ~2×. DFlash instead uses a deeper draft model with bidirectional attention that unmasks a whole block of γ∈[8,16] tokens in one pass. The single pass is 2–4× slower than EAGLE's — but it replaces EAGLE's entire draft phase and drafts more, so it wins (~3× on Qwen3-8B / B200).
  • Training detail worth stealing (Baseten): condition on fused hidden states from 5–6 evenly spaced target layers; freeze embed + lm_head; random anchors; CE loss with exponential position decay exp(−(i−1)/γ) — earlier draft tokens matter more.
  • Fine-tuned models need retrained drafters. Fal (§07): an off-the-shelf MTP/DFlash head trained on the base model under-performs on the fine-tuned target; warm-started retraining on task data fixed acceptance. And drafters for copy-heavy workloads (OCR, ASR) get another free option: prompt-lookup decoding (2311.03511) — zero-cost drafts from repeated spans.
  • Self-speculative, no extra params: Draft & Verify layer-skipping (2309.08168), Kangaroo's double early exit (2404.18911), speculative streaming (reuse the target’s own recent hidden states as drafts), QuantSpec (2502.10424). Per-request speculation-length RL: RESDO. Hybrid draft-large/verify-small: Speculative Cascades (2507.20127).
  • The catch that papers bury: speculation degrades at large batch — draft FLOPs eat what bandwidth savings bought, and host-side overhead kills the win. Baseten's fix: run draft + verify as one fused forward so overhead is ~zero (§08).

04The training-free toolbox (PTQ) and where it breaks

AttackMethodsPractical note
weightsGPTQ (2210.17323), AWQ (2306.00978), bnb NF4, Marlin fused-dequant GEMMsW4A16 ≈ lossless on LLMs; if the dequant kernel is slow the win evaporates — the kernel is the technique
weights + activationsSmoothQuant W8A8 (2211.10438), OmniQuant, FP8 / NVFP4 / MXFP8 on BlackwellFP8 basically free on Ada/Hopper+; ≤4-bit activations need QAT/QAD (§2.4)
sparsityWanda (2306.12929), SparseGPT 2:4 semi-structured, SliceGPT2:4 gives real 1.3–1.7× on Ampere+; unstructured remains mostly paper speedups
low-rankSVD-LLM (2403.05170), QuESTshrink weight matrices without deleting weights
KV cacheKIVI 2-bit (2402.02750), KVQuant (2401.18079), eviction: H2O (2306.14048), StreamingLLM (2309.17453), SnapKV (2404.14468); lossless-verification: VeriCache (2605.17613)grows in importance with context length; eviction is lossy — 2026 answer is a verification layer over compressed cache
long-context prefillMInference (2407.02490), DeepSeek Sparse Attention (indexer + top-K)prefill-side win, orthogonal to everything else; watch the short-sequence crossover (§08: sparse can be slower when S<K)

Rule of thumb from the 2025–26 literature: optimize for the exact deployment format first (PTQ + a fast kernel), and only escalate to QAT/QAD when crossing into sub-4-bit weights, FP4/FP8 activations, diffusion latents, or long-context KV where quality collapse is real.

05Non-training engineering: the systems layer across model families

Everything below changes how a fixed set of weights is executed — no gradient steps. First the LLM-serving backbone (the canonical stack), then the per-family layers: kernels/quant, diffusion, audio, VLM, and edge.

5.1 The LLM serving backbone (the order practitioners try things)

  1. Profile correctly first. Roofline per op. The traps that generate fake numbers: JIT warmup on first call (librosa/numba ≈ 600 ms one-time), launch-overhead-dominated tiny kernels, thermal/power throttling on laptops, and batch-size dependence. Report min-of-runs, interleave A/B, sanity-check outputs — see the SGLang loop-repetition incident in §08.
  2. Kernels & graphs: fused attention — FlashAttention-2 (2307.09284), FlashAttention-3 (2407.08608), Triton, CUTLASS epilogue fusion (§06), CUDA Graphs to kill launch overhead, torch.compile, dequant-fused GEMMs (§5.2).
  3. Memory: PagedAttention (2309.06180, 2–4× throughput), KV virtualization, host/SSD KV offload (Mooncake 2407.00079), multi-tenant LoRA serving (S-LoRA; Baseten: 10k fine-tunes on one GPU).
  4. Scheduling: continuous/in-flight batching (Orca, OSDI'22), chunked prefill (Sarathi 2308.16369 / Sarathi-Serve 2403.02310), prefix caching via radix trees (SGLang 2312.07104), KV-cache-aware routing, prefill/decode disaggregation (DistServe 2401.09670) once the phases want different machines.
  5. Parallelism judgment calls: TP wins latency, EP wins throughput (Baseten on GPT-OSS 120B); DP-attention for long-context batch scaling; Ulysses/Ring sequence parallelism with communication–compute overlap (fal's Ulysses Unbound; DeepSeek-V3's DualPipe).
  6. Training-free decode acceleration: prompt-lookup drafting (2311.03511) — free candidates from repeated spans, absurdly effective for copy-heavy OCR/ASR output; tree attention over multi-candidate drafts; self-speculative layer skipping (Draft & Verify 2309.08168).
  7. Measure with SLOs, not peaks: TTFT (prefill), TPOT/ITL (decode), and goodput — throughput subject to a per-user speed floor. "16× more users at ≥300 tok/s" is the number that ships.

5.2 Kernels where compression lives or dies

kernel / enginewhat it iswhy it matters
Marlin (2408.11743)W4A16 GEMM kernels, near-ideal bandwidth on Ampere→BlackwellGPTQ/AWQ weights are only as fast as their dequant kernel — the kernel is the technique
AQLM (2401.06118) · QuIP# (2402.04396)additive / lattice codebook quantization + incoherence processing at 2–3 bitsthe sub-4-bit frontier needs bespoke GEMV kernels; PTQ here flirts with QAT territory (§02)
FlashInfer (2501.01005)JIT-customizable attention engine: heterogeneous KV layouts, load-balanced schedulingthe modern attention substrate under SGLang/vLLM/TRT-LLM comparisons
SageAttention → v3 (FP4, Blackwell)INT8/FP4 quantized attention, plug-and-playdrops into image/video DiT stacks too — the attention-side twin of FP4 GEMMs
fal MXFP8 quantizer · CUTLASS EVT epiloguesquantize at 6+ TB/s writing Blackwell's block-scaled layout directly; fuse norms/SwiGLU into the GEMM epiloguekills the pack step and the HBM round-trips that would otherwise eat the FP4 win (§06)

5.3 Diffusion & generative media (training-free)

levermethodsmeasured
step caching — adjacent denoising steps are redundantDeepCache (2312.00858) (U-Net high-level features) → TeaCache (2411.19108) (timestep-embedding signal decides when to reuse) → FasterCache (2410.19355) (dynamic reuse + CFG-Cache across the two guidance branches)2.3× SD1.5 · up to 4.4× Open-Sora-Plan at −0.07% VBench · 1.67× Vchitect-2.0 — composes with step distillation since it's train-free
sparse attention for video DiTsVSA (2505.13389) trainable-sparse (FastVideo), Baseten's custom Video Sparse Attentionpart of Wan 2.2's 133 s → 3 s (§08)
multi-GPU for one image/videoDistriFusion (2402.19481) — displaced patch parallelism using one-step-stale features; PipeFusion (2405.14430) — patch-level pipeline parallelism for DiTs, low-bandwidth (both shipped in xDiT)6.1× SDXL latency at 3840² on 8×A100
per-pass plumbingVAE tiling/slicing, torch.compile, FP8/FP4 DiTs, CFG folded into one batched forwardhygiene layers under every media endpoint

5.4 Audio — ASR & TTS

  • The engine, not the model: faster-whisper (CTranslate2 rewrite + int8) gave ~4× over the original PyTorch Whisper at identical weights — the classic proof that re-implementing the runtime is a first-class technique. Whistle/kaminari-style CUDA-graph engines are the same move in modern dress.
  • Streaming architecture is the latency spec: chunked encoders, VAD, first-sound latency budgets for TTS (fal's Chatterbox Turbo: sub-150 ms), streamed acoustic-token decode, KV reuse across turns.
  • Codec frame rate is the cost multiplier for speech-LLMs: 75 Hz residual-VQ tokens vs 12.5 Hz (Mimi-style) token rates is a ~6× decode-bill difference — decided at training time, paid at serving time (§02).
  • Front-end hygiene: decode+resample+mono in one pass (torchcodec vs librosa: 2×) — it's ~2% of wall; fix once, then forget it (dante lesson).

5.5 Vision-language / multimodal

  • Prefill dominates VLM cost (hundreds–thousands of image tokens): shrink or skip them at serve time — FastV (2403.06764) prunes image tokens in deeper layers, no training; the tuned variant is PruMerge (§02 catalog).
  • Encoder-output caching: don't re-run the ViT across turns; cache visual embeddings behind the same prefix-cache machinery as text.
  • Resolution/tiling budgets and MRoPE-friendly windowing are scheduling knobs with direct prefill cost; bounded-KV sliding windows enable long streaming OCR (kaminari territory).

5.6 Edge & heterogeneous compute

  • llama.cpp / GGUF k-quants + graph rewrites; exllama2 GEMV kernels; KTransformers (SOSP'25) CPU/GPU hybrid MoE — experts in system RAM, attention on the GPU; Apple unified-memory stacks (MLX); ternary kernels for BitNet.
  • The edge is where the whole stack must compose: batch-1, bandwidth-starved → quantized weights + paged KV + self-speculation all hit simultaneously, and every unfused op costs more than it does on HBM-rich datacenter parts.

06Case study — fal: sub-second Ideogram V4, 6.3× with no visible quality loss

blog.fal.ai, Jul 2026. 2.75 s → 0.44 s at 1K, framed as cost = (# forward passes) × (cost per pass) with every factor attacked:

cheaper passes: NVFP4 + epilogue fusion

a 4-bit GEMM is worthless if an unfused RMSNorm sends the output back to HBM. Blackwell epilogue details: head_dim=128 fuses in one thread's fragment; head_dim=256 straddles subtiles → they wrote a three-touch "revisit" epilogue keeping the first half in registers. SwiGLU fused by permuting B's columns at weight-pack time so gate/up pairs land adjacent.

kernels decide whether quantization pays

quality repair: QAD through the quantizer

naive FP4 washed out color — and DiTs are less forgiving than LLMs: velocity error compounds across denoising steps. Fix: distill the bf16 teacher into the FP4 student with velocity-MSE, gradients flowing through the rounding via a straight-through estimator. Failed attempts stated plainly: post-hoc color correction can't un-bake latent errors; naive frozen-quant distillation lowered loss but not quality — "diffusion training loss does not predict image quality."

weights move to where rounding doesn't hurt

fewer branches: CFG distillation

teacher runs both branches; student is only the conditional branch trained to predict the already-guided velocity. Inference: one forward at guidance=1. Halves the transformer cost; beyond the published QAD paper, which stops at recovery.

the CFG branch deleted from the serving path

fewer passes: timestep distillation

DMD2-family, but their own GAN-free data-free distribution-matching objective first (closest public cousins TDM / CDM), then a GAN stage as the primary loss with zero discriminator stabilizers — stable because the starting student was already strong. Order matters: few-step bf16 student first, QAD into FP4 on top.

few-step · single-branch · FP4 · fused

The "Fast" tier collapses the two guidance branches; "Instant" additionally collapses step count — 6.3× over the bf16 many-step base. Every technique attacks a different term of the same equation, and they stack.

07Case study — fal: ~1000 tok/s and 16× throughput on the Ideogram prompt expander

blog.fal.ai, Jul 2026. An LLM pipeline end-to-end: expand user prompts (~600 tokens) in <2 s so the prompt expander stops being the bottleneck next to the (already distilled) image model.

  • Size ≠ serving speed: at low concurrency, 0.8B through 27B dense models were almost equally slow — memory bandwidth. A 35B-A3B MoE in FP8 was the sweet spot: prompt expansion needs world knowledge; the 4B variant produced worse images and broken JSON despite few-shots.
  • Distillation pragmatics: distilled the 397B teacher's style/format; online KD, offline KD and PEFT all converged to the same model → they shipped PEFT. Froze MoE expert weights during fine-tuning: training them made the model "ambitious beyond its capability," breaking image coherence.
  • The speculation ladder, with measured rungs: 328 tok/s baseline (SGLang+FlashInfer) → 500 with the stock MTP head (not fine-tuned, acceptance ~2.5) → 468 with off-the-shelf DFlash (wrong distribution) → warm-start retrain of DFlash on 250K task samples: cold start failed (data not diverse enough; loading pretrained draft weights and continuing training fixed it) → 700 tok/s, 3.9 acceptance → DSpark (DeepSeek's DFlash + Markov refinement heads): reimplemented the trainer in TorchSpec to dodge its ~38 TB-of-hidden-states-on-disk requirement, warm-started from their own DFlash checkpoint, patched vLLM and SGLang to serve it at all → 830 tok/s at 4.6 acceptance in production.
  • Plumbing > parallelism: TP=4 broke the sharded Markov heads' collectives — replicating the draft-head weights instead pushed past 1000 tok/s — but production stayed TP=1 because TP=4 wasted GPU cost. Result: 2.6× at single-user, 16× total throughput at a fixed 300 tok/s/user floor on one B200.

08Case studies — Baseten: GLM-5, DFlash, video, honest benchmarks

  • Fastest GLM-5 API (SOTA on Artificial Analysis: 186+ tok/s, lowest TTFT). Custom DeepSeek-Sparse-Attention kernels: the lightning indexer must scan the whole sequence, so at short context sparse attention is slower than dense — production fix: skip the indexer below K, fuse the indexer's high-precision projections (they can't run FP8) through the overlap path. Zero-overhead MTP: draft and target run as one fused forward — "an often-overlooked aspect of speculative decoding is minimizing the overhead." Plus NVFP4 on Blackwell, hand-profiled MoE dispatch kernels, KV-aware routing.
  • DFlash implementation (~3× latency and throughput vs 2× EAGLE, Qwen3-8B on one B200; GSM8K 654 TPS mean). Best public explanation of block-diffusion drafting + the training recipe (§03).
  • "Fastest video inference, ever" (Wan 2.2): 5-second video, 133 s → 3 s on one B200 — custom Video Sparse Attention + AdaLN/RMSNorm/gated-residual fusion + timestep distillation + FP4. Convergent with fal's Ideogram stack, which is itself the finding.
  • GPT-OSS 120B at 500+ tok/s: the TP-vs-EP judgment call — TP for latency SLOs, EP for raw throughput; shipped TP4EP1 on TensorRT-LLM's Blackwell MoE backend.
  • Benchmark hygiene: their DFlash eval excluded SGLang because it produced loop-repetition outputs that inflated acceptance stats and biased length distributions — filtered results "were not directly comparable." The stack's numbers lie unless you eyeball the generations.
  • Research track: "Still" (single-pass amortized KV-cache compaction), RadixMLP (intra-batch deduplication — the same tokens computed once per batch step, Aug 2026), neural KV-cache compaction for long-running agents, "BYO SWE-grep" (RL-trained fast code-search sub-agents), production speculative-decoding on TRT-LLM, and 10,000 fine-tuned LoRA models served from a single GPU.

09Case study — vLLM: the open engine room, same recipes

blog.vllm.ai, Dec 2025–Aug 2026. The closed-shop stacks of §06–§08 have a public mirror — same techniques, reproducible numbers, and the best place to watch the field institutionalize:

10Where it converges: the 2026 standard playbook

1 · one stack, many terms

both shops independently shipped: FP4/FP8 hardware-native quant + fused epilogues so the quant win isn't eaten by HBM round-trips + a trained-for-your-model drafter + SLO-gated throughput.

2 · "post-training for serving cost" is a real job

QAD, CFG/step distillation, draft-head retraining — small GPU-hour fine-tunes whose entire purpose is the deployment bill.

3 · order of composition matters

distill steps first, quantize the few-step student second (QAD twice). The naive order fights compounding error.

4 · zero-overhead plumbing is a technique

fused draft+verify forward; replicated vs sharded draft heads. Each was worth >1.5× by itself.

5 · warm-start, don't cold-start

drafters and students inherit pretrained weights; task data fine-tunes them. Cold training on 250K samples failed; continuing from the public checkpoint didn't.

6 · quality gates, always

throughput claims only count next to a quality column (MMLU/WER/NED) and eyeballed outputs — and min-of-runs on throttling-prone hardware.

11Read-through for kaminari & dante

moves with the best published evidence

  • DFlash/DSpark-style drafter trained on our own OCR outputs — copy-heavy text → very high acceptance; prompt-lookup as a free baseline before training anything
  • QAD, not just int4 PTQ, if we push sub-4-bit or worry about OCR quality — and the literature predicts W4A16 lands near-lossless (our pending int4-vs-bf16 NED experiment should check exactly that)
  • GQA down-conversion + short fine-tune on the Qwen3-VL thinker → proportional KV cut
  • KV int4/int8 for the streaming scheduler with a verification escape hatch (VeriCache)
  • Zero-overhead MTP plumbing: fuse draft+verify into one graph — the qwen3-mtp branch should measure host overhead explicitly, per Baseten

what does not transfer

  • CFG/timestep distillation is diffusion-only leverage (relevant if verso's flow-matching TTS ever needs a turbo tier — then DMD2-family is the road)
  • Long-context sparse attention (DSA/MInference) needs contexts long enough to clear the indexer overhead — irrelevant at OCR-window sizes
  • Disaggregated prefill/decode and multi-node: single-GPU by design, out of scope
  • Concise-reasoning post-training buys little for OCR (output is already minimally verbose)

12The low-resource chapter: this playbook on a 6 GB laptop GPU

The case studies measured on B200/B300 fleets — but almost every §02/§04 lever is a fine-tune-scale job, and fine-tune scale fits in 6 GB with room left for discipline:

jobbudget at 6 GBthe trick
LoRA student KD (≤2B)~3.5–4 GBteacher inference-only (1.7B bf16 ≈ 3.4 GB) — or offline target generation: two passes, never co-resident
full fine-tune of a 0.6B decoder~5 GB, tight8-bit Adam (states 1.2 GB not 4.8) + grad checkpointing + micro-batch ≤2, seq ≤512
depth-prune + LoRA recovery (Minitron-lite)fitsdelete middle layers, recover 1–2 epochs against the pre-prune self (§02.1)
QAT / QAD-lite STE finetune (0.6–1.7B, W4 fake-quant)fits with LoRAforward carries the quant noise, backward bypasses it — §06's recipe at laptop scale
GPTQ / 2:4 SparseGPT one-shotsfits (solves layer-by-layer, CPU-friendly)calibration = a few hundred samples of your own data
drafter training (EAGLE / DFlash block head)fits at 0.6–2B targetsoffline hidden states; warm-start from public draft weights (§07's lesson)
serving target + drafter together~4.5 GB (2B bf16 + 4-bit drafter)vLLM/SGLang spec-decode paths; acceptance length is the experiment

12.1 The experiments worth running first (open cells, not saturated ones)

  • Vocab-diet KD on a 152k-vocab ASR decoder: the lm_head alone is ~155M params — a quarter of a 0.6B model. Distill into a 32k-row student on the deployment's text distribution; the output GEMM gets ~5× cheaper and the WER-vs-TPOT Pareto basically draws itself. distil-whisper proved vocab-cutting in principle (§02.1); nobody has published it for the Qwen3-ASR family.
  • Task-domain drafters for copy-heavy decoders (ASR/OCR): drafters are trained for chat; transcription/OCR output is template-y — acceptance ≥5–6 is plausible where chat gets 2–4 (§03), and prompt-lookup gives a free baseline before training anything.
  • Consumer 2:4 sparsity: the sparse-tensor-core literature benchmarks A100+; one-shot SparseGPT + cuSPARSELt on a 3050-class part is near-unexplored — and the sparsity×W4 interaction (do they stack?) is an open cell.
  • QAD int4 recovery gated by a task metric (WER/NED, not perplexity): fal's §06 recipe transplanted from velocity-MSE to token logits is a genuinely open study.
  • Self-gating loops: TTS→ASR round-trip WER as a human-free quality gate — the synthesizer generates the data, the transcriber scores it. Where two of your models validate each other, you can run sweeps with nobody in the loop (Baseten's "Still" KV compaction uses exactly this trick: ASR re-transcribes the compacted cache).

12.2 Protocol — cheap hardware punishes sloppy measurement twice

  • Power/thermal caps are the enemy: a 35 W laptop part throttles within ~2 consecutive benchmark runs — fresh process, few runs, interleaved A/B, report min, checkpoint like preemption is imminent (it is).
  • Disk is a first-class constraint (10–40 GB free is normal): audit the HF cache before downloading — and beware config-only stubs: a "cached" model can be 84 KB of metadata with no weights.
  • Data without downloads: synthesize pairs locally (cached TTS generates audio, cached ASR + forced alignment generates transcripts).

12.3 When to rent: market rates, Aug 2026 (USD and NGN per GPU-hour, ₦1,338.68/USD)

GPUVast.ai marketplaceRunPod communityRunPod secureDataCrunch
RTX 3090 24 GB$0.05–0.20 → ₦67–268———
RTX 4090 24 GB$0.06–0.35 → ₦80–469$0.34 → ₦455$0.69–0.74 → ₦924–991—
RTX 5090 32 GB$0.33–1.00 → ₦442–1,339varies——
A100 80 GB$0.65–0.90 → ₦870–1,205$1.19–1.39 → ₦1,593–1,861$1.39–1.49 → ₦1,861–1,995—
H100 80 GB$1.20–1.85 → ₦1,606–2,477$1.99–2.69 → ₦2,664–3,601$2.89–2.99 → ₦3,869–4,003~$2.29 → ₦3,066

The spread is the story: same-class H100 listings span $1.49–6.98/hr across providers, and marketplace prices move hourly (budget ₦1,400–1,500/USD for card fees and parallel-market reality). Working rule: iterate locally, rent only to generalize a finding — one overnight 4090 (~₦2k–10k) finishes a QAD or depth-prune sweep; the H100 tier is for headline benchmarks after the laptop says the result is real. Vast interruptible + checkpointing is just §12.2's thermal protocol with extra steps.

13The audio lane: the fal recipe on Stable Audio 3, and mini-DSpark for speech tokens

13.1 Running the §06 stack on stable-audio-3-small — what's left to attack

Stability already ran much of the playbook (§06's own lineage): the Small-Music/SFX post-trained checkpoints ship as a 433M-param flow-matching DiT over 256-dim SAME latents (4096× compression, 44.1 kHz stereo, 266M-param SAME-S autoencoder), pre-distilled and adversarially post-trained to 8 PingPong steps at CFG 1.0 — guidance is already folded and steps already collapsed from the base checkpoints' ~50 steps at CFG≈7. So the cost equation reduces to two live terms and one honest alternative:

  • Cheap passes — quantize + QAD: a 433M DiT + 266M autoencoder quantize easily; int8 W8A8 with a short STE quantization-aware distillation (§06's recipe — the velocity field will not forgive quant noise any less than images) fits a 6 GB laptop. FP4 is Blackwell-only; on Ampere hardware int8 is §06's analogue.
  • Fewer passes — 8 → 4 steps: TDM/CDM-style steps-aware distillation on top of an already-distilled model, quality-gated by FAD / embedding distance. Stability's own docs concede below-8 trades quality; the live question is whether a second distillation pass, warm-started into the student's own few-step distribution (§09's warm-start lesson applied to samplers), closes that gap.
  • Reproduce the full stack from small-music-base: the pedagogically honest run — CFG-folding + DMD-family step-distillation + QAD, all on a 433M DiT. That's a fine-tune-scale job: ~₦33k–80k on rented 4090s for 3–7 days; pipeline smoke tests run on laptop VRAM.

13.2 Which small models can carry a mini-DSpark — and yes, someone has done it for speech

DSpark = DFlash's block-diffusion drafter + Markov refinement heads (§03). Its interface is architecture-agnostic: fused hidden states from the target, frozen embed + lm_head, standard speculative verification. What matters is whether the target's output token distribution is draftable. The local zoo:

modeltoken streammini-DSpark verdict
Qwen3-ASR-0.6B/1.7Btranscript textstrong fit. SpecASR 3.04–3.79× on LLM-ASR (2507.18181); CTC self-speculative drafts 4.4× iRTF (2603.11243); SpeechSpec ran DFlash/EAGLE3 on Qwen3-Omni / Qwen2-Audio at 1.5–1.65× batch-1, WER unchanged
Qwen3-TTS-0.6B/1.7B12 Hz discrete audio tokensprecedent exists. SpeechSpec reports DSpark/DFlash on CosyVoice2/3's speech-token LLMs at 1.2–3.1× across batch. Draft only the semantic/first codebook — acoustic RVQ layers are high-entropy and naïve acceptance degrades quality (said plainly by the codec-MTP paper, 2410.13839; VADUSA 2410.21951 for the AR-TTS baseline)
GLM-OCR (1.33B, ships 1× MTP)copy-heavy textbest bet. Document layout is near-deterministic → acceptance ≥5 is plausible, and nobody has published it (§11's plan)
LFM2-Audio-1.5Btext (audio understanding)suitable, small ceiling. The conv-heavy LFM2 hybrid breaks nothing — the drafter only consumes hidden states — but conv-dominated decode is already cheap per token, shrinking the bandwidth-bound problem spec-dec exists to solve. Expect ~1.2–1.6× at batch 1, and only if draft+verify is fused so host overhead doesn't eat it (§08's zero-overhead MTP lesson)
gemma-4-E2B-it-qat-mobiletextgood lab target: quantized student + drafter = §01's compositionality (QuantSpec-style) on one card
whisper-smalltranscript textproven but modest: 2.2× with distil-whisper assisted decoding; encoder-decoder drafter support is thin
stable-audio-3 · mimi · htdemucscontinuous latents / non-ARincompatible by construction — there is no token stream to verify; their lever is §06's (steps × branches × passes), not §03's

Cost of a mini-DSpark: the drafter is ~one block + Markov heads reading the target's fused hidden states; 50–250K domain samples (fal needed 250K and warm-started, §07), and hidden-state dumps at 0.6B scale land in the tens of GB — nowhere near fal's 38 TB problem. Training is hours on a 4090 (₦3k–10k) or overnight on a laptop-class card, with the acceptance measurement + batch-scaling curve on the rented GPU. And there's the open cell, free for the taking: SpeechSpec published batch-1 numbers only — acceptance × batch × bit-width on ASR/OCR token streams (§01's regime shift meeting §12's laptop constraint) is uncharted.

14Source trail

clustersources
fal: Ideogram V4 FP4+QAD+CFG+timestep stack; prompt expander / DSparkServing sub-second Ideogram v4 · 1000 tok/s with DSpark · Epilogue fusion · MXFP8 quantizer on Blackwell
Baseten: GLM-5, DFlash, video, GPT-OSS, benchmarkingFastest GLM-5 API · DFlash · GPT-OSS 120B · How to benchmark · Research: "Still" KV compaction
speculative lineage2211.17192 · 2302.01318 · 2401.10774 (Medusa) · 2402.05109 (Hydra) · 2401.15077 (EAGLE) / 2406.16858 / 2503.01840 (EAGLE-2/3) · 2602.06036 (DFlash) · 2412.19464 (DeepSeek-V3 MTP) · 2602.06019 (post-hoc MTP self-distillation) · 2309.08168 (Draft & Verify) · 2404.18911 (Kangaroo) · 2502.10424 (QuantSpec) · 2311.03511 (prompt lookup) · 2507.20127 (speculative cascades)
quantization: PTQ → QAT → QAD2210.17323 (GPTQ) · 2306.00978 (AWQ) · 2211.10438 (SmoothQuant) · 2407.11062 (EfficientQAT) · 2506.09104 (UPQ) · 2601.20088 (NVFP4 QAD) · GuidedQuant (ICML 2025)
KV cache & long context2309.06180 (PagedAttention) · 2412.19442 (KV survey) · 2402.02750 (KIVI) · 2401.18079 (KVQuant) · 2306.14048 / 2309.17453 / 2404.14468 (H2O / StreamingLLM / SnapKV) · 2407.02490 (MInference) · 2605.17613 (VeriCache)
distillation & compression1503.02531 · 1603.09785 · 2307.13386 (GKD) · 2305.08445 (MiniLLM) · 2310.06694 (Sheared-LLaMA) · 2407.14679 (Minitron) · 2311.00430 (distil-whisper) · 2305.13245 (GQA) · 2405.04434 (MLA) · 2205.14135 (CALM) · 2404.16710 (LayerSkip) · 2502.18600 (Chain of Draft) · 2412.19792 (InfAlign)
diffusion step distillation2303.01469 (CM) · 2310.04378 (LCM) · 2202.00512 (progressive) · 2311.18828 / 2405.14867 (DMD/DMD2) · 2311.17042 (ADD) · 2503.06674 (TDM) · 2502.06085 (MeanFlow) · 2207.12598 (guidance distillation)
diffusion caching & parallel inference2312.00858 (DeepCache) · 2411.19108 (TeaCache) · 2410.19355 (FasterCache) · 2505.13389 (VSA) · 2402.19481 (DistriFusion) · 2405.14430 (PipeFusion) · xDiT
serving systems & kernels2309.06180 (vLLM) · 2308.16369 / 2403.02310 (Sarathi) · 2401.09670 (DistServe) · 2407.00079 (Mooncake) · 2312.07104 (SGLang) · 2407.08608 (FA3) · 2408.11743 (Marlin) · 2401.06118 (AQLM) · 2402.04396 (QuIP#) · 2501.01005 (FlashInfer) · 2410.02367 / 2505.11594 (SageAttention) · KTransformers SOSP'25 · 2311.03285 (S-LoRA)
multimodal / VLM efficiency2403.06764 (FastV) · 2403.15388 (LLaVA-PruMerge) · fal DSpark · Baseten GLM-5
vLLM engineering blog: spec-dec toolchain, KV platform layer, Blackwell stacks, CI disciplineInside vLLM: Anatomy · DSpark adaptive verification · 25K TPS/GPU (Qwen3.5) · GLM-5.2 on 24×B300 · FP8 KV state · TurboQuant study · Decode Context Parallelism · KV offloading connector · Hybrid SSM disaggregation · DiffusionGemma · CI & benchmarking discipline
Baseten research trackRadixMLP intra-batch dedup · "Still" KV compaction · Repeated KV for agents · BYO SWE-grep
GPU rental rates (Aug 2026) + FXvast.ai marketplace snapshots (gpu.org, gpuhosted review) · RunPod pricing · DataCrunch H100 rate · FX: open.er-api.com USD→NGN 1,338.68 (2026-08-29)
audio lane: speculative decoding for speech/ASR/TTS; Stable Audio 3 architecture2410.21951 (VADUSA) · 2410.13839 (codec MTP + spec-dec) · 2507.18181 (SpecASR) · 2603.11243 (CTC self-speculative) · SpeechSpec (DFlash/DSpark on Qwen3-Omni & CosyVoice) · Whisper assisted decoding · stable-audio-3 docs (433M DiT · 8-step PingPong · CFG 1.0)