The inference-efficiency playbook — post-training and serving for faster models
A working map of what an inference/ML-efficiency engineer actually does: the cost model, the six families of post-training that buy serving speed, the training-free compression toolbox, the systems layer — and the concrete 2026 production numbers from fal's research blog and Baseten's model-performance posts that prove the recipe works end to end. Includes the full catalog of post-training-for-efficiency methods (§02) the non-training engineering map across LLMs, diffusion, audio and VLMs (§05), the open engine room's production numbers from vLLM's engineering blog (§09), and how to run the whole post-training-for-efficiency loop on a 6 GB laptop — including Aug-2026 cloud-rental rates for the runs that need more (§12) — plus a worked application to the audio lane: the fal recipe on Stable Audio 3, and mini-DSpark drafters for speech-token models (§13).
01The mental model everything hangs on
Every technique attacks one of four multiplicative terms:
latency/cost ≈ (tokens generated) × (bytes moved per token) × (time per byte) ÷ (parallelism amortized over batch)
- Decode is memory-bound. One token at a time; every step re-reads all weights plus the whole KV cache. Arithmetic intensity is ~O(1) per matmul. Anything that shrinks bytes read per token (quantization, sparsity, KV compression, fewer KV heads) is nearly free speedup at small batch.
- Prefill is compute-bound. O(n²) attention, full matmuls. Attacked with chunked prefill, sparse attention, prefix caching — a different toolbox.
- Batch changes the regime. At large batch, decode approaches compute-bound: quantization's win shrinks, KV capacity and scheduling dominate. A trick that wins at batch 1 can lose at 64.
- Token count is a hidden dimension. Generating 2000 reasoning tokens to say "yes" is 2000× the decode cost. Post-training can shrink this directly — most people forget it.
02Post-training routes that buy inference speed
Changing the weights by training so the deployed artifact is cheaper. Six families:
2.1 Distillation into a smaller student
- Classic soft-target KD (Hinton, 1503.02531) → sequence-level KD (Kim & Rush, 1603.09785) → on-policy distillation that fixes the train/inference mismatch (GKD, 2307.13386; DistiLLM, 2401.10491) and reverse-KD so students stop covering modes they can't (MiniLLM, 2305.08445).
- Prune-then-adapt: Sheared-LLaMA (2310.06694) — structured 3D pruning + continued training beats from-scratch small models at 25% of the compute; Minitron (2407.14679) — prune width/depth and MoE experts, then a short KD re-alignment (~3% of original training cost).
- The canonical inference-aware ASR example: distil-whisper (2311.00430) — decoder-only distillation of Whisper (6→2 decoder layers, vocab cut, synthetic alignment retraining): 4.9× faster at <1% WER loss, plus a cascaded text-only correction model.
- Fal's pragmatics (see §07): online KD, offline KD and plain PEFT landed at the same quality on their task — so ship PEFT, it's the cheapest.
2.2 Training the model to emit fewer tokens
- Concise-reasoning / overthinking control: token-budget penalties in RL (2403.05587), CCoT curriculum (2412.20993), L1 length-controlled DPO (2503.04697), Chain of Draft (2502.18600) — think in ≤5-word drafts, ~90% token cut at near-parity accuracy.
- Vocabulary work: MiniLLM-style vocab distillation shrinks the lm_head (a 151k-row embedding is real decode cost); tokenizer retraining cuts token counts 20–30% for CJK/multilingual.
- Adaptive computation: CALM early-exit adapters (2205.14135) and Meta's LayerSkip (2404.16710) — train draft-at-bottom/verify-at-top inside one model; self-speculation by construction.
2.3 Training the decoding algorithm itself (MTP & draft heads)
The most active 2024–2026 front — detailed in §03. Post-hoc head-training jobs cost GPU-hours, not runs: Medusa (2401.10774), EAGLE-3 (2503.01840), DFlash (2602.06036), DSpark; DeepSeek-V3 (2412.19464) proved MTP heads trained in the main run serve as a free draft module; post-hoc MTP self-distillation (2602.06019) retrofits the trick onto existing checkpoints (>3× decode, <5% accuracy loss).
2.4 Quantization-aware post-training
You cannot PTQ your way below 4 bits cleanly — see §04 for the ladder: learned rounding (LSQ 1810.01875, GMLQ 2102.03586), EfficientQAT (2407.11062: ~25% of usual QAT cost), progressive UPQ FP16→INT4→INT2 (2506.09104), end-loss-guided PTQ GuidedQuant (ICML 2025), and Quantization-Aware Distillation (2601.20088): KL from the bf16 teacher into an NVFP4 student — the current default at the frontier when you can't rerun the original SFT/RL recipe.
2.5 Architecture surgery + fine-tune recovery
- MHA→GQA uptraining (2305.13245): shrink KV heads with a small fine-tune; KV cache shrinks proportionally.
- Sliding-window + global interleaves (Gemma recipe) → bounded KV for streaming.
- MLA low-rank KV latents (DeepSeek-V2, 2405.04434): ~93% KV reduction — wants training in the loop.
- Layer/block pruning with cheap recovery: LLM-Pruner (2305.15286), ShortGPT (2310.11102 — delete 25% of layers in minutes), SliceGPT (2401.15024 — orthogonalize then slice, no retrain).
- MoE compression: expert pruning/merging (I-MoE 2405.04534, Minitron expert slimming).
2.6 Inference-aware post-training: optimize the procedure, not the likelihood
- InfAlign (2412.19792): align against the actual test-time policy (best-of-N), not single-sample quality.
- Diffusion step-distillation is the same philosophy with a sampler instead of a decoder: consistency models (2303.01469), progressive distillation (2202.00512), DMD/DMD2 (2311.18828 / 2405.14867), adversarial ADD → SDXL-Turbo (2311.17042), MeanFlow (2502.06085). Deployed artifact runs 4 steps, not 50 — and the objective was chosen because of the sampler. See §06 for the 2026 production version.
- Model merging (DARE-TIES, soups) as cheap capability recovery after an efficiency fine-tune ate generality.
2.7 The full catalog — every post-training method that buys inference speed
Grouped by which term of the §01 cost equation it attacks. A method is listed only if it needs gradient steps on real data — pure training-free compression lives in the §04 toolbox. Every entry links to its representative paper.
| family | method | what inference gains | key refs |
|---|---|---|---|
| student size | soft-logit KD → student | smaller model ≈ proportional decode + KV cut | 1503.02531 |
| sequence-level KD (synthetic teacher data) | same, without logit access | 1603.09785 | |
| on-policy KD (student samples, teacher scores) | kills exposure bias on long generations | GKD 2307.13386 · DistiLLM 2401.10491 | |
| reverse-KD generative distillation | smaller student, no mode-averaging mush | MiniLLM 2305.08445 | |
| speculative KD (teacher as the drafter and target) | student + built-in acceptance lift in one job | 2404.09728 | |
| task distillation of a whole pipeline (ASR) | distil-whisper: 4.9× at <1% WER loss | 2311.00430 | |
| prune-then-adapt (width/depth/head/experts + KD recovery) | cheap small model from a big one | Sheared-LLaMA 2310.06694 · Minitron 2407.14679 | |
| layer/block deletion + short retrain | depth cut w/ recovery; SliceGPT needs no retrain | ShortGPT 2310.11102 · LLM-Pruner 2305.15286 · SliceGPT 2401.15024 | |
| MoE expert pruning / merging / shrinking | smaller active set + denser expert tables | I-MoE 2405.04534 · Minitron | |
| tokens generated | length-controlled RL / DPO (budget or penalty in the reward) | linear latency cut on reasoning-heavy traffic | 2403.05587 · L1 2503.04697 · CCoT 2412.20993 |
| concise-thought formats (Chain-of-Draft) | ~90% fewer CoT tokens at parity | 2502.18600 | |
| latent / continuous reasoning | thinking in fewer, denser steps | Coconut 2412.06769 | |
| vocabulary / tokenizer shrink for the deployment language mix | fewer decode steps per unit of meaning | MiniLLM (vocab) · tokenizer retrain | |
| early-exit training (per-layer confidence) | easy tokens skip deep layers | CALM 2205.14135 · LayerSkip 2404.16710 | |
| draft / verify structure | parallel decoding heads (frozen backbone) | self-drafting, ~1–2× | Medusa 2401.10774 · Hydra 2402.05109 |
| feature-space autoregressive draft head | higher acceptance; ~2× ceiling | EAGLE 2401.15077 · -2 · -3 2503.01840 | |
| MTP module trained with the main model | draft built-in; free at serving | DeepSeek-V3 2412.19464 | |
| post-hoc MTP via online self-distillation | retrofit existing ckpts, >3× decode | 2602.06019 | |
| block-diffusion parallel drafter (+ Markov refine heads) | ~3×; DFlash 4.6 acceptance | DFlash 2602.06036 · DSpark (DeepSeek) | |
| double early-exit adapter (self-speculative) | no draft model, lossless | Kangaroo 2404.18911 · Draft&Verify 2309.08168 | |
| warm-started drafter retrain on task outputs | acceptance recovery for fine-tuned targets | fal DSpark post | |
| number format | learned rounding / quant (QAT) | int8→int4 weights+acts deployable-native | LSQ 1810.01875 · GMLQ |
| budgeted / progressive QAT recipes | QAT at ~25% cost; INT2 instruction models | EfficientQAT 2407.11062 · UPQ 2506.09104 | |
| quantization-aware distillation (teacher → quantized student) | recover NVFP4/FP4 quality post-SFT/RL | QAD 2601.20088 · GuidedQuant ICML'25 | |
| rotation + quant fine-tuning (incoherence processing) | makes 2–4 bit tractable | QuaRot 2404.00456 · SpinQuant | |
| ternary / 1.58-bit training | extreme bandwidth cut (BW GEMM) | BitNet b1.58 2402.17764 | |
| sparsity-aware fine-tune (2:4 semi-structured) | real 1.3–1.7× on Ampere+ silicon | Maini et al. FGS; NVIDIA blog | |
| attention / KV structure | MHA→GQA uptraining | KV heads ÷N w/ small fine-tune | GQA 2305.13245 |
| MLA low-rank KV-latent conversion | ~93% KV cut | DeepSeek-V2 2405.04434 | |
| sliding-window + global interleaving fine-tune | bounded KV for streaming | Gemma 2/3 recipe | |
| KV-quantization-aware fine-tuning | tolerates 2–4 bit cache at long context | KVQuant-style + in-loop fake-quant | |
| sampler / procedure | CFG / guidance distillation into one branch | halves DiT passes per step | 2207.12598 · fal Ideogram §06 |
| progressive step distillation | steps ÷2 per stage | 2202.00512 | |
| consistency models / LCM / TDM / MeanFlow | 1–8 step samplers | 2303.01469 · LCM 2310.04378 · TDM 2503.06674 · MeanFlow 2502.06085 | |
| distribution matching (DMD/DMD2; GAN-first variants) | few-step, data-free; the shipped recipe | DMD 2311.18828 · DMD2 2405.14867 | |
| adversarial distillation (ADD) | 1-step image gen (SDXL-Turbo) | 2311.17042 | |
| inference-aware alignment (train vs best-of-N policy) | quality at the deployed procedure | InfAlign 2412.19792 | |
| multimodal / recovery / economics | VLM visual-token compression w/ tuning | ÷14 image prefill tokens | LLaVA-PruMerge 2403.15388 |
| single-step distilled TTS / ultra-small speech LMs | sub-150 ms first-sound | ConsistencyTTS · Chatterbox Turbo | |
| model merging for capability recovery (soup/TIES/DARE) | restores generality lost to efficiency tunes — no training | 2204.06504 · TIES 2306.01708 · DARE 2311.03099 | |
| LoRA / adapter economics (fine-tune per tenant, share base) | marginal model cost ≈ 0 per customer | LoRA 2106.09685 · S-LoRA 2311.03285 |
03Speculative decoding: the 2026 frontier
Exact greedy/sampling acceptance (Leviathan 2211.17192, Chen 2302.01318) preserves the target distribution; the engineering race is all about acceptance length and drafting overhead:
(original specdec)→ parallel heads on frozen backbone
(Medusa, 2401.10774)→ autoregressive draft in feature space
(EAGLE-3, 2503.01840 — ~2× ceiling)→ MTP heads trained with the model
(DeepSeek-V3; Qwen3.6 ships them)→ block-diffusion parallel draft
(DFlash, 2602.06036 — ~3×)→ DFlash + Markov refine heads
(DSpark, DeepSeek — 4.6 acceptance)
- DFlash's argument: EAGLE is autoregressive — drafting 8 tokens costs 8 draft passes and errors accumulate, capping ~2×. DFlash instead uses a deeper draft model with bidirectional attention that unmasks a whole block of γ∈[8,16] tokens in one pass. The single pass is 2–4× slower than EAGLE's — but it replaces EAGLE's entire draft phase and drafts more, so it wins (~3× on Qwen3-8B / B200).
- Training detail worth stealing (Baseten): condition on fused hidden states from 5–6 evenly
spaced target layers; freeze embed + lm_head; random anchors; CE loss with exponential position decay
exp(−(i−1)/γ)— earlier draft tokens matter more. - Fine-tuned models need retrained drafters. Fal (§07): an off-the-shelf MTP/DFlash head trained on the base model under-performs on the fine-tuned target; warm-started retraining on task data fixed acceptance. And drafters for copy-heavy workloads (OCR, ASR) get another free option: prompt-lookup decoding (2311.03511) — zero-cost drafts from repeated spans.
- Self-speculative, no extra params: Draft & Verify layer-skipping (2309.08168), Kangaroo's double early exit (2404.18911), speculative streaming (reuse the target’s own recent hidden states as drafts), QuantSpec (2502.10424). Per-request speculation-length RL: RESDO. Hybrid draft-large/verify-small: Speculative Cascades (2507.20127).
- The catch that papers bury: speculation degrades at large batch — draft FLOPs eat what bandwidth savings bought, and host-side overhead kills the win. Baseten's fix: run draft + verify as one fused forward so overhead is ~zero (§08).
04The training-free toolbox (PTQ) and where it breaks
| Attack | Methods | Practical note |
|---|---|---|
| weights | GPTQ (2210.17323), AWQ (2306.00978), bnb NF4, Marlin fused-dequant GEMMs | W4A16 ≈ lossless on LLMs; if the dequant kernel is slow the win evaporates — the kernel is the technique |
| weights + activations | SmoothQuant W8A8 (2211.10438), OmniQuant, FP8 / NVFP4 / MXFP8 on Blackwell | FP8 basically free on Ada/Hopper+; ≤4-bit activations need QAT/QAD (§2.4) |
| sparsity | Wanda (2306.12929), SparseGPT 2:4 semi-structured, SliceGPT | 2:4 gives real 1.3–1.7× on Ampere+; unstructured remains mostly paper speedups |
| low-rank | SVD-LLM (2403.05170), QuEST | shrink weight matrices without deleting weights |
| KV cache | KIVI 2-bit (2402.02750), KVQuant (2401.18079), eviction: H2O (2306.14048), StreamingLLM (2309.17453), SnapKV (2404.14468); lossless-verification: VeriCache (2605.17613) | grows in importance with context length; eviction is lossy — 2026 answer is a verification layer over compressed cache |
| long-context prefill | MInference (2407.02490), DeepSeek Sparse Attention (indexer + top-K) | prefill-side win, orthogonal to everything else; watch the short-sequence crossover (§08: sparse can be slower when S<K) |
Rule of thumb from the 2025–26 literature: optimize for the exact deployment format first (PTQ + a fast kernel), and only escalate to QAT/QAD when crossing into sub-4-bit weights, FP4/FP8 activations, diffusion latents, or long-context KV where quality collapse is real.
05Non-training engineering: the systems layer across model families
Everything below changes how a fixed set of weights is executed — no gradient steps. First the LLM-serving backbone (the canonical stack), then the per-family layers: kernels/quant, diffusion, audio, VLM, and edge.
5.1 The LLM serving backbone (the order practitioners try things)
- Profile correctly first. Roofline per op. The traps that generate fake numbers: JIT warmup on first call (librosa/numba ≈ 600 ms one-time), launch-overhead-dominated tiny kernels, thermal/power throttling on laptops, and batch-size dependence. Report min-of-runs, interleave A/B, sanity-check outputs — see the SGLang loop-repetition incident in §08.
- Kernels & graphs: fused attention — FlashAttention-2 (2307.09284), FlashAttention-3 (2407.08608), Triton, CUTLASS epilogue fusion (§06), CUDA Graphs to kill launch overhead, torch.compile, dequant-fused GEMMs (§5.2).
- Memory: PagedAttention (2309.06180, 2–4× throughput), KV virtualization, host/SSD KV offload (Mooncake 2407.00079), multi-tenant LoRA serving (S-LoRA; Baseten: 10k fine-tunes on one GPU).
- Scheduling: continuous/in-flight batching (Orca, OSDI'22), chunked prefill (Sarathi 2308.16369 / Sarathi-Serve 2403.02310), prefix caching via radix trees (SGLang 2312.07104), KV-cache-aware routing, prefill/decode disaggregation (DistServe 2401.09670) once the phases want different machines.
- Parallelism judgment calls: TP wins latency, EP wins throughput (Baseten on GPT-OSS 120B); DP-attention for long-context batch scaling; Ulysses/Ring sequence parallelism with communication–compute overlap (fal's Ulysses Unbound; DeepSeek-V3's DualPipe).
- Training-free decode acceleration: prompt-lookup drafting (2311.03511) — free candidates from repeated spans, absurdly effective for copy-heavy OCR/ASR output; tree attention over multi-candidate drafts; self-speculative layer skipping (Draft & Verify 2309.08168).
- Measure with SLOs, not peaks: TTFT (prefill), TPOT/ITL (decode), and goodput — throughput subject to a per-user speed floor. "16× more users at ≥300 tok/s" is the number that ships.
5.2 Kernels where compression lives or dies
| kernel / engine | what it is | why it matters |
|---|---|---|
| Marlin (2408.11743) | W4A16 GEMM kernels, near-ideal bandwidth on Ampere→Blackwell | GPTQ/AWQ weights are only as fast as their dequant kernel — the kernel is the technique |
| AQLM (2401.06118) · QuIP# (2402.04396) | additive / lattice codebook quantization + incoherence processing at 2–3 bits | the sub-4-bit frontier needs bespoke GEMV kernels; PTQ here flirts with QAT territory (§02) |
| FlashInfer (2501.01005) | JIT-customizable attention engine: heterogeneous KV layouts, load-balanced scheduling | the modern attention substrate under SGLang/vLLM/TRT-LLM comparisons |
| SageAttention → v3 (FP4, Blackwell) | INT8/FP4 quantized attention, plug-and-play | drops into image/video DiT stacks too — the attention-side twin of FP4 GEMMs |
| fal MXFP8 quantizer · CUTLASS EVT epilogues | quantize at 6+ TB/s writing Blackwell's block-scaled layout directly; fuse norms/SwiGLU into the GEMM epilogue | kills the pack step and the HBM round-trips that would otherwise eat the FP4 win (§06) |
5.3 Diffusion & generative media (training-free)
| lever | methods | measured |
|---|---|---|
| step caching — adjacent denoising steps are redundant | DeepCache (2312.00858) (U-Net high-level features) → TeaCache (2411.19108) (timestep-embedding signal decides when to reuse) → FasterCache (2410.19355) (dynamic reuse + CFG-Cache across the two guidance branches) | 2.3× SD1.5 · up to 4.4× Open-Sora-Plan at −0.07% VBench · 1.67× Vchitect-2.0 — composes with step distillation since it's train-free |
| sparse attention for video DiTs | VSA (2505.13389) trainable-sparse (FastVideo), Baseten's custom Video Sparse Attention | part of Wan 2.2's 133 s → 3 s (§08) |
| multi-GPU for one image/video | DistriFusion (2402.19481) — displaced patch parallelism using one-step-stale features; PipeFusion (2405.14430) — patch-level pipeline parallelism for DiTs, low-bandwidth (both shipped in xDiT) | 6.1× SDXL latency at 3840² on 8×A100 |
| per-pass plumbing | VAE tiling/slicing, torch.compile, FP8/FP4 DiTs, CFG folded into one batched forward | hygiene layers under every media endpoint |
5.4 Audio — ASR & TTS
- The engine, not the model: faster-whisper (CTranslate2 rewrite + int8) gave ~4× over the original PyTorch Whisper at identical weights — the classic proof that re-implementing the runtime is a first-class technique. Whistle/kaminari-style CUDA-graph engines are the same move in modern dress.
- Streaming architecture is the latency spec: chunked encoders, VAD, first-sound latency budgets for TTS (fal's Chatterbox Turbo: sub-150 ms), streamed acoustic-token decode, KV reuse across turns.
- Codec frame rate is the cost multiplier for speech-LLMs: 75 Hz residual-VQ tokens vs 12.5 Hz (Mimi-style) token rates is a ~6× decode-bill difference — decided at training time, paid at serving time (§02).
- Front-end hygiene: decode+resample+mono in one pass (torchcodec vs librosa: 2×) — it's ~2% of wall; fix once, then forget it (dante lesson).
5.5 Vision-language / multimodal
- Prefill dominates VLM cost (hundreds–thousands of image tokens): shrink or skip them at serve time — FastV (2403.06764) prunes image tokens in deeper layers, no training; the tuned variant is PruMerge (§02 catalog).
- Encoder-output caching: don't re-run the ViT across turns; cache visual embeddings behind the same prefix-cache machinery as text.
- Resolution/tiling budgets and MRoPE-friendly windowing are scheduling knobs with direct prefill cost; bounded-KV sliding windows enable long streaming OCR (kaminari territory).
5.6 Edge & heterogeneous compute
- llama.cpp / GGUF k-quants + graph rewrites; exllama2 GEMV kernels; KTransformers (SOSP'25) CPU/GPU hybrid MoE — experts in system RAM, attention on the GPU; Apple unified-memory stacks (MLX); ternary kernels for BitNet.
- The edge is where the whole stack must compose: batch-1, bandwidth-starved → quantized weights + paged KV + self-speculation all hit simultaneously, and every unfused op costs more than it does on HBM-rich datacenter parts.
06Case study — fal: sub-second Ideogram V4, 6.3× with no visible quality loss
blog.fal.ai, Jul 2026. 2.75 s → 0.44 s at 1K, framed as cost = (# forward passes) × (cost per pass) with every factor attacked:
cheaper passes: NVFP4 + epilogue fusion
a 4-bit GEMM is worthless if an unfused RMSNorm sends the output back to HBM. Blackwell epilogue details: head_dim=128 fuses in one thread's fragment; head_dim=256 straddles subtiles → they wrote a three-touch "revisit" epilogue keeping the first half in registers. SwiGLU fused by permuting B's columns at weight-pack time so gate/up pairs land adjacent.
kernels decide whether quantization pays
quality repair: QAD through the quantizer
naive FP4 washed out color — and DiTs are less forgiving than LLMs: velocity error compounds across denoising steps. Fix: distill the bf16 teacher into the FP4 student with velocity-MSE, gradients flowing through the rounding via a straight-through estimator. Failed attempts stated plainly: post-hoc color correction can't un-bake latent errors; naive frozen-quant distillation lowered loss but not quality — "diffusion training loss does not predict image quality."
weights move to where rounding doesn't hurt
fewer branches: CFG distillation
teacher runs both branches; student is only the conditional branch trained to predict the already-guided velocity. Inference: one forward at guidance=1. Halves the transformer cost; beyond the published QAD paper, which stops at recovery.
the CFG branch deleted from the serving path
fewer passes: timestep distillation
DMD2-family, but their own GAN-free data-free distribution-matching objective first (closest public cousins TDM / CDM), then a GAN stage as the primary loss with zero discriminator stabilizers — stable because the starting student was already strong. Order matters: few-step bf16 student first, QAD into FP4 on top.
few-step · single-branch · FP4 · fused
The "Fast" tier collapses the two guidance branches; "Instant" additionally collapses step count — 6.3× over the bf16 many-step base. Every technique attacks a different term of the same equation, and they stack.
07Case study — fal: ~1000 tok/s and 16× throughput on the Ideogram prompt expander
blog.fal.ai, Jul 2026. An LLM pipeline end-to-end: expand user prompts (~600 tokens) in <2 s so the prompt expander stops being the bottleneck next to the (already distilled) image model.
- Size ≠ serving speed: at low concurrency, 0.8B through 27B dense models were almost equally slow — memory bandwidth. A 35B-A3B MoE in FP8 was the sweet spot: prompt expansion needs world knowledge; the 4B variant produced worse images and broken JSON despite few-shots.
- Distillation pragmatics: distilled the 397B teacher's style/format; online KD, offline KD and PEFT all converged to the same model → they shipped PEFT. Froze MoE expert weights during fine-tuning: training them made the model "ambitious beyond its capability," breaking image coherence.
- The speculation ladder, with measured rungs: 328 tok/s baseline (SGLang+FlashInfer) → 500 with the stock MTP head (not fine-tuned, acceptance ~2.5) → 468 with off-the-shelf DFlash (wrong distribution) → warm-start retrain of DFlash on 250K task samples: cold start failed (data not diverse enough; loading pretrained draft weights and continuing training fixed it) → 700 tok/s, 3.9 acceptance → DSpark (DeepSeek's DFlash + Markov refinement heads): reimplemented the trainer in TorchSpec to dodge its ~38 TB-of-hidden-states-on-disk requirement, warm-started from their own DFlash checkpoint, patched vLLM and SGLang to serve it at all → 830 tok/s at 4.6 acceptance in production.
- Plumbing > parallelism: TP=4 broke the sharded Markov heads' collectives — replicating the draft-head weights instead pushed past 1000 tok/s — but production stayed TP=1 because TP=4 wasted GPU cost. Result: 2.6× at single-user, 16× total throughput at a fixed 300 tok/s/user floor on one B200.
08Case studies — Baseten: GLM-5, DFlash, video, honest benchmarks
- Fastest GLM-5 API (SOTA on Artificial Analysis: 186+ tok/s, lowest TTFT). Custom DeepSeek-Sparse-Attention kernels: the lightning indexer must scan the whole sequence, so at short context sparse attention is slower than dense — production fix: skip the indexer below K, fuse the indexer's high-precision projections (they can't run FP8) through the overlap path. Zero-overhead MTP: draft and target run as one fused forward — "an often-overlooked aspect of speculative decoding is minimizing the overhead." Plus NVFP4 on Blackwell, hand-profiled MoE dispatch kernels, KV-aware routing.
- DFlash implementation (~3× latency and throughput vs 2× EAGLE, Qwen3-8B on one B200; GSM8K 654 TPS mean). Best public explanation of block-diffusion drafting + the training recipe (§03).
- "Fastest video inference, ever" (Wan 2.2): 5-second video, 133 s → 3 s on one B200 — custom Video Sparse Attention + AdaLN/RMSNorm/gated-residual fusion + timestep distillation + FP4. Convergent with fal's Ideogram stack, which is itself the finding.
- GPT-OSS 120B at 500+ tok/s: the TP-vs-EP judgment call — TP for latency SLOs, EP for raw throughput; shipped TP4EP1 on TensorRT-LLM's Blackwell MoE backend.
- Benchmark hygiene: their DFlash eval excluded SGLang because it produced loop-repetition outputs that inflated acceptance stats and biased length distributions — filtered results "were not directly comparable." The stack's numbers lie unless you eyeball the generations.
- Research track: "Still" (single-pass amortized KV-cache compaction), RadixMLP (intra-batch deduplication — the same tokens computed once per batch step, Aug 2026), neural KV-cache compaction for long-running agents, "BYO SWE-grep" (RL-trained fast code-search sub-agents), production speculative-decoding on TRT-LLM, and 10,000 fine-tuned LoRA models served from a single GPU.
09Case study — vLLM: the open engine room, same recipes
blog.vllm.ai, Dec 2025–Aug 2026. The closed-shop stacks of §06–§08 have a public mirror — same techniques, reproducible numbers, and the best place to watch the field institutionalize:
- Speculative decoding became a toolchain, not a trick. DSpark adaptive verification sizes each request's draft-check budget from the draft's own confidence — one configuration holds the throughput/latency frontier from batch 1 to 256, the direct answer to §03's "speculation dies at big batch." P-EAGLE drafts multiple tokens in one pass; EAGLE-3.1 adds post-norm hidden-state feedback; Speculators plus a hidden-state-extraction API make drafter retraining an offline job on task data — fal's §07 workflow, but pip-installable. AMD ships the same story: EAGLE-3 trained + served with Quark at 1.79–2.00× on Kimi-K2.5 / MiniMax-M2.5.
- KV cache is the platform layer. FP8 KV + attention is the shipped default (per-layer skip lists, NIAH before/after checks); the first TurboQuant study says 4-bit KV wins at long context but FP8 remains the safe pick; Decode Context Parallelism shards KV along the sequence across GPUs for ~3× long-context agentic throughput vs plain TP; Mooncake / PegaFlow external KV stores, an async KV-offloading connector, KV-aware routing, elastic expert parallelism and even Mamba conv-state transfer for hybrid SSMs (disaggregation for SSM-attention hybrids) make disaggregation plumbing, not research.
- The headline stacks converge on the §10 recipe: Qwen3.5-397B-A17B NVFP4 on GB200 NVL72 disaggregated → 25K total TPS/GPU; GLM-5.2-NVFP4 on 24×B300 → mean TPOT 40→17 ms via P/D disaggregation + MTP + Model Runner V2; DeepSeek at 2.2k tok/s/H200 on Wide-EP. Day-0 support now ships with model-native MTP, sparse attention (DSA, MiniMax sparse, Kimi-K3 KDA-aware prefix caching) and even the first diffusion LM (DiffusionGemma) reusing the speculative-decoding verification machinery.
- Media on the same engine: vLLM-Omni — TeaCache/Cache-DiT diffusion step caching (1.5–2× for free), staged TTS serving with CUDA graphs, and distributed layerwise offload running a 124 GB DiT on 64 GB HBM. The Qwen3-Omni serving postmortem lists its wins as hot-path cleanup and launch-overhead removal — §05.1 by name.
- Discipline, institutionalized: nightly performance + accuracy evals across every accelerator gate releases, and Inside vLLM: Anatomy of a High-Throughput LLM Inference System is the field's best single engineering document.
10Where it converges: the 2026 standard playbook
1 · one stack, many terms
both shops independently shipped: FP4/FP8 hardware-native quant + fused epilogues so the quant win isn't eaten by HBM round-trips + a trained-for-your-model drafter + SLO-gated throughput.
2 · "post-training for serving cost" is a real job
QAD, CFG/step distillation, draft-head retraining — small GPU-hour fine-tunes whose entire purpose is the deployment bill.
3 · order of composition matters
distill steps first, quantize the few-step student second (QAD twice). The naive order fights compounding error.
4 · zero-overhead plumbing is a technique
fused draft+verify forward; replicated vs sharded draft heads. Each was worth >1.5× by itself.
5 · warm-start, don't cold-start
drafters and students inherit pretrained weights; task data fine-tunes them. Cold training on 250K samples failed; continuing from the public checkpoint didn't.
6 · quality gates, always
throughput claims only count next to a quality column (MMLU/WER/NED) and eyeballed outputs — and min-of-runs on throttling-prone hardware.
11Read-through for kaminari & dante
moves with the best published evidence
- DFlash/DSpark-style drafter trained on our own OCR outputs — copy-heavy text → very high acceptance; prompt-lookup as a free baseline before training anything
- QAD, not just int4 PTQ, if we push sub-4-bit or worry about OCR quality — and the literature predicts W4A16 lands near-lossless (our pending int4-vs-bf16 NED experiment should check exactly that)
- GQA down-conversion + short fine-tune on the Qwen3-VL thinker → proportional KV cut
- KV int4/int8 for the streaming scheduler with a verification escape hatch (VeriCache)
- Zero-overhead MTP plumbing: fuse draft+verify into one graph — the qwen3-mtp branch should measure host overhead explicitly, per Baseten
what does not transfer
- CFG/timestep distillation is diffusion-only leverage (relevant if verso's flow-matching TTS ever needs a turbo tier — then DMD2-family is the road)
- Long-context sparse attention (DSA/MInference) needs contexts long enough to clear the indexer overhead — irrelevant at OCR-window sizes
- Disaggregated prefill/decode and multi-node: single-GPU by design, out of scope
- Concise-reasoning post-training buys little for OCR (output is already minimally verbose)
12The low-resource chapter: this playbook on a 6 GB laptop GPU
The case studies measured on B200/B300 fleets — but almost every §02/§04 lever is a fine-tune-scale job, and fine-tune scale fits in 6 GB with room left for discipline:
| job | budget at 6 GB | the trick |
|---|---|---|
| LoRA student KD (≤2B) | ~3.5–4 GB | teacher inference-only (1.7B bf16 ≈ 3.4 GB) — or offline target generation: two passes, never co-resident |
| full fine-tune of a 0.6B decoder | ~5 GB, tight | 8-bit Adam (states 1.2 GB not 4.8) + grad checkpointing + micro-batch ≤2, seq ≤512 |
| depth-prune + LoRA recovery (Minitron-lite) | fits | delete middle layers, recover 1–2 epochs against the pre-prune self (§02.1) |
| QAT / QAD-lite STE finetune (0.6–1.7B, W4 fake-quant) | fits with LoRA | forward carries the quant noise, backward bypasses it — §06's recipe at laptop scale |
| GPTQ / 2:4 SparseGPT one-shots | fits (solves layer-by-layer, CPU-friendly) | calibration = a few hundred samples of your own data |
| drafter training (EAGLE / DFlash block head) | fits at 0.6–2B targets | offline hidden states; warm-start from public draft weights (§07's lesson) |
| serving target + drafter together | ~4.5 GB (2B bf16 + 4-bit drafter) | vLLM/SGLang spec-decode paths; acceptance length is the experiment |
12.1 The experiments worth running first (open cells, not saturated ones)
- Vocab-diet KD on a 152k-vocab ASR decoder: the lm_head alone is ~155M params — a quarter of a 0.6B model. Distill into a 32k-row student on the deployment's text distribution; the output GEMM gets ~5× cheaper and the WER-vs-TPOT Pareto basically draws itself. distil-whisper proved vocab-cutting in principle (§02.1); nobody has published it for the Qwen3-ASR family.
- Task-domain drafters for copy-heavy decoders (ASR/OCR): drafters are trained for chat; transcription/OCR output is template-y — acceptance ≥5–6 is plausible where chat gets 2–4 (§03), and prompt-lookup gives a free baseline before training anything.
- Consumer 2:4 sparsity: the sparse-tensor-core literature benchmarks A100+; one-shot SparseGPT + cuSPARSELt on a 3050-class part is near-unexplored — and the sparsity×W4 interaction (do they stack?) is an open cell.
- QAD int4 recovery gated by a task metric (WER/NED, not perplexity): fal's §06 recipe transplanted from velocity-MSE to token logits is a genuinely open study.
- Self-gating loops: TTS→ASR round-trip WER as a human-free quality gate — the synthesizer generates the data, the transcriber scores it. Where two of your models validate each other, you can run sweeps with nobody in the loop (Baseten's "Still" KV compaction uses exactly this trick: ASR re-transcribes the compacted cache).
12.2 Protocol — cheap hardware punishes sloppy measurement twice
- Power/thermal caps are the enemy: a 35 W laptop part throttles within ~2 consecutive benchmark runs — fresh process, few runs, interleaved A/B, report min, checkpoint like preemption is imminent (it is).
- Disk is a first-class constraint (10–40 GB free is normal): audit the HF cache before downloading — and beware config-only stubs: a "cached" model can be 84 KB of metadata with no weights.
- Data without downloads: synthesize pairs locally (cached TTS generates audio, cached ASR + forced alignment generates transcripts).
12.3 When to rent: market rates, Aug 2026 (USD and NGN per GPU-hour, ₦1,338.68/USD)
| GPU | Vast.ai marketplace | RunPod community | RunPod secure | DataCrunch |
|---|---|---|---|---|
| RTX 3090 24 GB | $0.05–0.20 → ₦67–268 | — | — | — |
| RTX 4090 24 GB | $0.06–0.35 → ₦80–469 | $0.34 → ₦455 | $0.69–0.74 → ₦924–991 | — |
| RTX 5090 32 GB | $0.33–1.00 → ₦442–1,339 | varies | — | — |
| A100 80 GB | $0.65–0.90 → ₦870–1,205 | $1.19–1.39 → ₦1,593–1,861 | $1.39–1.49 → ₦1,861–1,995 | — |
| H100 80 GB | $1.20–1.85 → ₦1,606–2,477 | $1.99–2.69 → ₦2,664–3,601 | $2.89–2.99 → ₦3,869–4,003 | ~$2.29 → ₦3,066 |
The spread is the story: same-class H100 listings span $1.49–6.98/hr across providers, and marketplace prices move hourly (budget ₦1,400–1,500/USD for card fees and parallel-market reality). Working rule: iterate locally, rent only to generalize a finding — one overnight 4090 (~₦2k–10k) finishes a QAD or depth-prune sweep; the H100 tier is for headline benchmarks after the laptop says the result is real. Vast interruptible + checkpointing is just §12.2's thermal protocol with extra steps.
13The audio lane: the fal recipe on Stable Audio 3, and mini-DSpark for speech tokens
13.1 Running the §06 stack on stable-audio-3-small — what's left to attack
Stability already ran much of the playbook (§06's own lineage): the Small-Music/SFX post-trained checkpoints ship as a 433M-param flow-matching DiT over 256-dim SAME latents (4096× compression, 44.1 kHz stereo, 266M-param SAME-S autoencoder), pre-distilled and adversarially post-trained to 8 PingPong steps at CFG 1.0 — guidance is already folded and steps already collapsed from the base checkpoints' ~50 steps at CFG≈7. So the cost equation reduces to two live terms and one honest alternative:
- Cheap passes — quantize + QAD: a 433M DiT + 266M autoencoder quantize easily; int8 W8A8 with a short STE quantization-aware distillation (§06's recipe — the velocity field will not forgive quant noise any less than images) fits a 6 GB laptop. FP4 is Blackwell-only; on Ampere hardware int8 is §06's analogue.
- Fewer passes — 8 → 4 steps: TDM/CDM-style steps-aware distillation on top of an already-distilled model, quality-gated by FAD / embedding distance. Stability's own docs concede below-8 trades quality; the live question is whether a second distillation pass, warm-started into the student's own few-step distribution (§09's warm-start lesson applied to samplers), closes that gap.
- Reproduce the full stack from
small-music-base: the pedagogically honest run — CFG-folding + DMD-family step-distillation + QAD, all on a 433M DiT. That's a fine-tune-scale job: ~₦33k–80k on rented 4090s for 3–7 days; pipeline smoke tests run on laptop VRAM.
13.2 Which small models can carry a mini-DSpark — and yes, someone has done it for speech
DSpark = DFlash's block-diffusion drafter + Markov refinement heads (§03). Its interface is architecture-agnostic: fused hidden states from the target, frozen embed + lm_head, standard speculative verification. What matters is whether the target's output token distribution is draftable. The local zoo:
| model | token stream | mini-DSpark verdict |
|---|---|---|
| Qwen3-ASR-0.6B/1.7B | transcript text | strong fit. SpecASR 3.04–3.79× on LLM-ASR (2507.18181); CTC self-speculative drafts 4.4× iRTF (2603.11243); SpeechSpec ran DFlash/EAGLE3 on Qwen3-Omni / Qwen2-Audio at 1.5–1.65× batch-1, WER unchanged |
| Qwen3-TTS-0.6B/1.7B | 12 Hz discrete audio tokens | precedent exists. SpeechSpec reports DSpark/DFlash on CosyVoice2/3's speech-token LLMs at 1.2–3.1× across batch. Draft only the semantic/first codebook — acoustic RVQ layers are high-entropy and naïve acceptance degrades quality (said plainly by the codec-MTP paper, 2410.13839; VADUSA 2410.21951 for the AR-TTS baseline) |
| GLM-OCR (1.33B, ships 1× MTP) | copy-heavy text | best bet. Document layout is near-deterministic → acceptance ≥5 is plausible, and nobody has published it (§11's plan) |
| LFM2-Audio-1.5B | text (audio understanding) | suitable, small ceiling. The conv-heavy LFM2 hybrid breaks nothing — the drafter only consumes hidden states — but conv-dominated decode is already cheap per token, shrinking the bandwidth-bound problem spec-dec exists to solve. Expect ~1.2–1.6× at batch 1, and only if draft+verify is fused so host overhead doesn't eat it (§08's zero-overhead MTP lesson) |
| gemma-4-E2B-it-qat-mobile | text | good lab target: quantized student + drafter = §01's compositionality (QuantSpec-style) on one card |
| whisper-small | transcript text | proven but modest: 2.2× with distil-whisper assisted decoding; encoder-decoder drafter support is thin |
| stable-audio-3 · mimi · htdemucs | continuous latents / non-AR | incompatible by construction — there is no token stream to verify; their lever is §06's (steps × branches × passes), not §03's |
Cost of a mini-DSpark: the drafter is ~one block + Markov heads reading the target's fused hidden states; 50–250K domain samples (fal needed 250K and warm-started, §07), and hidden-state dumps at 0.6B scale land in the tens of GB — nowhere near fal's 38 TB problem. Training is hours on a 4090 (₦3k–10k) or overnight on a laptop-class card, with the acceptance measurement + batch-scaling curve on the rented GPU. And there's the open cell, free for the taking: SpeechSpec published batch-1 numbers only — acceptance × batch × bit-width on ASR/OCR token streams (§01's regime shift meeting §12's laptop constraint) is uncharted.
14Source trail
| cluster | sources |
|---|---|
| fal: Ideogram V4 FP4+QAD+CFG+timestep stack; prompt expander / DSpark | Serving sub-second Ideogram v4 · 1000 tok/s with DSpark · Epilogue fusion · MXFP8 quantizer on Blackwell |
| Baseten: GLM-5, DFlash, video, GPT-OSS, benchmarking | Fastest GLM-5 API · DFlash · GPT-OSS 120B · How to benchmark · Research: "Still" KV compaction |
| speculative lineage | 2211.17192 · 2302.01318 · 2401.10774 (Medusa) · 2402.05109 (Hydra) · 2401.15077 (EAGLE) / 2406.16858 / 2503.01840 (EAGLE-2/3) · 2602.06036 (DFlash) · 2412.19464 (DeepSeek-V3 MTP) · 2602.06019 (post-hoc MTP self-distillation) · 2309.08168 (Draft & Verify) · 2404.18911 (Kangaroo) · 2502.10424 (QuantSpec) · 2311.03511 (prompt lookup) · 2507.20127 (speculative cascades) |
| quantization: PTQ → QAT → QAD | 2210.17323 (GPTQ) · 2306.00978 (AWQ) · 2211.10438 (SmoothQuant) · 2407.11062 (EfficientQAT) · 2506.09104 (UPQ) · 2601.20088 (NVFP4 QAD) · GuidedQuant (ICML 2025) |
| KV cache & long context | 2309.06180 (PagedAttention) · 2412.19442 (KV survey) · 2402.02750 (KIVI) · 2401.18079 (KVQuant) · 2306.14048 / 2309.17453 / 2404.14468 (H2O / StreamingLLM / SnapKV) · 2407.02490 (MInference) · 2605.17613 (VeriCache) |
| distillation & compression | 1503.02531 · 1603.09785 · 2307.13386 (GKD) · 2305.08445 (MiniLLM) · 2310.06694 (Sheared-LLaMA) · 2407.14679 (Minitron) · 2311.00430 (distil-whisper) · 2305.13245 (GQA) · 2405.04434 (MLA) · 2205.14135 (CALM) · 2404.16710 (LayerSkip) · 2502.18600 (Chain of Draft) · 2412.19792 (InfAlign) |
| diffusion step distillation | 2303.01469 (CM) · 2310.04378 (LCM) · 2202.00512 (progressive) · 2311.18828 / 2405.14867 (DMD/DMD2) · 2311.17042 (ADD) · 2503.06674 (TDM) · 2502.06085 (MeanFlow) · 2207.12598 (guidance distillation) |
| diffusion caching & parallel inference | 2312.00858 (DeepCache) · 2411.19108 (TeaCache) · 2410.19355 (FasterCache) · 2505.13389 (VSA) · 2402.19481 (DistriFusion) · 2405.14430 (PipeFusion) · xDiT |
| serving systems & kernels | 2309.06180 (vLLM) · 2308.16369 / 2403.02310 (Sarathi) · 2401.09670 (DistServe) · 2407.00079 (Mooncake) · 2312.07104 (SGLang) · 2407.08608 (FA3) · 2408.11743 (Marlin) · 2401.06118 (AQLM) · 2402.04396 (QuIP#) · 2501.01005 (FlashInfer) · 2410.02367 / 2505.11594 (SageAttention) · KTransformers SOSP'25 · 2311.03285 (S-LoRA) |
| multimodal / VLM efficiency | 2403.06764 (FastV) · 2403.15388 (LLaVA-PruMerge) · fal DSpark · Baseten GLM-5 |
| vLLM engineering blog: spec-dec toolchain, KV platform layer, Blackwell stacks, CI discipline | Inside vLLM: Anatomy · DSpark adaptive verification · 25K TPS/GPU (Qwen3.5) · GLM-5.2 on 24×B300 · FP8 KV state · TurboQuant study · Decode Context Parallelism · KV offloading connector · Hybrid SSM disaggregation · DiffusionGemma · CI & benchmarking discipline |
| Baseten research track | RadixMLP intra-batch dedup · "Still" KV compaction · Repeated KV for agents · BYO SWE-grep |
| GPU rental rates (Aug 2026) + FX | vast.ai marketplace snapshots (gpu.org, gpuhosted review) · RunPod pricing · DataCrunch H100 rate · FX: open.er-api.com USD→NGN 1,338.68 (2026-08-29) |
| audio lane: speculative decoding for speech/ASR/TTS; Stable Audio 3 architecture | 2410.21951 (VADUSA) · 2410.13839 (codec MTP + spec-dec) · 2507.18181 (SpecASR) · 2603.11243 (CTC self-speculative) · SpeechSpec (DFlash/DSpark on Qwen3-Omni & CosyVoice) · Whisper assisted decoding · stable-audio-3 docs (433M DiT · 8-step PingPong · CFG 1.0) |