lab

Qwen3-ASR — Latency & Bottleneck Analysis

Where the 2.65 seconds go when transcribing 60 s of music-with-vocals on a laptop RTX 3050. Every stage, module and layer timed individually with CUDA events; decode attention — not lm_head — turns out to be the dominant cost.

01Method & setup

propertyvalue
modelQwen3-ASR-0.6B, fp16, greedy, max_new_tokens 512
stacktransformers 4.57.6 backend (baseline @ 7c6daf7), torch 2.13.0+cu130
gpuRTX 3050 6GB Laptop (40W), driver 580.119.02
clipbad_guy_60s.wav — 60.0s, vocals + BGM, 44.1 kHz → 16 kHz
prompt / output795 prompt tokens (779 audio) → 130 generated tokens, EOS-stopped
instrumentationCUDA events around each stage and module; coarse pass (layer-level wraps, near-zero overhead) for absolutes, fine pass (inner attn/mlp/norm wraps) for module splits; warmup 2 runs; values normalized to the coarse total

02Stage waterfall — total wall 2.74s

GPU decode eats 88.5% of the wall clock; the whole encoder+prefill is 8.4%; the CPU front-end (librosa resample, steady-state) is 2.2%.

Wall-clock stages, 60s clip
total 2743 ms · hover for exact values
cpu load+resample: 59.4ms (2.2%) mel+tokenize: 17.4ms (0.6%) cpu→gpu: 8.1ms (0.3%) encoder+prefill: 231.0ms (8.4%) decode: 2426.9ms (88.5%) 2.2% resample 8.4% prefill 88.5% decode
CPU front-end (85ms) GPU prefill (231ms) GPU decode (2427ms)
stagems% wall% gpubound by
cpu load + resample (librosa)59.42.2%—CPU; first call in a process pays ~600ms numba JIT
cpu mel + tokenize (processor)17.40.6%—CPU (STFT, 128 mel bins)
cpu → gpu transfer8.10.3%—H2D copies
gpu encoder + decoder prefill231.08.4%8.7%attention + conv stack (see §04)
gpu decode (130 tokens)2426.988.5%91.3%per-token latency — see §03
Takeaway Decode is the game: 91% of GPU time. Optimizing prefill or the encoder can only ever shave ~8% of wall time; every decode optimization pays off 130× per clip.

03Decode — 18.58 ms per generated token

Attention is 48.6% of decode — more than MLP and lm_head combined. The 28-layer decoder stack dominates (84.9%); lm_head is only 13.5%.

Per-token breakdown (normalized to coarse total)
18.58 ms/token · 2415 ms per 60s clip · hover for values
decoder attn (28 layers): 9.033ms (48.6%) decoder mlp (28 layers): 5.114ms (27.5%) lm_head: 2.504ms (13.5%) norms (56): 1.189ms (6.4%) resid+layer-other: 0.428ms (2.3%) rotary mrope: 0.262ms (1.4%) embed+final_norm: 0.047ms (0.3%) 48.6% attn 27.5% mlp 13.5% lm_head
attention MLP lm_head norms
modulems/token% decodems per 60s clip% wall
decoder attention (28 layers)9.03348.6%117442.8%
decoder MLP (28 layers)5.11427.5%66524.3%
lm_head (1024→151936)2.50413.5%32511.9%
RMSNorms (56/layer-pair)1.1896.4%1555.7%
residuals + layer glue0.4282.3%562.0%
rotary (MRoPE, fp32)0.2621.4%341.2%
embed + final_norm0.0470.3%60.2%
The surprise The paper analysis predicted lm_head ≈ 93% of per-token memory traffic — true in bytes, but on this GPU attention still wins the wall clock: it moves only ~0.3 GB/token but takes 9.0 ms because ~300 tiny kernels are launched per token across 28 layers. lm_head is a single fat GEMM running at 83% memory efficiency (2.50 ms vs 2.07 ms ideal for 311 MB); attention runs at ~20% of what its memory traffic allows. This is a latency/launch-bound bottleneck, not FLOP- or bandwidth-bound.

04Prefill & encoder — 236 ms

The encoder's conv stack is a surprise: 51 ms (59%) of the audio tower goes to 3 conv2d + conv_out on small-channel tiles, not the transformer layers.

Prefill module breakdown
236.1 ms total · hover for values
audio_tower 87.0ms conv2d×3 + conv_out + pos: 51.2ms (59%) encoder attn (18): 14.5ms (17%) encoder fc1+fc2 (18): 17.3ms (20%) encoder norms+resid: 4.5ms (5%) decoder layers 126.6ms dec attn (28): 70.3ms (55%) dec mlp (28): 38.6ms (30%) dec norms (56): 15.7ms (12%) dec resid: 2.4ms (2%) lm_head 21.1ms lm_head over 795 prompt positions: 21.1ms bar widths to scale, 236.1ms = 1000 units
modulems% prefillnote
conv2d×3 + conv_out + pos-emb (non-layer)51.221.7%chunked conv on 60×1×128×100 tiles; 480-ch kernel
encoder layers (18)35.815.2%attn 14.5 · fc 17.3 · norms 4.5
decoder layers (28)126.653.6%attn 70.3 (55%) · mlp 38.6 (30%) · norms 15.7 (12%)
lm_head (all 795 positions)21.18.9%teacher-forced logits for the whole prompt
rotary + embed + final_norm0.90.4%once per chunk
Takeaway Even if prefill became free, RTF only drops from 0.044 to 0.041. The encoder is not the hot path — but its conv stack is 59% of itself, and the per-sample loop in get_audio_features() would matter only for batched requests.

05Per-layer spread — decode

Attention: 0.329–0.468 ms/layer/token (mean 0.341, layer 0 the outlier). MLP: 0.190–0.201 — perfectly uniform, as expected for identical GEMMs.

28 decoder layers, ms per token
attention on top, MLP below · height ∝ time
attn mlp
hover a bar

06Bottleneck ranking

Ranked by wall-clock impact on the standard 60s clip (130 generated tokens).

Wall-time impact
decode components scale with token count; CPU stages are fixed per clip · total 2743ms
1decode attention (28L)
1174 ms · 42.8%
2decode MLP (28L)
665 ms · 24.3%
3decode lm_head
325 ms · 11.9%
4decode norms (56)
155 ms · 5.7%
5prefill decoder attn
70 ms · 2.6%
6CPU resample (librosa)
59 ms · 2.2%
7encoder conv stack
45 ms · 1.7%
8prefill decoder mlp
39 ms · 1.4%
9encoder layers (rest)
37 ms · 1.4%
10rotary (MRoPE, per step)
34 ms · 1.2%
11mel + transfer + embed
27 ms · 1.0%
#stage / modulems/clip% wallbottleneck typebest lever
1decode attention117442.8%launch/latency-bound (~300 kernels/token)CUDA graphs / torch.compile / fused attn
2decode MLP66524.3%memory-bound (529MB/token)persistent GEMMs, fp8, graph capture
3decode lm_head32511.9%memory-bound (311MB/token, 83% eff.)fp8/int8 lm_head, speculative decode
4decode norms1555.7%168 tiny fp32-roundtrip kernels/tokenfused RMSNorm, stay fp16
5prefill decoder attn702.6%compute (parallel over 795 tokens)fine as-is; sdpa already
6CPU resample592.2%CPU, serial before GPUtorchcodec (2× step), overlap with decode
7encoder conv stack451.7%small-channel conv tilesfuse 3 convs, bigger tiles
8prefill decoder mlp391.4%compute—
9encoder layers371.4%computebatch the per-sample loop
10rotary341.2%fp32 recompute per stepcache freqs, fuse into q/k proj

07Root causes

Why attention is 48.6% of decode Per token, each of the 28 layers launches ~10 kernels for attention alone: q/k/v projections, reshapes/transposes, GQA repeat_kv, QKT, mask add, fp32 softmax, AV, transpose, .contiguous(), o_proj. At seqlen ≈ 925, every one of these is a tiny (16×128×925) op — the GPU spends its time on launch latency and sync stalls, not math (13.6 MFLOP vs ~9 ms = ~1500× over FLOP-bound). KV reads (~0.3 GB/token) would only justify ~2 ms. This is the #1 target for a minimal runtime: capture the whole per-token step as one CUDA graph, or fuse attention into 1–2 kernels.
Why lm_head is only #4 despite being 93% of bytes A single 1024×151936 GEMM reads 311 MB fp16 per token and runs at 83% of memory-bandwidth efficiency. It is the largest single kernel, but the memory system is not the critical resource — kernel count is. fp8 lm_head still wins (2.5 → ~1.3 ms, plus VRAM headroom), and speculative decoding with the 0.6B itself as draft would skip most lm_head calls.
Why the CPU front-end is only 2.2% librosa's resampler takes ~60ms steady-state for a 60s clip (the first call in a process pays ~600ms numba JIT — a profiling trap). Swapping it for torchcodec (ffmpeg decode + resample + mono in one pass) halves it to ~30ms — a 2.0–2.1× step speedup, but only ~1% of total wall time. Front-end choice does NOT change transcription (verified bit-identical model inputs).
Norms are real money at this size 56 RMSNorms per token (2 per layer × 28), each doing an fp32 up-cast, variance reduction, and down-cast ≈ 3 kernels — 168 kernels/token, 1.19 ms (6.4%). Fusing the four norms of a layer into the attention/MLP kernels (or a single fused RMSNorm) is nearly free and cuts ~2/3 of that.

08Optimization roadmap — measured targets

stepchangeexpected ms/token (now 18.58)expected RTF (now 0.044)effort
1torch.compile on thinker (declared fullgraph-capable)18.6 → ~11–130.044 → ~0.031low
2hand decode loop + CUDA graphs (static 1-token step)→ ~9–10→ ~0.027medium
3fp8/int8 lm_head→ ~7–8→ ~0.023medium
4torchcodec front-end (done) — 2× step, ~1% overallwall −30 ms fixed→ ~0.021done
5speculative decode (draft = 0.6B) — optional later→ ~4–5→ ~0.012high
Sanity check Baseline steady-state: 2.64 s per 60 s clip (RTF 0.0442, results.md). Every optimization is measured against scripts/bench_asr.py with the same 60 s protocol; profiling is reproducible via scripts/profile_asr.py (coarse + fine passes, CUDA events).