Qwen3-ASR — Latency & Bottleneck Analysis
Where the 2.65 seconds go when transcribing 60 s of music-with-vocals on a laptop RTX 3050. Every stage, module and layer timed individually with CUDA events; decode attention — not lm_head — turns out to be the dominant cost.
01Method & setup
| property | value |
|---|---|
| model | Qwen3-ASR-0.6B, fp16, greedy, max_new_tokens 512 |
| stack | transformers 4.57.6 backend (baseline @ 7c6daf7), torch 2.13.0+cu130 |
| gpu | RTX 3050 6GB Laptop (40W), driver 580.119.02 |
| clip | bad_guy_60s.wav — 60.0s, vocals + BGM, 44.1 kHz → 16 kHz |
| prompt / output | 795 prompt tokens (779 audio) → 130 generated tokens, EOS-stopped |
| instrumentation | CUDA events around each stage and module; coarse pass (layer-level wraps, near-zero overhead) for absolutes, fine pass (inner attn/mlp/norm wraps) for module splits; warmup 2 runs; values normalized to the coarse total |
02Stage waterfall — total wall 2.74s
GPU decode eats 88.5% of the wall clock; the whole encoder+prefill is 8.4%; the CPU front-end (librosa resample, steady-state) is 2.2%.
| stage | ms | % wall | % gpu | bound by |
|---|---|---|---|---|
| cpu load + resample (librosa) | 59.4 | 2.2% | — | CPU; first call in a process pays ~600ms numba JIT |
| cpu mel + tokenize (processor) | 17.4 | 0.6% | — | CPU (STFT, 128 mel bins) |
| cpu → gpu transfer | 8.1 | 0.3% | — | H2D copies |
| gpu encoder + decoder prefill | 231.0 | 8.4% | 8.7% | attention + conv stack (see §04) |
| gpu decode (130 tokens) | 2426.9 | 88.5% | 91.3% | per-token latency — see §03 |
03Decode — 18.58 ms per generated token
Attention is 48.6% of decode — more than MLP and lm_head combined. The 28-layer decoder stack dominates (84.9%); lm_head is only 13.5%.
| module | ms/token | % decode | ms per 60s clip | % wall |
|---|---|---|---|---|
| decoder attention (28 layers) | 9.033 | 48.6% | 1174 | 42.8% |
| decoder MLP (28 layers) | 5.114 | 27.5% | 665 | 24.3% |
| lm_head (1024→151936) | 2.504 | 13.5% | 325 | 11.9% |
| RMSNorms (56/layer-pair) | 1.189 | 6.4% | 155 | 5.7% |
| residuals + layer glue | 0.428 | 2.3% | 56 | 2.0% |
| rotary (MRoPE, fp32) | 0.262 | 1.4% | 34 | 1.2% |
| embed + final_norm | 0.047 | 0.3% | 6 | 0.2% |
04Prefill & encoder — 236 ms
The encoder's conv stack is a surprise: 51 ms (59%) of the audio tower goes to 3 conv2d + conv_out on small-channel tiles, not the transformer layers.
| module | ms | % prefill | note |
|---|---|---|---|
| conv2d×3 + conv_out + pos-emb (non-layer) | 51.2 | 21.7% | chunked conv on 60×1×128×100 tiles; 480-ch kernel |
| encoder layers (18) | 35.8 | 15.2% | attn 14.5 · fc 17.3 · norms 4.5 |
| decoder layers (28) | 126.6 | 53.6% | attn 70.3 (55%) · mlp 38.6 (30%) · norms 15.7 (12%) |
| lm_head (all 795 positions) | 21.1 | 8.9% | teacher-forced logits for the whole prompt |
| rotary + embed + final_norm | 0.9 | 0.4% | once per chunk |
get_audio_features()
would matter only for batched requests.
05Per-layer spread — decode
Attention: 0.329–0.468 ms/layer/token (mean 0.341, layer 0 the outlier). MLP: 0.190–0.201 — perfectly uniform, as expected for identical GEMMs.
06Bottleneck ranking
Ranked by wall-clock impact on the standard 60s clip (130 generated tokens).
| # | stage / module | ms/clip | % wall | bottleneck type | best lever |
|---|---|---|---|---|---|
| 1 | decode attention | 1174 | 42.8% | launch/latency-bound (~300 kernels/token) | CUDA graphs / torch.compile / fused attn |
| 2 | decode MLP | 665 | 24.3% | memory-bound (529MB/token) | persistent GEMMs, fp8, graph capture |
| 3 | decode lm_head | 325 | 11.9% | memory-bound (311MB/token, 83% eff.) | fp8/int8 lm_head, speculative decode |
| 4 | decode norms | 155 | 5.7% | 168 tiny fp32-roundtrip kernels/token | fused RMSNorm, stay fp16 |
| 5 | prefill decoder attn | 70 | 2.6% | compute (parallel over 795 tokens) | fine as-is; sdpa already |
| 6 | CPU resample | 59 | 2.2% | CPU, serial before GPU | torchcodec (2× step), overlap with decode |
| 7 | encoder conv stack | 45 | 1.7% | small-channel conv tiles | fuse 3 convs, bigger tiles |
| 8 | prefill decoder mlp | 39 | 1.4% | compute | — |
| 9 | encoder layers | 37 | 1.4% | compute | batch the per-sample loop |
| 10 | rotary | 34 | 1.2% | fp32 recompute per step | cache freqs, fuse into q/k proj |
07Root causes
repeat_kv, QKT, mask add, fp32 softmax, AV,
transpose, .contiguous(), o_proj. At seqlen ≈ 925, every one of these is a tiny
(16×128×925) op — the GPU spends its time on launch latency and sync stalls, not math
(13.6 MFLOP vs ~9 ms = ~1500× over FLOP-bound). KV reads (~0.3 GB/token) would only justify
~2 ms. This is the #1 target for a minimal runtime: capture the whole per-token step as one
CUDA graph, or fuse attention into 1–2 kernels.
08Optimization roadmap — measured targets
| step | change | expected ms/token (now 18.58) | expected RTF (now 0.044) | effort |
|---|---|---|---|---|
| 1 | torch.compile on thinker (declared fullgraph-capable) | 18.6 → ~11–13 | 0.044 → ~0.031 | low |
| 2 | hand decode loop + CUDA graphs (static 1-token step) | → ~9–10 | → ~0.027 | medium |
| 3 | fp8/int8 lm_head | → ~7–8 | → ~0.023 | medium |
| 4 | torchcodec front-end (done) — 2× step, ~1% overall | wall −30 ms fixed | → ~0.021 | done |
| 5 | speculative decode (draft = 0.6B) — optional later | → ~4–5 | → ~0.012 | high |
scripts/bench_asr.py with the same 60 s protocol; profiling is
reproducible via scripts/profile_asr.py (coarse + fine passes, CUDA events).