lab

Can CTC make Qwen3-ASR faster?

Yes—but the most promising version is not “train another tiny language model and put a CTC head on it.” For Qwen3-ASR-0.6B on an RTX 3050 6GB, the best research bet is a tiny, non-autoregressive CTC transcript head attached to the audio features the model already computes, followed by strict block verification in the existing Qwen decoder. It fits the memory budget and attacks the measured 2.43 s autoregressive hot loop. It is also a research project, not the first optimization to ship.

Evidence labels used below: measured here published result analytical bound hypothesis—benchmark required

01Verdict in plain English

bottom line

Plausibility: high for a prototype, medium for a robust multilingual speedup, low for a guaranteed paper-sized speedup without careful runtime work. CTC is unusually well matched to ASR because the output is tightly anchored to already-known audio. The Qwen audio tower already emits about 13 feature vectors per second, while the measured 60 s clip emits only 130 text tokens. That temporal slack is exactly what CTC needs.

Build first

Measure an oracle verifier before training anything. Feed known-correct future tokens in blocks of 4/8/16/full transcript and measure verification cost. This tells us whether the RTX 3050 can profit from speculation at all.

Best learned draft

Freeze Qwen’s audio tower; train a small byte/character or pruned-token CTC head. It drafts the whole transcript in one parallel pass and adds tens of MB, not another 0.6B model.

Ship before research

Benchmark native Transformers + torch.compile, static KV, and CUDA-graph/fusion work first. The current profile is launch-bound; these changes are cheaper and stack with CTC later.

The realistic first target is 1.5–2.7× end-to-end on the profiled long clip if a strict CTC draft yields a long accepted prefix. The analytical best case—nearly the entire transcript accepted after one draft and one verification—is roughly 4.8–6.9×, or RTF 0.0066–0.0094, but that is a ceiling scenario, not a forecast. A decoder-side MTP/Medusa/CTC-drafter is easier to make lossless but more likely to land around 1.3–2.0× end-to-end on this small target.

the crucial distinction

A confidence threshold that returns the CTC transcript directly is a quality–latency trade. Strict token equality against Qwen’s greedy argmax is lossless greedy speculative decoding. The 4.4× CTC-ASR result in the literature uses the former at its fastest point and reports a 12% relative WER increase. Do not call that operating point “exact.”

02The measured machine, model, and bottleneck

This report is anchored to the project’s CUDA-event profile, not a generic server GPU. The target is Qwen3-ASR-0.6B, fp16 greedy decoding, on a 40 W RTX 3050 Laptop GPU with 6GB VRAM. The test clip is 60 s of music with vocals; it creates 795 prompt tokens (779 audio) and 130 output tokens.

stagemeasured timesharemeaning for CTC/speculation
CPU load, resample, mel, transfer84.9 ms3.1%not accelerated by token speculation
audio encoder + decoder prefill231.0 ms8.4%CTC can reuse audio features; target verification still needs decoder work
autoregressive decode, 130 tokens2426.9 ms88.5%the addressable section
total profile wall time2743 ms100%RTF 0.0457; about 21.9× real time
decoder attention, 28 layers
9.03 ms/tok
decoder MLP
5.11 ms/tok
lm_head, 151,936 vocab
2.50 ms/tok
norm, glue, rotary, embedding
1.93 ms/tok

The surprise is that lm_head is 93% of decoder weight traffic but only 13.5% of measured decode time. Attention wins because roughly 300 small kernels are launched per output token. This makes the loop launch/latency-bound, which is favorable to multi-token verification: one block forward can turn many tiny one-token operations into fewer, wider operations.

Amdahl ceiling = 2743 / (2743 − 2426.9) = 8.68×

Even a magical zero-cost decoder cannot exceed 8.68× without also improving the front end and prefill. Ordinary 2× decode acceleration yields only 1.79× end-to-end; 3× decode yields 2.44×. Any proposal must be judged against this measured denominator.

03Speculative decoding: what work actually disappears?

Normal greedy decoding calls the 28-layer target once for every new token. A speculative round obtains several cheap guesses, scores those known candidate positions in one causal target pass, commits the longest prefix whose candidate tokens equal Qwen’s argmax, then emits the target correction at the first mismatch.

audio + committed text→ cheap draft y₁…yₖ→ one Qwen block verification→ accept exact prefix→ target correction

The target is still causally correct: logit position i sees only the committed prefix and candidate tokens before i. The win comes from fewer target invocations and better GPU utilization per invocation—not from deleting the target’s authority.

Greedy exactness

Qwen3-ASR’s project baseline is greedy, so the simplest contract is enough. Accept draft token y[i] only when it equals argmax(target_logits[i]). At the first mismatch, discard it and the suffix, append the target argmax, crop speculative KV state to the committed prefix, and begin another round. Subject to the same processors and numerics, the transcript equals ordinary greedy decoding.

Sampling exactness is harder with acoustic CTC

Classical speculative sampling accepts a proposed token with min(1, p(y)/q(y)) and samples a correction from normalized [p−q]₊. That requires the draft and target to expose aligned next-token distributions over the same vocabulary and conditioning history. A whole-transcript acoustic CTC model defines a sequence probability by summing alignments, not a convenient autoregressive q(yᵢ|y<i). Therefore the recommended acoustic CTC path is exact for greedy decoding; exact stochastic decoding would require a much more complicated sequence-level proposal correction or a conventional AR/MTP draft.

speed model for one speculative round

Let A be committed tokens per round, D draft time, and V(k) target verification time for a block of k candidates. Baseline time for those tokens is A × 18.58 ms. The round is profitable only if:

D + V(k) + rollback overhead < A × 18.58 ms

Acceptance rate alone is insufficient. A draft that accepts five tokens but spends 70 ms producing them may lose to five ordinary 18.58 ms steps only after verification is added. Measure time-per-committed-token, not draft accuracy in isolation.

04CTC from first principles

Connectionist Temporal Classification solves a mismatch: speech gives a long sequence of acoustic frames, but training data gives only a shorter transcript with no frame-to-character alignment. CTC adds a blank symbol ∅ and considers every frame-level path that collapses to the transcript.

The collapse rule

For a frame path, first merge consecutive repeated labels, then remove blanks. Several paths can yield the same text:

[∅, c, c, ∅, a, ∅, t] → [c, a, t]
[c, ∅, a, a, ∅, t, t] → [c, a, t]

A blank between repeated labels is significant. To output “book,” CTC needs a path such as b o ∅ o k; adjacent o o without a separator would collapse to one o.

The probability and loss

At each of T audio steps the head emits a distribution over vocabulary plus blank. Conditioned on the encoder features, CTC factorizes a path across time, then sums the probability of all paths that collapse to target y:

P(y | x) = Σπ : B(π)=y ∏t=1…T P(πₜ | hₜ)
LCTC = −log P(y | x)

Enumerating paths is exponential, so the forward–backward dynamic program walks an expanded label sequence with blanks inserted between labels. Its time is O(TU) for transcript length U; there is no required frame alignment. PyTorch’s CTCLoss implements this log-space dynamic program.

Inference

  • Greedy CTC: take per-frame argmax and collapse. It is one parallel head pass plus a cheap scan.
  • Prefix beam CTC: keep multiple collapsed prefixes while combining blank/non-blank probabilities. It improves accuracy but adds serial CPU/GPU work.
  • Confidence: frame entropy, blank probability, sequence score, or beam margin can decide how much of the draft deserves verification.

Why CTC is fast

No previous output token is needed to produce the next frame distribution. All frame logits are computed together; collapse is linear time.

Why CTC makes errors

The conditional independence assumption is weak for punctuation, spelling, language-model context, and token boundaries. The encoder carries context, but the output head itself does not autoregress.

05“CTC draft model” can mean three different systems

designinput to draftdraft unitrelationship to this project
Acoustic CTC transcript headaudio encoder featureswhole transcript in parallelrecommended experiment reuses Qwen audio compute and avoids an AR draft
Decoder-side CTC draftertarget decoder hidden states at current text stepseveral future token slots plus blanks/repeatsthe NeurIPS 2024 “CTC-drafter”; one draft transformer layer + token tree
Ordinary CTC ASR as a separate draft modelraw/mel audiowhole transcriptsimple conceptually, but duplicates the encoder and spends scarce 6GB VRAM

What the decoder-side CTC-drafter paper actually did

Wen, Gui, and Feng froze 7B–33B Vicuna targets, fed the target’s hidden states into one transformer draft layer, trained future positions with sequence-level CTC loss on distilled outputs, formed top-k candidate trees, removed blanks/repeats, modified the tree attention mask, and verified candidates in the base model. It reported 3.40–3.56 accepted tokens per step and 2.20–2.78× speedup. Training used four RTX 3090 GPUs for about two days—roughly 192 GPU-hours.

This is relevant as an algorithmic template, but it did not use audio CTC and it did not test a sub-billion-parameter target. Its target is expensive enough to amortize a sophisticated draft head. Our target’s whole decode step is only 18.58 ms, so draft and tree overhead consume a larger fraction.

What the 2026 CTC-encoder ASR paper did

Saon et al. used a 440M CTC encoder inside a speech-aware LLM. If every CTC frame entropy was below a threshold, it returned the CTC transcript directly. Otherwise the causal LLM scored all CTC tokens in one forward pass; tokens above a likelihood threshold were accepted and AR decoding resumed from the accepted prefix. The fastest reported Open ASR point improved inverse RTF by 4.4× with 12% relative WER degradation. At high-accuracy operating points, LLM verification sometimes improved WER through complementary CTC/LLM errors.

That paper proves the broad idea is plausible. Its confidence acceptance is deliberately approximate, and its encoder was trained as CTC from the start. Qwen’s audio tower was not, so a small head may need more than a linear probe.

06Why Qwen3-ASR is a promising—but awkward—host

The useful properties

  • The 60 s prompt has 779 acoustic positions for 130 generated BPE tokens: about six acoustic positions per output token.
  • The audio encoder runs once in the Transformers generation path; subsequent steps are decoder-only. A head can consume its existing output rather than re-encode audio.
  • ASR is low-entropy compared with open-ended chat. SpecASR reports 3.04–3.79× over AR, and ParaASR reports an average five accepted tokens out of six proposals, supporting longer drafts in speech.
  • The project uses greedy decoding, so strict argmax-prefix verification can preserve the baseline transcript without implementing rejection sampling.

The awkward properties

  • Huge vocabulary: 151,936 tokens. A dense 1024→151,936 fp16 head is about 311MB. Applying a similar head to all 779 audio positions creates roughly 118M logits (about 236MB fp16) and ~121B MACs for one clip.
  • The target is already small: a second 0.6B draft is not cheap enough. Qwen has no smaller released ASR sibling below 0.6B.
  • Audio features are semantic tokens, not guaranteed CTC frames: the tower uses convolutional downsampling and windowed bidirectional transformer layers, then projects into the text hidden size. Local alignment may be learnable, but it was not the pretrained objective.
  • Language and formatting tokens: Qwen emits language X<asr_text>, punctuation, casing, song text, and multilingual BPE. Ground-truth transcripts may not match the exact teacher normalization needed for high speculative acceptance.
  • Current generation API: GenerationMixin.generate hides cache transactions. Efficient block verification and rollback will probably need a custom loop, static cache, or native speculator integration.
vocabulary recommendation

Do not begin with a full 151k CTC projection. Start with UTF-8 bytes plus blank for universal coverage, or a corpus-derived subset of Qwen token IDs plus byte fallback. Bytes make the head tiny and can be decoded to text then retokenized by Qwen. A pruned native-token alphabet avoids retokenization changes but needs a clear OOV path. Benchmark both.

07Hardware-specific gain estimates

These are scenario calculations from the measured 2743 ms profile, not new GPU benchmarks. The non-decode floor is 316 ms. For the acoustic CTC path, assume the already-paid audio features are reused and CTC draft + target block verification adds 80–250 ms beyond the existing prefill. That interval must be replaced with an oracle measurement.

strict CTC outcome on 130 tokensestimated totalend-to-end speedupRTFinterpretation
whole draft accepted0.40–0.57 s4.8–6.9×0.0066–0.0094analytical ceiling scenario; very unlikely on every clip
75% accepted as useful prefix1.00–1.17 s2.3–2.7×0.0167–0.0196strong, plausible target after distillation
50% accepted as useful prefix1.61–1.78 s1.5–1.7×0.0268–0.0297still useful if runtime is clean
25% accepted2.22–2.39 s1.15–1.24×0.0369–0.0398probably not worth the complexity
early mismatch / no useful prefix>2.74 s<1×>0.0457draft and verification are pure overhead

Whole-transcript strict verification has a sharp failure mode: one early token mismatch prevents accepting later tokens even when the strings realign. Blockwise verification (for example 8–16 tokens), adaptive block length, and recycling the aligned suffix are likely better robust designs. SpecASR’s main contribution is exploiting exactly this ASR realignment behavior.

What I would forecast before the oracle benchmark

pathlikely end-to-end rangeconfidence
native HF + compile/static runtime work1.3–2.0×medium; official A100 batch-4 result is 2.4×, but batch-1 3050 differs
decoder MTP / Medusa / decoder-CTC head1.3–2.0×; stretch ~2.4×low–medium until small-block verification is timed
strict acoustic CTC + block verification1.5–2.7×; best-case 4.8–6.9×low before CTC acceptance is measured
relaxed whole-hypothesis CTC gate2–5×medium that it is fast; quality change must be chosen, not hidden
separate AR 0.6B draft for 1.7B targetpotentially useful, but out of current scopeproject has explicitly selected 0.6B as the target
do not multiply marketing numbers

A 2× compiled runtime and a 2× speculator rarely compose to 4×. Compilation makes the baseline target step cheaper, which raises the acceptance needed to amortize drafting. Re-run the profitability equation after every runtime optimization.

08Recommended architecture

mel features→ Qwen audio tower once→ audio embeddings↗ tiny CTC head→ draft text→ Qwen block verifier→ exact transcript

Draft head, version 0

  • Tap audio_tower(...).last_hidden_state after the existing projector, shape approximately [batch, T, 1024] for 0.6B.
  • Apply LayerNorm → Linear(1024, alphabet+blank). Begin with a byte alphabet; if the probe underfits, add one depthwise temporal convolution or one small bidirectional transformer/Conformer block.
  • Freeze Qwen entirely. Detach audio features before the CTC head. This makes the head cheap to train and keeps the baseline model immutable.
  • Train on teacher-greedy Qwen transcripts for acceptance alignment. Mix ground-truth CTC loss only if you are willing to trade exact match rate for potentially better WER.
  • Greedy-collapse for the first runtime. Add a tiny prefix beam only if its WER/acceptance gain exceeds its latency.

Verifier, version 0

  • Tokenize the CTC text with the exact Qwen tokenizer and prepend/force the exact language marker policy used by the baseline.
  • Verify in blocks of k ∈ {4, 8, 16}. Run the target causally over each candidate block using the committed KV cache.
  • Commit equal argmax tokens. At mismatch, append the target correction, crop the KV cache, and recycle any draft suffix that realigns after the correction.
  • Choose the next k from CTC confidence: longer for low-entropy spans; shorter around uncertain frames, punctuation, language switches, or repeated labels.
  • Keep a fallback to the current eager greedy loop. Every optimization should be switchable and compared token-for-token.

Optional fast mode

After strict mode is stable, add an explicitly named approximate mode:

  1. Return CTC directly only when calibrated sequence confidence exceeds a threshold.
  2. Otherwise score the full CTC hypothesis once in Qwen.
  3. Accept tokens above a tuned likelihood threshold and fall back to AR from the first failure.

This mirrors the 2026 self-speculative CTC-ASR paper and may reach much lower RTF. It must expose its threshold, WER delta, and confidence calibration; it cannot inherit the “same output” claim from strict mode.

09Training data, VRAM, time, and dollar cost

Data strategy

goaldatalabelswhat it proves
pipeline proof10–50 h, one language/domainexisting transcripts + Qwen teacher outputalignment, head capacity, verifier profitability
credible English head100–1,000 h diverse speechmostly teacher-distilled; held-out human truthacceptance across clean/noisy/music/accent sets
Qwen-like multilingual robustnessseveral thousand hours, balanced languagesteacher output plus curated truthstill far below Qwen pretraining scope; expect tail-language gaps

For speculation, teacher labels matter more than ordinary WER training: punctuation, casing, normalization, language tags, and BPE choices must imitate Qwen’s greedy decisions. Generate the teacher transcript once and cache it. Keep human transcripts for WER evaluation and optionally a secondary loss.

Memory

  • Qwen3-ASR-0.6B-hf weights are about 1.56GB on disk. The project’s fp16 runtime estimate is about 1.7GB plus working memory.
  • A 1024→257 byte+blank head has only ~0.26M weights; even a 5k-symbol head is ~5.1M weights (~10MB fp16).
  • A full 151,937-way CTC head is ~155.6M weights (~311MB fp16) and produces a large T×V logit tensor. It still fits inference in 6GB, but it is a poor first experiment and more awkward to train.
  • Frozen-head training can fit the 6GB card with batch 1 and gradient accumulation. Full-model fine-tuning is not a sensible local plan; top-layer tuning should use checkpointing/LoRA or a 24–48GB cloud GPU.

Estimated training budget

These ranges include iteration and evaluation, not just one clean epoch. Actual time depends more on audio loading, padding, sequence length, and whether encoder features are cached than on the tiny head.

experimentcompute estimateRunPod reference cost*comment
linear byte/char probe, 10–100 h1–5 GPU-h on 24GB class$0.27–$3.45cache audio features; multiple seeds dominate
small temporal head, 100–1,000 h5–30 GPU-h$1.35–$20.70A5000/4090; frozen Qwen tower
unfreeze top audio layers / LoRA20–100 GPU-h$5.40–$6924–48GB preferred; watch catastrophic drift
decoder-side CTC/MTP research run20–100+ GPU-h$5.40–$69+paper precedent was ~192 RTX-3090 GPU-h on larger LLMs
teacher labels for 1,000 h on current 3050~44 GPU-h at measured RTF 0.044local time; cloud likely lowerskip when labels can be produced during data prep on faster GPU

*Price snapshot checked 2026-08-16: RunPod lists A5000 at $0.27/h, L4 at $0.39/h, A40 at $0.44/h, 3090 at $0.50/h, and 4090 at $0.69/h. Storage, idle setup time, availability, and tax are excluded.

Inference cost

The local baseline uses 0.044 GPU-hours per hour of audio. At $0.69/GPU-hour that is an intentionally conservative $0.030 per audio-hour if a rented 4090 were no faster than the laptop 3050; it should be faster in practice. A 2× end-to-end improvement halves compute cost to roughly $0.015/audio-hour under that conservative assumption. Locally, 40 W × 0.044 h is about 1.76Wh of GPU energy per audio-hour, excluding the rest of the laptop.

10What else can be used instead?

pathtrainingextra VRAMquality contractfit for this 3050
native HF + torch.compilenonecompile workspace/cachesame model; check numerical token parityfirst attacks launch-bound decoder; official A100 batch-4 ASR result reports ~2.4×
static KV + CUDA graph/custom loopnonepreallocated cacheexact greedy if implemented correctlyfirst directly removes Python/cache/kernel-launch overhead
fused RMSNorm/residual/RoPE/attention pathnonenegligibletolerance or token parity must be testedhigh relevance because attention + norms + glue are ~57% of decode
INT8/INT4 weight-only runtimecalibration optionallowerWER may moveuseful for fit/bandwidth; less direct on launch-bound attention. RTX 3050 has no native FP8
MTP / Medusa headsyes, frozen backbone possiblesmall to moderatestrict verification can be exact greedygood second research path; ASR evidence includes 1.4–1.5× Whisper-Medusa and 5/6 acceptance in ParaASR
decoder-side CTC drafteryesone small decoder layerstrict tree verification can preserve greedyinteresting, but more engineering than MTP and published gains are on much larger targets
early-exit / LayerSkip self-draftbest with early-exit trainingalmost noneverification can preserve outputmemory-friendly; early Qwen layers may not be acoustically/textually predictive without training
EAGLE-style feature draftyessmall head; shares embedding/headexact speculative sampling supportedstrong general method, but target LM head still costs 2.5ms per draft token
small separate AR ASRdistill/fine-tunelargestrict verify exactpoor under 6GB unless tiny and encoder-sharing; no official smaller Qwen-ASR
token-map / n-gram draftingno neural trainingtiny CPU mapstrict verify exactexcellent cheap baseline for commands/repetitive domains; published 1.27–1.37×
relaxed CTC direct returnCTC headtinyapproximate; thresholded WER tradelargest likely speed, appropriate only behind an explicit fast-mode flag
replace AR with NAR/diffusion/editor ASRsubstantialmodel dependentnew model, not Qwen-equivalentportfolio/research direction, not a minimal Qwen optimization

My ranking for this project

  1. Upgrade and compile the runtime; build the oracle verifier. Cheapest evidence and prerequisite for every speculative path.
  2. Static cache/CUDA graph and targeted fusion. The profile already identifies a launch-bound decoder.
  3. Acoustic CTC linear probe + strict block verification. Best novel research/portfolio angle and lowest draft cost.
  4. MTP/Medusa head as a controlled fallback. Easier token alignment, but repeated LM-head work and smaller expected upside.
  5. Token-map draft. Build early if the real product domain is commands, names, or repetitive templates.
  6. Separate draft model or full NAR replacement. Only after simpler paths fail or the project goal changes.

11An implementation path you can reproduce yourself

Prediction checkpoint: before reading each phase, write down what result would make you stop. The fastest optimization project is the one that kills weak ideas early.

Phase A — prove verification economics without training

  1. Move from the wrapper’s model.generate() call to a minimal greedy loop that exposes prompt KV, one-token decode, cache length, and logits.
  2. Implement verify_block(committed_cache, draft_ids). It performs one target forward over all draft IDs and returns target argmax IDs for those positions plus the speculative cache.
  3. Make the one-token shift explicit: the distribution that verifies draft[i] comes from the preceding causal position, not the logit emitted at the same input position. Preserve the prompt’s next-token logit for the first candidate; the last candidate position supplies the bonus next-token distribution.
  4. Use the baseline transcript itself as an oracle draft. Time block sizes 1, 2, 4, 8, 16, 32, and full output.
  5. Inject mismatches at controlled positions and verify cache cropping/correction against the baseline token sequence.
  6. Stop if even oracle time-per-committed-token cannot beat 18.58 ms by at least 25%; runtime work is then the priority.
@dataclass(frozen=True)
class Verification:
    accepted: tuple[int, ...]
    correction: int | None
    committed_length: int

def accept_greedy(
    draft_ids: tuple[int, ...],
    target_argmax: tuple[int, ...],
) -> Verification:
    """return the exact matching prefix and first target correction."""
    for index, (draft, target) in enumerate(zip(draft_ids, target_argmax)):
        if draft != target:
            return Verification(draft_ids[:index], target, index + 1)
    return Verification(draft_ids, None, len(draft_ids))

Phase B — create the CTC dataset

  1. Select short/long, clean/noisy, music, accents, punctuation, language switches, and repeated-word clips.
  2. Run the frozen baseline once. Save audio ID, duration, language policy, exact raw teacher string, Qwen token IDs, normalized human transcript, and baseline WER.
  3. Split by speaker/source before distillation to avoid leakage.
  4. For bytes, encode the exact teacher UTF-8 bytes; for pruned Qwen IDs, create the training vocabulary from train only and define byte fallback.

Phase C — train the smallest possible probe

  1. Expose audio-tower hidden states without invoking text generation. Cache them to disk if storage permits.
  2. Train LayerNorm + Linear with CTCLoss(blank_id, zero_infinity=True). Assert input_length ≥ target_length; log failures rather than silently dropping languages.
  3. Track CTC WER, exact utterance match, Qwen-token longest-prefix ratio, first-mismatch position, and entropy calibration. Exact-match/prefix metrics predict speed better than WER.
  4. Only if the linear probe misses the acceptance target, add one temporal block. Do not unfreeze Qwen first.

Phase D — integrate strict speculation

  1. Run the audio tower once and retain both projected audio embeddings and CTC draft.
  2. Build prompt KV once. Verify candidate blocks and transactionally commit/crop KV.
  3. Use an immutable state record: committed IDs, target cache length, remaining CTC IDs, audio positions, and EOS status. Every function returns the new state.
  4. Handle EOS, language prefix, repetition processors, maximum length, empty audio, and invalid UTF-8 before optimizing.
  5. Add suffix recycling and confidence-adaptive k only after the fixed-k implementation is token-identical.

Phase E — optimize and optionally relax

  1. Compile the target verifier separately for a small set of static block sizes; capture CUDA graphs if shapes and addresses are stable.
  2. Quantize the CTC head freely, then evaluate target weight-only quantization separately.
  3. Introduce direct-CTC or likelihood acceptance only under an approximate-mode flag with a selected WER budget.

12The benchmark contract

A speedup is credible only when transcript quality, exactness, warmup, and the denominator are fixed.

dimensionminimum report
latencycold and warm wall time; TTFT; encoder/prefill; CTC; draft; verify; fallback AR; CPU overhead
token efficiencyproposed tokens, accepted tokens, committed tokens/round, first mismatch, verification rounds, draft waste
qualityWER/CER by dataset/language/noise; exact transcript match to baseline; punctuation/case-normalized WER separately
memorypeak allocated and reserved VRAM; model/head/cache sizes; OOM boundary by duration
powerGPU power limit and clock state; energy/audio-hour if available
statisticswarmup count, at least 20 timed clips or stable repeated runs, median/p50/p95, tokens and audio duration

Required ablations

  • eager baseline vs native HF baseline vs compile/static baseline
  • oracle draft vs real CTC draft—separates verifier ceiling from model quality
  • byte vs pruned-Qwen vocabulary
  • linear vs one temporal block
  • ground-truth vs teacher-distilled labels
  • fixed block sizes 4/8/16/full vs entropy-adaptive
  • strict equality vs likelihood threshold vs direct CTC gate
  • clean speech vs music/noise vs multilingual/code-switching
success gates

Advance from oracle to training only if 8–16 token verification is at least 1.5× cheaper per committed token than AR. Advance from linear CTC to a temporal head only if the oracle ceiling is good and the probe reaches at least ~50% median useful-prefix coverage. Ship strict CTC only if p95 latency improves and exact baseline-token parity is 100%. Ship approximate mode only with an explicit, accepted WER budget.

13The decision tree

If compile gives ≥1.7×

Keep it as the new baseline, then rerun oracle speculation. The draft must now beat a cheaper target; do not assume the old CTC economics survive.

If oracle blocks give <1.25×

Stop speculative work on this runtime. Focus on CUDA graphs, fusion, native Transformers, or C++/GGML execution.

If oracle is strong but CTC prefix is weak

Try teacher distillation, pruned native tokens, one temporal block, then MTP/Medusa. The verifier is valuable even if CTC is not.

If CTC WER is good but strict prefix is weak

The drafts are semantically right but tokenization/normalization differs. Align teacher text, use target IDs, or add suffix realignment; do not judge by WER alone.

If strict CTC reaches ≥1.5×

Optimize fixed shapes and add adaptive blocks. This is a credible portfolio result because it joins model training, exact decoding, cache transactions, and hardware profiling.

If the user accepts small WER drift

Add calibrated direct-CTC/likelihood gates. Expect the largest latency gain, but publish the entire WER–RTF frontier rather than one flattering point.

final recommendation

Do it—but as the third step, after native compile/runtime work and an oracle verifier. Build the CTC head on shared audio features, start with bytes or a pruned target-token alphabet, train against Qwen’s greedy outputs, and preserve exactness with fixed 8-token block verification. The experiment is small enough to afford, novel enough to teach a lot, and well aligned with the measured hot path. Do not begin by training a separate small ASR model, and do not promise 4× until the oracle and first-mismatch distributions earn it.

Sources and evidence trail

  1. decode-lab: Qwen3-ASR latency & bottleneck analysis — local CUDA-event measurements used for every hardware-specific calculation.
  2. Qwen3-ASR official repository and technical report — architecture, inference modes, model claims, and evaluation setup.
  3. Qwen3-ASR-0.6B native Transformers model card — 1.56GB checkpoint, native Transformers path, and official torch.compile batch-4 A100 result.
  4. Graves et al., Connectionist Temporal Classification (2006) — original CTC formulation.
  5. Leviathan, Kalman & Matias, Fast Inference via Speculative Decoding (ICML 2023) — exact speculative decoding and sampling foundation.
  6. Wen, Gui & Feng, CTC-based Draft Model (NeurIPS 2024) — decoder-side CTC drafter, training recipe, acceptance, speed, and compute precedent.
  7. Saon et al., Self-Speculative Decoding for LLM-based ASR with CTC Encoder Drafts (2026) — whole-transcript acoustic CTC drafting, confidence gates, LLM verification, 4.4×/WER tradeoff.
  8. Wei et al., SpecASR (2025) — adaptive ASR draft length, suffix recycling, sparse trees, and 3.04–3.79× results.
  9. Lin et al., ParaASR (2026) — five MTP branches, 5.0/6 average accepted length, and long-context ASR results.
  10. Segal-Feldman et al., Whisper in Medusa’s Ear (2024) — ASR-specific multi-head prediction and measured speed/WER tradeoffs.
  11. Ho et al., Token Map Drafting (2025) — model-free domain n-gram speculation and 1.27–1.37× CPU ASR speedups.
  12. EAGLE and LayerSkip — feature drafting and early-exit self-speculation alternatives.
  13. RunPod GPU pricing — hourly price snapshot used for the rough dollar ranges.

Research cutoff: 2026-08-16. Published speedups are not directly comparable across GPUs, batch sizes, datasets, models, acceptance rules, and quality budgets. All Qwen/RTX 3050 projections above are explicitly marked as calculations or hypotheses until benchmarked.