The loop that turns practice into skill: predict, run, benchmark, explain. Plus the gotchas ledger — every trap this course has hit so far, so you hit them once and once only.
Every item is a gate. Tick them off as you go; the checks survive reloads (saved in your browser). Skip one and the week doesn’t count.
(rows, cols), per launch.-inf for max passes, 0 for sum passes, never the other way).element_ty.do_bench — never time.time().kernel.asm['ptx'] when the numbers lie.num_warps, num_stages. 4–8 configs, key=[...], grid lambda.logs.txt.Grid, ownership, masks, expected output — written down before the launch. When you can do this reliably, you can write kernels.
Memory passes, occupancy, launch tax. The benchmark number should not surprise you. This is the skill that matters in real optimization — nobody cares that it’s correct, they care why it’s 2.1× slower than the theory says.
Strides first, then masks, then dtype accumulation, then PTX. The whistle rule: fast-but-wrong output invalidates everything.
| week | work | gate |
|---|---|---|
| now | Finish L3 homework #2 (minmax) + #3 (layernorm), then L4 (block pointers) | blank-page: softmax, rmsnorm |
| +1 | Port kernels/rmsnorm.py to the T4, bench it, autotune it (L6 applied) | parity + one do_bench number |
| +2 | Write gemv.py for lm_head — the actual decode bottleneck | parity + launch-overhead share computed |
| +3 | L7 kernel → kernels/attention.py (sliding window + GQA + KV-cache slice) | parity vs torch reference |
| +4 | Whistle port (phase 2) — the real deployment | whistle codec + waveform gate |
| trap | the fix / why |
|---|---|
tl.arange bounds must be constexpr | The tray size is printed on the recipe card — compile time only. |
| Runtime loop bounds crash the CPU interpreter | N_ITERS=triton.cdiv(K, BLOCK_K) as a constexpr at launch. Harmless on GPU. |
Masked columns in softmax load as 0 | Must be -inf: exp(-inf - m) = 0. A 0 score polls the row sum — silent, off-by-a-little. |
Padded K columns in attention score 0 | Same rule, score tile: s = tl.where(cols < N, s, -inf). Shipped this bug — max_err 1e-1. |
| Strides guessed, not derived | An (H, I) view of an (I, H) matrix has strides (1, H). Guessing produced NaN in two lessons. |
Block-pointer padding: "zero" or "nan" only | No -inf — softmax still needs manual masks. Know which tool pads what. |
boundary_check isn’t free | Pay only for the raggedness you actually have; require divisible shapes in the hot path. |
| fp16/bf16 accumulation | Compute fp32, cast back with .to(out_ptr.dtype.element_ty) — the whistle rule. |
order=(1, 0) lies about the layout | Order lists the fastest axis first. Wrong order = correct values, terrible perf (or a crash). |
| Autotune’s first call benchmarks | Keep 4–8 configs; the config space explodes (3×3×3×3×3 = 243). Grid becomes a lambda. |
time.time() around a kernel | L2-cache mirages and pipeline noise. do_bench clears L2, warms up, reports a median. |
| Launch overhead ≈ 5–10 µs | At batch-1 decode sizes that’s 30–50% of a kernel’s life. Fuse launches; fusion beats faster kernels. |
Ghost rows: m = -inf → NaN | Pin with m = tl.where(offs_m < M, m, 0.0) — those rows are never stored. |
silu(x) = x * sigmoid(x) · norm = tl.rsqrt | One instruction instead of two; the fused forms are free once the data is loaded. |
| Surprising benchmark numbers | Your assumptions are wrong, not the GPU. Read kernel.asm['ptx'] and count the loads. |
Once a week, one kernel from memory in an empty file. The list grows with the course:
gemv 1×4096→32000, autotuned (lesson 6 homework)Rule: if you can’t write it, the lesson wasn’t absorbed — redo it, don’t move on. This is the test that keeps you LLM-independent.