lab

Kernel Notes & Checklist

The loop that turns practice into skill: predict, run, benchmark, explain. Plus the gotchas ledger — every trap this course has hit so far, so you hit them once and once only.

1The kernel checklist — one kernel per week

Every item is a gate. Tick them off as you go; the checks survive reloads (saved in your browser). Skip one and the week doesn’t count.

progress: 0/12

2The three skills — the LLM-independence bar

1 · Predict correctness before running

Grid, ownership, masks, expected output — written down before the launch. When you can do this reliably, you can write kernels.

2 · Predict performance before benchmarking

Memory passes, occupancy, launch tax. The benchmark number should not surprise you. This is the skill that matters in real optimization — nobody cares that it’s correct, they care why it’s 2.1× slower than the theory says.

3 · Debug when numbers lie

Strides first, then masks, then dtype accumulation, then PTX. The whistle rule: fast-but-wrong output invalidates everything.

The LLM policy

3The plan — fast without skipping

weekworkgate
nowFinish L3 homework #2 (minmax) + #3 (layernorm), then L4 (block pointers)blank-page: softmax, rmsnorm
+1Port kernels/rmsnorm.py to the T4, bench it, autotune it (L6 applied)parity + one do_bench number
+2Write gemv.py for lm_head — the actual decode bottleneckparity + launch-overhead share computed
+3L7 kernel → kernels/attention.py (sliding window + GQA + KV-cache slice)parity vs torch reference
+4Whistle port (phase 2) — the real deploymentwhistle codec + waveform gate

4The gotchas ledger — every trap so far

trapthe fix / why
tl.arange bounds must be constexprThe tray size is printed on the recipe card — compile time only.
Runtime loop bounds crash the CPU interpreterN_ITERS=triton.cdiv(K, BLOCK_K) as a constexpr at launch. Harmless on GPU.
Masked columns in softmax load as 0Must be -inf: exp(-inf - m) = 0. A 0 score polls the row sum — silent, off-by-a-little.
Padded K columns in attention score 0Same rule, score tile: s = tl.where(cols < N, s, -inf). Shipped this bug — max_err 1e-1.
Strides guessed, not derivedAn (H, I) view of an (I, H) matrix has strides (1, H). Guessing produced NaN in two lessons.
Block-pointer padding: "zero" or "nan" onlyNo -inf — softmax still needs manual masks. Know which tool pads what.
boundary_check isn’t freePay only for the raggedness you actually have; require divisible shapes in the hot path.
fp16/bf16 accumulationCompute fp32, cast back with .to(out_ptr.dtype.element_ty) — the whistle rule.
order=(1, 0) lies about the layoutOrder lists the fastest axis first. Wrong order = correct values, terrible perf (or a crash).
Autotune’s first call benchmarksKeep 4–8 configs; the config space explodes (3×3×3×3×3 = 243). Grid becomes a lambda.
time.time() around a kernelL2-cache mirages and pipeline noise. do_bench clears L2, warms up, reports a median.
Launch overhead ≈ 5–10 µsAt batch-1 decode sizes that’s 30–50% of a kernel’s life. Fuse launches; fusion beats faster kernels.
Ghost rows: m = -inf → NaNPin with m = tl.where(offs_m < M, m, 0.0) — those rows are never stored.
silu(x) = x * sigmoid(x) · norm = tl.rsqrtOne instruction instead of two; the fused forms are free once the data is loaded.
Surprising benchmark numbersYour assumptions are wrong, not the GPU. Read kernel.asm['ptx'] and count the loads.

5Blank-page tests — the mastery ladder

Once a week, one kernel from memory in an empty file. The list grows with the course:

Current ladder (L1–L4)
  • ✔ vector add (lesson 1) — done
  • ▸ per-row softmax, masked (lesson 3)
  • ▸ fused rmsnorm + residual (lesson 3)
  • ▸ minmax normalize (lesson 3 homework)
  • ▸ tiled matmul with block pointers (lesson 4)
  • ▸ gemv 1×4096→32000, autotuned (lesson 6 homework)

Rule: if you can’t write it, the lesson wasn’t absorbed — redo it, don’t move on. This is the test that keeps you LLM-independent.