Triton from zero · decode-lab

The Triton Tutorial

A ground-up Triton course, taught as self-contained HTML lessons. Written for someone with zero prior knowledge — mental models, analogies, interactive visualizations, and runnable companion scripts (Colab T4 GPU or local CPU interpreter).

Lesson 1 of 7

The Triton Mental Model

Why GPUs exist, what a kernel is, the “write one, run many” shift, tiles, masks, constexpr — and your first vector add.

open lesson →
Lesson 2 of 7

2D Tiles & Matrix Multiply

2D grids, broadcasting, strides, and a complete tiled matmul in ~25 lines — the exact kernel shape decode-lab’s GEMV and attention kernels are built from.

open lesson →
Lesson 3 of 7

Reductions: Softmax & RMSNorm

Block sums, the two-pass softmax dance, and decode-lab’s real fused RMSNorm kernel — the first kernel on the surgery map.

open lesson →
Lesson 4 of 7

Block Pointers

tl.make_block_ptr, tl.advance, and boundary checks — masks for free, and the matmul rewritten in ~10 fewer lines.

open lesson →
Lesson 5 of 7

Fused Kernels

One load, a register chain, one store. The fused SwiGLU MLP — decode-lab’s swiglu.py pattern — and when not to fuse.

open lesson →
Lesson 6 of 7

Performance

num_warps, num_stages, @triton.autotune, and honest benchmarking with do_bench — plus the launch-bound truth about decode.

open lesson →
Lesson 7 of 7 · capstone

Mini FlashAttention

QKᵀ scores, causal mask, row softmax and PV in one kernel — with the online-softmax rescaling trick, the core of decode-lab’s future attention.py.

open lesson →

Each lesson ships with a runnable Python companion — run it on Colab (!pip install -q triton, T4 GPU) or locally via Triton’s CPU interpreter. Old lesson versions are archived under artifacts/archive/.