A home for measured experiments, technical breakdowns, and long-form learning artifacts.
Reports, breakdowns and notes on inference, speech, audio and image generation, plus a few digests on how to do research.
Fused rotary and norms, a fixed token layout, per-axis attention and a skipped chunk: 2.42× faster at the original settings. A practical 25 % overlap option reaches 5.93 s (3.74×), with 0.082/0.042 dB mean truth-SDR loss. Baseline, quality gate, before/after traces and an explicit tradeoff review.
A 2.16 B audio model profiled on a single Colab TPU v6e-1 and on an L4, with the same JAX port on both: 405 ms for 120 s of audio on one v6e chip (297× real time), 6.3× faster than the official PyTorch runtime on an L4. The SAME-L decoder costs more than the whole 24-block DiT on TPU, the port beats torch on the same GPU at every duration (3.97× at 5 s, 1.15× at 120 s), and the XProf capture puts the accelerator at ~100 % device-busy during the sampler — the only real idle is ~12 ms of host-side work.
FLUX.2 klein-4B at 1024² on a single TPU v6e-1, from a frozen 944.67 ms to 345.9 ms per image (2.73×, 1.06 → 2.89 img/s, 17.5 % → 47.5 % MFU), every kept step gated against the PyTorch oracle. Resident weights that were silently re-copied every call, a Pallas flash kernel that cut 2 GB-per-call fp32 score buffers, a fused qk-prep kernel where three XLA rewrites only moved the cost — and the measured dead ends: batching, sub-pixel VAE, and int8 that plateaus just short of the quality floor.
Stable Audio 3 medium on a single Colab TPU v6e-1, from 401.4 ms to 163.4 ms, 2.46× faster and ahead of TensorRT fp8 on an RTX PRO 6000. Quality gated against official PyTorch; the full optimization log with teaching notes lives in an appendix.
The first TPU latency, MFU and residency numbers for FLUX.2 klein-4B: 0 FAIL across all 16 parity rungs on a third backend, 262 ms per 512² image with everything resident on a single 32 GiB v6e chip (3.8 images/s, 34 % of HBM still free), and 35.5 s per request if the chip has to reload the weights instead. Plus the finding that cost a day: an fp32 matmul on TPU is silently bf16 unless you raise the precision — and the text encoder, at 3 % of peak MFU, eats 40 % of the wall clock for 6 % of the FLOPs.
Per-module FLOPs, bytes moved, arithmetic intensity and collective cost for two generative inference workloads at their real shapes. Ridge point 560 FLOP/byte; FSDP-8 costs more than the entire math floor while Ulysses is nearly free; and the analytical model underestimates the measured L4 by 2.4-2.7× on the DiT and 8.1× on the text encoder, decomposed into m=1 prologues and small-shape GEMM inefficiency. Eleven tables, every row printed by one script.
Testing an NNX checkpointing library I wrote and never tested: the committed version could not save a single model on current jax/flax/safetensors, an nnx.List stack loaded with random weights and no error — and how streaming reads plus an abstract builder cut a load from 3× the checkpoint in host RAM to 1.1×, with shard-per-process paths for a TPU pod.
A 228 M-parameter axial transformer over a 60-band mel-split STFT, costed on a 6 GB laptop GPU — and why flash_attn=True silently lands on the slowest available attention kernel.
Pick questions worth answering, test claims, design around real constraints, and act without waiting to be asked. Includes notes on taste in inference engineering and a section on agency.
Why so few scientists do significant work, the 10-20 problems list, Friday "great thoughts", working with doors open, and the courage to tackle essential questions.
The 70-year meta-pattern in AI: general methods that leverage search and learning scaled with compute inevitably crush human-engineered heuristics.
The craft of problem discovery, building tools for thought, Anki for deep retention, networked science, and navigating the research valley of despair.
Six mental models for invention: radical simplification, structural isomorphism, state-space inversion, generalization, and playful tinkering.
The cost model, the full catalog of post-training families that buy serving speed (distillation, MTP/draft heads, QAD, architecture surgery, step distillation), the non-training engineering map across LLMs/diffusion/audio/VLMs, and the 2026 production numbers from fal and Baseten that prove the recipe end to end.
A digest of @gabriel1's tweets on effective job search: proof-of-work demos instead of resumes, direct outreach past HR, risk-reversal trial offers, interviews you control — and the career-as- decisions mindset behind it all.
Full codebase & systems breakdown of Resemble AI's Chatterbox suite (Base TTS, Multilingual V2/V3, Turbo 350M, Nano 110M, Voice Conversion): the T3 autoregressive LLM, S3Gen flow matching, MeanFlow distillation, iSTFTNet neural source filter vocoding, empirical benchmarks, and the optimization roadmap.
Full repo walkthrough: the native Qwen3-VL port, static-KV compiled decode, int4 sidecar quantization, three schedulers, the 3.49×/9.35× optimization ladder — and what it would take to swap Qwen3-VL for GLM-OCR.
Analogy → plain English → pseudocode, grounded in the paper (arXiv 2607.05147) and the DeepSpec reference code: the parallel backbone + Markov/RNN sequential head, the confidence head with sequential-temperature calibration, the hardware-aware prefix scheduler (with the Appendix A losslessness counterexample), training losses, the inference loop, results — and a minimal plan to train a DSpark drafter for a local VLM.
A 12-week plan for speech and multimodal inference engineering on constrained hardware: case studies, evidence inventory, and job search operating habits.
Exact greedy and sampling algorithms, draft contracts, acceptance math, cache state, and applications across LLMs, VLMs, ASR, audio, images, and actions.
A stage-by-stage and layer-by-layer account of a 60-second ASR run, with the measured hot path and an optimization roadmap.
A hardware-specific feasibility study of CTC drafting, exact speculation, alternatives, and an oracle-first implementation plan.
The SAME autoencoder, the diffusion transformer and the 8-step sampling loop, costed against a 6 GB laptop GPU — plus why a sub-3 GB model can still OOM, and how these models are served at scale.
A frontier open music model that will not fit a 6 GB GPU: verified parameter inventory, the mixture-of-transformers and flow-matching inference path, and why the blocker is memory rather than speed.
A practical checklist, deliberate-practice loop, LLM-independence bar, and gotchas ledger for reading and writing performance kernels.