How to build research & engineering taste
Taste is knowing what is worth doing before you have proof. Agency is doing it without waiting to be asked. This guide is about growing both, with notes on how they show up in inference engineering.
Updated
What taste is
Skill is being able to carry out a plan. Taste is choosing the plan, and noticing early when it has gone wrong. A useful way to think about it: taste is a compressed model of consequences. You have seen enough decisions play out that you can guess how a new one will go before you pay for it.
That model comes from a specific kind of experience. You make a call, you see what happened, and you work out why. Years of work where you never check your predictions build confidence, not taste.
Research taste asks two things: does this question matter, and can an experiment actually settle it? Engineering taste asks whether a design meets its constraints and stays understandable when it changes or breaks. Both are about reasoning under uncertainty. They differ in what counts as winning. A research prototype can be thrown away, but the evidence it produces has to hold up. A production system can use a boring method and still be excellent.
Novelty, elegance, and scale are not goals. They are worth something only when they serve the problem.
Pick problems you can learn from
Before spending months on something, write one paragraph that answers these four questions. If you can't, you are not ready to commit.
- What changes if it works? Name the capability, the explanation, or the practical improvement. A 3% gain can matter a lot at scale. A big benchmark gain can mean nothing for the use you have in mind.
- What don't you know? Separate "we don't understand the mechanism" from "we haven't built it yet". For the first, write down the competing explanations and the observation that would tell them apart.
- Why you? What access do you have (data, users, hardware, tools, expertise, or a failure others overlooked)? Pick a question your resources can actually test.
- What is the cheapest test that would decide it? Set a budget, a baseline, and a result that would make you stop or turn. A negative result only means something if the experiment could have detected the effect you were looking for.
Hamming's advice applies here (lab digest): an important problem is one that matters and has a way in. If the big problem has no way in yet, a tractable piece of it is often the right place to start.
Contrarian ideas need evidence, not just conviction. Write down why sensible people reject the idea and what you know that might change their minds. Often a familiar problem with neglected users or unusual constraints (a small GPU, a TPU, a low-resource language) has more room for discovery than a fashionable bet against the consensus.
Also ask how the problem ages. Sutton's bitter lesson (lab digest) is a warning against clever tricks that more compute will make obsolete. A good test: will this still matter when the model is 10× bigger, or the hardware is 3× faster?
Design around the real constraints
Start from the workload, the correctness requirements, the latency and memory budgets, and the people who will maintain the thing. Design principles help you choose between tradeoffs. None of them lets you skip measuring.
Put complexity only where it pays
Delete machinery nobody uses. Add an abstraction when it names a stable concept and reduces what callers need to know. If every new use needs another flag or special case, the boundary is in the wrong place. Sandi Metz's case for tolerating duplication is a warning against forcing unlike things together. It is not a rule that you must wait for three uses.
Find the bottleneck before fixing it
Estimate how much computation and data movement the job needs, then profile a representative workload. Memory bandwidth, compute, communication, allocation, or launch overhead can each dominate depending on batch size and sequence length. A faster kernel helps only if it moves the metric you care about.
Make ownership and failure explicit
Write down who owns each buffer, which shapes and dtypes an interface accepts, and what happens on cancellation or partial failure. Catch invalid computation before it corrupts results. For failures you can recover from, pick bounded retries, isolation, or a visible degraded mode based on what the service needs. Crashing the whole process is not a recovery strategy for everything.
Make the next decision easier
Expose enough measurements to explain slow or wrong behaviour. Keep configuration tied to choices people actually make. Take on a dependency when the work it saves and its reliability outweigh the cost of integrating and maintaining it. The real test of a design is how safely someone else can change it.
Study the mechanism, not the name
Each of these examples links an algorithmic idea to a systems constraint. The speedups they report belong to particular hardware, workloads, and baselines. What transfers is the reasoning.
FlashAttention: move less data
FlashAttention (2022) computes exact attention in tiles, so it never writes the full attention matrix out to GPU high-bandwidth memory. It cuts memory traffic. It does not change the arithmetic, which is still quadratic in sequence length. "Exact" means it is not an approximation of attention. Rounding can still differ slightly.
FlashAttention-3 (2024) shows why you should re-check your assumptions when the hardware changes. It relies on Hopper's asynchronous execution and low-precision units. The habit to keep is looking at the memory hierarchy and the execution schedule again on every new chip.
PagedAttention: allocate memory like an OS
PagedAttention and vLLM (2023) borrow virtual-memory paging for the growing key-value cache. This cuts fragmentation and lets requests share cache blocks. The paper reports 2–4× serving throughput over the baselines it tested, at similar latency. That is a comparison against the systems of the time, not a promise of what turning on paging will do for you today.
Speculative decoding: spend compute to save time
Speculative decoding (2023) has a cheap model propose several tokens, which the target model checks in one parallel pass. The accept/reject rule keeps the output distribution the same as the target model's, given the algorithm's assumptions. Whether it helps depends on acceptance rate, draft cost, verification cost, and load. At high batch sizes the GPU is already busy, and the extra work can erase the gain.
For any technique, ask three things: what resource was scarce, what changed, and when would this stop helping? That answer is worth more than knowing the technique's name.
Read papers and code to test a claim
Keep two questions apart: how important is the question, and how strong is the evidence? A promising idea can rest on weak support. Careful work can be valuable just by establishing a limit, without proposing anything new.
For each paper, write down:
- Claim: What exactly got better, and under which conditions?
- Comparison: Were the baselines strong, with similar data, tuning, compute, and evaluation budgets?
- Evidence: Do the ablations isolate the proposed cause? Could leakage, cherry-picked samples, or noise explain the result?
- Limits: Where does the method fail, and what would you need to reproduce before relying on it?
Then follow one result through the code: data loading, preprocessing, model execution, metric calculation. Check masks, padding, precision, synchronization, and evaluation settings. Record the commit, dependencies, data version, and hardware, since you need them to interpret any reproduction.
Look at individual errors, not just averages. In speech recognition, a better average word error rate can hide a regression on one language or one recording condition. In serving, a good average latency can hide the requests that miss their deadline. Pick slices and percentiles that match the real use, with enough samples to trust them.
Taste in inference engineering
Inference work is a good place to train taste, because the feedback is fast and numeric. You can predict a latency, run the code, and find out if you were right within minutes. It also has its own traps. These are the habits worth building.
Estimate before you measure
Know the physical limits of your hardware and do the arithmetic before profiling. Two estimates cover most cases.
- Decode floor at batch 1. Each generated token reads every weight once. So the minimum time per token is roughly
weight bytes ÷ memory bandwidth. An 8B-parameter model in bf16 is about 16 GB. On a GPU with ~3.35 TB/s of HBM bandwidth that gives ~4.8 ms per token, or a ceiling of ~200 tokens/s for a single stream. If you measure 40 tokens/s, the gap is overhead you can find. If you measure 190, stop tuning kernels and change the problem (quantize, batch, or speculate). - KV cache size. Per token:
2 × layers × kv_heads × head_dim × bytes per element. Multiply by context length and concurrent requests. This number, not the weights, often decides how many users fit on one device.
The roofline model ties these together. Divide peak FLOPs by memory bandwidth to get the chip's ridge point in FLOPs per byte. An operation whose arithmetic intensity falls below that point is memory-bound, and above it compute-bound. Decode at small batch sits far below the ridge. Prefill and large-batch decode move toward it. That single picture explains why batching, quantization, and fusion help where they do.
Measure the metric users feel
"Faster" is not a metric. Name one: time to first token, time per output token, end-to-end latency at p50 and p99, throughput at a latency target (goodput), peak memory, or cost per million tokens. These pull against each other. Batching improves throughput and hurts per-request latency. Always report the conditions with the number: hardware, dtype, batch size, input and output lengths, and whether compilation and warmup were excluded.
Distrust your timer
GPU and TPU work runs asynchronously. A Python timer without a device sync measures how fast you queued work, not how fast it ran. The first call often includes compilation, autotuning, or memory allocation. JAX and torch.compile recompile when shapes change, so a variable-length workload can look fine in a benchmark and recompile constantly in production. Warm up, synchronize, use real input shapes, and read a profiler trace before believing any number.
Check what actually runs
A flag is a request, not a guarantee. Libraries fall back silently when a dtype, shape, or head dimension isn't supported. In the Mel-Band RoFormer report, setting flash_attn=True ended up on the slowest available attention path. The trace shows which kernel ran. The config file only shows which one you asked for.
Use Amdahl's law to decide what to skip
Break the request into stages (text encoder, main model, decoder, host-side work) and measure each one. A 3× speedup on a stage that takes 10% of the time saves under 7%. The klein-4B TPU report (944 → 346 ms) and the Stable Audio 3 medium report (401 → 163 ms) are both stage-by-stage stories for this reason. Also know when to stop. Once you are near the hardware floor, the next gain has to come from changing the work itself, not from doing the same work faster.
Treat quality as part of the result
Quantization, reduced precision, fewer diffusion steps, and speculative drafts all trade something. Decide beforehand how you will compare outputs: bitwise equality, a numerical tolerance, a task metric, or a perceptual metric such as LPIPS for images. Check that metric on real prompts, including hard ones, not only on the benchmark average. A speedup that quietly damages the long tail of outputs is a bug.
Compare against a strong baseline
Beating a naive generate() loop proves little. Compare against a well-configured serving engine (vLLM, SGLang, TensorRT-LLM, or the platform's own stack) on the same hardware and workload. If you can't beat it, that is useful too: use it, and spend your effort where it can't reach.
Match the optimization to the regime
| If the profile shows | Usually try | Usually won't help |
|---|---|---|
| Memory-bound decode, small batch | Weight quantization, batching, speculative decoding | Faster matmul kernels |
| Compute-bound prefill or diffusion steps | Better kernels, lower-precision matmuls, fewer steps, smaller shapes | More batching |
| Gaps between many small kernels | Fusion, CUDA graphs, compilation, moving host work off the hot path | Quantization |
| Out of memory at target concurrency | KV cache quantization, paging, prefix sharing, shorter contexts | Faster kernels |
For a longer treatment see the inference-efficiency playbook.
Agency
Agency is the habit of treating an outcome as yours to produce, rather than waiting for permission, instructions, or ideal conditions. Taste tells you what is worth doing. Agency is what gets it done. It is also how taste grows, because every action you take and check is another data point for your judgment.
What it looks like
- Turn vague goals into a next action. "Make the model faster" becomes "profile one request on the target GPU this afternoon". When you are stuck, the next step is usually too big.
- Go to the source. When a library behaves strangely, read its code, open the trace, or write a ten-line reproduction. Don't wait for someone to explain it. Most "mysterious" behaviour is a few files away.
- Build the tool you're missing. If you keep squinting at the same numbers, write the benchmark harness, the dashboard, or the comparison script. Good tools make the next hundred decisions cheaper.
- Ask early, and ask specifically. Agency is not doing everything alone. "I tried X and Y, saw Z, and think the cause is W. Am I missing something?" gets a useful answer in minutes. "It doesn't work" gets nothing.
- Finish. The last 10% — edge cases, docs, the write-up, upstreaming the fix — is where much of the value is and where most people stop. Unpublished results mostly don't exist for anyone else.
- Own the failure too. When something breaks, say so plainly, find the cause, and fix the process that let it through.
Agency with judgment
Agency is not recklessness. Sort decisions by how hard they are to undo. Reversible ones (a branch, an experiment, a prototype) should be made quickly and alone. Irreversible or shared ones (a production deploy, a public API, deleting data, someone else's code) deserve a check with the people affected. High agency means you move fast on the first kind and communicate well on the second. It does not mean ignoring other people's ownership.
Agency without taste is thrashing: lots of activity aimed at the wrong target. If you are shipping a lot and little of it matters, slow down and go back to problem choice.
In inference work
The high-agency move is usually to go one layer down. If a serving framework is slow for your model, read the scheduler. If a kernel falls back, find the condition and fix or report it. If a model has no port for your hardware (a TPU, a small laptop GPU), port it and write down what you learned. Several reports in this lab exist because the default answer was "that isn't supported" and it turned out to be a few days of work.
Training it
Once a week, ask: what am I waiting on, and what would I do if nobody was going to give me an answer? Then do the smallest version of that. Keep a short list of things you complained about and never acted on. Pick one and fix it.
Train your judgment with short feedback loops
Pick one small decision each week. Write down a prediction before you test it, compare alternatives, and note what changed your mind. Reimplement a component when doing so exposes an assumption you need to understand. A full rewrite is rarely necessary.
A worked example: slow local inference
Say a model is too slow on your laptop. First measure time to first token, decode speed, and peak memory on prompts you actually use. Keep model loading separate from steady-state running.
Now predict. Work out the decode floor from the weight size and your memory bandwidth. If you are far above it, the problem is overhead: find it in a trace. If you are close to it and suspect weight traffic, predict what quantization should do (roughly halving bytes should roughly halve time per token), then compare against the same model and workload at higher precision. Measure output quality as well as speed and memory. If latency barely moves, open the profile: an unsupported kernel, a dequantization cost, or a different bottleneck may explain it.
The research question is whether your explanation survived the experiment. The engineering question is whether the gain is worth the quality loss and the work.
Keep a decision log
| Before the test | After the test |
|---|---|
| Claim, alternatives, and how confident you are | Result, uncertainty, and anything surprising |
| Workload, baseline, budget, and success threshold | Measured gain, costs, and where it failed |
| Evidence that would change your mind | What you changed, and when to look again |
Review the log each month and look for repeated mistakes: weak baselines, early abstractions, optimistic timelines, or treating a microbenchmark win as a system win. Never edit old predictions. For long-range forecasts, pick an observable outcome and a date. "This direction will win" is too vague to score.
Before your next big commitment, answer three questions: Why this problem? Why this design? What would change my mind? If the answers are fuzzy, make the next experiment smaller.
Further reading
On taste and research
- Richard Hamming, You and Your Research (1986) — important problems need a way in. Lab digest.
- Chris Olah, Research Taste Exercises — concrete exercises for training taste deliberately.
- John Schulman, An Opinionated Guide to ML Research (2020) — choosing problems, keeping a notebook, goal-driven versus idea-driven work.
- Paul Graham, How to Do Great Work (2023) — curiosity, choosing a direction, and acting on it.
- Lab digests: Nielsen, Principles of Effective Research · Shannon, Creative Thinking · Sutton, The Bitter Lesson.
On engineering
- Sandi Metz, The Wrong Abstraction (2016) — when shared code forces incompatible needs together.
- John Ousterhout, A Philosophy of Software Design, 2nd ed. (2021) — reducing complexity by deciding what an interface exposes.
- Andrej Karpathy, A Recipe for Training Neural Networks (2019) — patience, visualization, and building up from a trivial baseline.
On inference
- Horace He, Making Deep Learning Go Brrrr From First Principles (2022) — compute, memory bandwidth, and overhead as the three regimes.
- Williams, Waterman & Patterson, Roofline (2009) — the model behind most back-of-the-envelope performance estimates.
- In this lab: the inference-efficiency playbook.