Why does changing batch size speed up a model? Why can a program struggle before memory is full? Why can throughput improve while users wait longer? These questions appear when models meet software and hardware.
Efficiency and systems research asks how to train and serve models within hardware, time, and cost limits. An unchanged formula can run at very different speeds: data movement, kernel implementation, memory allocation, request scheduling, and communication determine whether GPUs stay busy and how long users wait.
This field is closely tied to model design. Longer contexts increase cache requirements, concurrent requests change memory pressure, and lower precision may affect quality. Read an optimization together with its workload: which memory does it save, which wait does it reduce, and how would another task change the result?
- Distribute model and training states that do not fit on one device.
- Reduce data movement and redundant work to use existing hardware more effectively.
- Measure tradeoffs among model quality, response time, concurrency, and service cost.
Measure the bottleneck before choosing an optimization. Lower memory use, higher throughput, and shorter user waits are different outcomes.
- 01Weight quantization
- 02KV cache
- 03Continuous batching
- 04Pipeline parallelism
Research milestones
- 2020
ZeRO: partition redundant training state
Shard optimizer states, gradients, and parameters that would otherwise be replicated in data-parallel training.
- 2022
FlashAttention: change data movement
Preserve exact attention while tiling computation to reduce transfers between GPU memory levels.
- 2023
SmoothQuant: make low precision practical
Handle activation outliers to support INT8 weights and activations during inference.
- 2023
PagedAttention: manage a growing KV cache
Organize caches in blocks to reduce fragmentation and duplication and support more concurrent requests.
Key concepts
- Throughput and latency
- Throughput measures work completed per unit time; latency measures how long one request waits. Larger batches can improve throughput while increasing queues.
- Prefill / decode
- Prefill processes the input context; decode generates output step by step. Their compute and memory-access patterns differ, so one speed number is insufficient.
- KV cache
- Stored attention keys and values for processed tokens avoid repeated computation during generation. Context length and concurrent requests consume cache memory.
- Quantization
- Lower-precision representations reduce storage and some computation costs. Speedups depend on hardware and kernel support.
Compute, memory access and execution overhead
Read Horace He's Making Deep Learning go Brrrr for compute, memory, and execution overhead. Record inputs, hardware, execution mode, and time for a small program, then investigate its bottleneck.
Use Machine Learning Systems for a broader view. For quantization, compression, and efficient computing, see MIT 6.5940, 2024.
Latency, throughput, memory and quality
Interactive services may prioritize time to first token and tail latency; offline batches may prioritize throughput. Memory and cost also constrain deployment. Write these conditions down and retain an evaluation of task quality.
Measure different input lengths or batch sizes, then consult the relevant PyTorch Performance Tuning Guide advice. Remeasure under the same workload after a change.
Distributed training and inference
How To Scale Your Model introduces distributed computation and hardware constraints. With a concrete bottleneck, follow communication, parallelism, memory, and compute.
This also tests your use of AI. It can write benchmark scripts, but you should explain what the test measures and whom the optimization affects. See working with AI.
System case studies, courses and research teams
Compare Purshow's systems notes, the systems catalog, and teams such as MIT HAN Lab to connect model design and deployment constraints.
Representative papers
These papers cover training state, kernel memory traffic, numerical precision, and serving caches. Identify each bottleneck before reading performance plots.