STRATA Reserve capacity
Strata/Blog

Measure before you scale: find your real bottleneck

Engineering3 Oct 20266 min read

Adding GPUs to an I/O-bound job buys you a more expensive version of the same problem. Four measurements tell you which axis to scale.

When a proof of concept works, the instinct is to throw more hardware at it. Sometimes that is right. Often the job is waiting on something other than the GPU, and more GPUs just wait in parallel.

1. GPU utilization over time, not a single snapshot

Run nvidia-smi dmon or your monitoring stack during a real run and look at the curve. Sustained utilization above 90% means the GPU is the bottleneck. A saw-tooth pattern that drops to near zero between steps usually means the GPU is starving for data.

2. Data loading time per step

Time how long your data loader takes to produce a batch compared with the forward and backward pass. If loading takes a third of the step, more GPUs will not help; more CPU workers, faster storage, or pre-processed data in a better format will.

3. Memory headroom

If memory is almost full and the batch size is tiny, you are memory-bound. Options before buying a bigger card: mixed precision, gradient checkpointing, a smaller optimizer state, or quantization for inference.

4. Communication share in multi-GPU runs

Profile how much of each step is spent in all-reduce or other collective operations. If it grows quickly as you add GPUs, you are network-bound and need faster interconnect, a larger per-GPU batch, or a different parallelism strategy.

Scale the axis that is actually full

SymptomBottleneckWhat helps
GPU at 95%+ all the timeComputeNewer or more GPUs
GPU utilization saw-toothInput pipelineCPU workers, storage, caching
Out-of-memory at small batchMemoryPrecision, checkpointing, bigger VRAM
Step time grows with node countNetworkInterconnect, parallelism strategy

An hour of measurement on a single rented GPU is the cheapest capacity planning you will ever do.

Need GPUs for this?Reserve capacity
More from the blog