A surprising share of an ML project's early life is CPU work: cleaning data, wiring up the pipeline, fixing shapes and off-by-one errors. Doing that on rented GPUs burns money on idle accelerators.
Validate on CPU
- The data pipeline end to end. Load, transform and batch a few hundred examples. Most bugs live here.
- A tiny model that overfits a tiny dataset. If a model with a few thousand parameters can't memorize 50 examples, something in the training loop is wrong.
- Checkpoint save and resume. Kill the run and restart it. You will need this on real hardware.
- Evaluation and logging. Make sure metrics are computed and written correctly before they cost GPU time.
Where CPU validation stops working
Move to a GPU when you need to test real throughput, memory use at full model size, mixed-precision numerics, or multi-GPU communication. These behave differently on accelerators and can only be measured there.
Then start with one GPU, by the hour
Your first GPU session should be short and specific: run the real model on one card, measure step time and memory, and estimate what the full job needs. That estimate is what you size a cluster or a reservation from.



