STRATA Reserve capacity
Strata/Blog

Notes on GPUs, pricing and running AI in production.

Guides · 5 Oct 2026

How to choose a GPU for your AI workload

H100, H200, B200 or an RTX 5090? The right card depends less on benchmarks and more on three numbers: model size, batch size and how long the job runs.

Read · 7 min →
Pricing · 4 Oct 2026

On-demand or reserved GPUs: when to commit

On-demand capacity is flexible and expensive per hour. Reserved capacity is cheaper and inflexible. The break-even depends on one number most teams never calculate.

Read · 6 min →
Engineering · 3 Oct 2026

Measure before you scale: find your real bottleneck

Adding GPUs to an I/O-bound job buys you a more expensive version of the same problem. Four measurements tell you which axis to scale.

Read · 6 min →
Engineering · 2 Oct 2026

Start without a GPU: what to validate first

The most expensive GPU hours are the ones spent finding bugs a laptop could have caught. Here is what to check before you rent anything.

Read · 5 min →
Inference · 1 Oct 2026

Getting more inference out of every GPU

Most production GPUs serving language models run far below their capacity. Five techniques close the gap without changing the model.

Read · 7 min →
Market · 30 Sep 2026

GPU rental prices in 2026: what moved and why

H100 rates fell to $1.70 an hour in late 2025 and climbed back above $3 by September 2026. What that swing means for anyone buying compute.

Read · 6 min →
Engineering · 29 Sep 2026

GPU optimization in Kubernetes: more work from every card

Kubernetes hands out GPUs one whole card at a time. Most workloads use a fraction of it. Sharing, right-sizing and autoscaling close the gap.

Read · 7 min →
Inference · 27 Sep 2026

Idle GPU capacity: the cost nobody sees

The gap between the GPU hours you pay for and the GPU work you get is often bigger than any price difference between providers.

Read · 6 min →
Pricing · 25 Sep 2026

Put a human where GPU spend changes

GPU cost overruns rarely come from one bad decision. They come from spend that grows by default. A few approval points fix that without slowing teams down.

Read · 5 min →
Engineering · 23 Sep 2026

Running GPU workloads on Kubernetes: what it does and what it doesn't

Kubernetes places GPU jobs. It does not tell you whether they needed that GPU. Knowing the difference saves money and debugging time.

Read · 6 min →