ILLATE
← Engineering notes

When a dedicated GPU beats the API bill, and when it doesn't

Owning a model isn't automatically cheaper. For short classification calls, one always-on L4 pays for itself at about 1.8 million calls a month. Below that, you move for control, not price.

5 Oct 2026L4 at $0.80/hourgpt-4.1-mini fine-tuned pricing
Break-even

Monthly cost against volume

Short classification calls · USD per month One L4, always onFine-tuned API model
$0 $1k $2k $3k $4k 0 2M 4M 6M 8M 10M calls per month API at $0.32 per 1,000 calls One L4, $584 break-even ≈ 1.8M calls

Hover a point for exact values. The same numbers are in the table at the end of this note.

Inputs

What goes into the lines

API sideValue
Input tokens per call900
Share served from cache80%
Output tokens per call10
Price per 1M tokens (in / cached / out)$0.80 / $0.20 / $3.20
Cost per 1,000 calls$0.32
Owned sideValue
GPU1 × NVIDIA L4
Price$0.80/hour × 730 h
Monthly cost, always on$584
Sized throughput per replica50 req/s
Capacity at that rate≈ 130M calls/month

The API side models a support-triage prompt that lists its categories, which is why so much of it caches. The owned side uses the throughput we measured with vLLM on one L4 (86.6 req/s peak), sized down to 50 req/s so latency stays low. One GPU at that rate covers about 130 million calls a month, far more than the break-even volume.

What it means

Three regimes

  • Under ~1.8M short calls a month. A dedicated GPU costs more than the API. A team at 1.5M calls pays about $480 a month to OpenAI. Moving still makes sense if the model is being retired or the data can't leave your cloud, and shared hosting (several clients' adapters on one GPU) brings the bill under the API price.
  • Above ~1.8M calls, or several tasks. One GPU serves many LoRA adapters at once, so every extra task or extra million calls is nearly free. At 10M calls the API line is at $3,200 and the GPU line is still $584.
  • Long outputs. Summaries and extraction with hundreds of output tokens make the API line far steeper, because output tokens cost four times input. For those, owning usually wins at much lower volume.

We run this calculation on your real traffic before proposing anything. If a cheaper API model passes your eval for less, the cost check says so.

Data

The chart as a table

Calls per monthAPIOne L4Cheaper
1M$320$584API
1.5M$480$584API
1.83M$584$584Break-even
4M$1,280$584GPU
10M$3,200$584GPU

GPU cost excludes engineering and monitoring time. API cost excludes retries and the migration you'll need when the model is retired.