When a dedicated GPU beats the API bill, and when it doesn't
Owning a model isn't automatically cheaper. For short classification calls, one always-on L4 pays for itself at about 1.8 million calls a month. Below that, you move for control, not price.
Monthly cost against volume
Hover a point for exact values. The same numbers are in the table at the end of this note.
What goes into the lines
| API side | Value |
|---|---|
| Input tokens per call | 900 |
| Share served from cache | 80% |
| Output tokens per call | 10 |
| Price per 1M tokens (in / cached / out) | $0.80 / $0.20 / $3.20 |
| Cost per 1,000 calls | $0.32 |
| Owned side | Value |
|---|---|
| GPU | 1 × NVIDIA L4 |
| Price | $0.80/hour × 730 h |
| Monthly cost, always on | $584 |
| Sized throughput per replica | 50 req/s |
| Capacity at that rate | ≈ 130M calls/month |
The API side models a support-triage prompt that lists its categories, which is why so much of it caches. The owned side uses the throughput we measured with vLLM on one L4 (86.6 req/s peak), sized down to 50 req/s so latency stays low. One GPU at that rate covers about 130 million calls a month, far more than the break-even volume.
Three regimes
- Under ~1.8M short calls a month. A dedicated GPU costs more than the API. A team at 1.5M calls pays about $480 a month to OpenAI. Moving still makes sense if the model is being retired or the data can't leave your cloud, and shared hosting (several clients' adapters on one GPU) brings the bill under the API price.
- Above ~1.8M calls, or several tasks. One GPU serves many LoRA adapters at once, so every extra task or extra million calls is nearly free. At 10M calls the API line is at $3,200 and the GPU line is still $584.
- Long outputs. Summaries and extraction with hundreds of output tokens make the API line far steeper, because output tokens cost four times input. For those, owning usually wins at much lower volume.
We run this calculation on your real traffic before proposing anything. If a cheaper API model passes your eval for less, the cost check says so.
The chart as a table
| Calls per month | API | One L4 | Cheaper |
|---|---|---|---|
| 1M | $320 | $584 | API |
| 1.5M | $480 | $584 | API |
| 1.83M | $584 | $584 | Break-even |
| 4M | $1,280 | $584 | GPU |
| 10M | $3,200 | $584 | GPU |
GPU cost excludes engineering and monitoring time. API cost excludes retries and the migration you'll need when the model is retired.