Your OpenAI fine-tune is being retired. What now?
Some fine-tuned models stop answering on 23 Oct 2026, and no new fine-tuning jobs run after 6 Jan 2027. This is the checklist we use to move a fine-tuned task without losing accuracy, including when the right answer is not to use us.
What is changing
Some fine-tuned models stop
Fine-tuned models built on retired base models stop answering requests. Calls to them return an error, so the feature built on them breaks, not degrades.
No new training
No new fine-tuning jobs. Models that keep running can't be retrained to fix a regression, add a label or follow a change in your data.
Dates from OpenAI's published deprecation notices. Your exact models and dates are on the deprecations page of the OpenAI docs; check it before planning.
Find every fine-tune you depend on
Fine-tuned model IDs start with ft:, so they are easy to find. Look in three places:
- Your code and config: search the repository for
ft:, including environment files and feature flags. - Your usage dashboard: list every model that served traffic in the last 30 days, with call volume.
- Your team: background jobs and internal tools often call a fine-tune nobody remembers.
For each one, write down what it does (classify, route, extract, tag), how many calls it takes a month, how fast it must answer, and whether you still have the training data and a labelled test set.
Three ways to replace it
| Option | Good when | The catch |
|---|---|---|
| A newer API model with a good prompt | The task is easy, volume is low, and a current small model already gets it right | You rent again: the next retirement starts the clock over. Test it; don't assume. |
| Fine-tuning on another hosted platform | You want to keep the fine-tuning workflow and someone else to run it | Still a platform's schedule and pricing. Check what happens to your weights if you leave. |
| A small open model you own | The task is narrow and repeated, and you want it to stop moving under you | Someone has to train, test and serve it. That is the work we do. |
Try the cheap option first. If a current small API model with a prompt passes your test set, use it. A parity check is how you find out, whichever option wins.
Prove the replacement is as good, before you switch
"It looked fine on a few examples" is how migrations go wrong. The test we use:
- Hold out a labelled test set the new model never trained on. A few hundred items is the minimum; over 1,000 gives tight intervals.
- Run the current model and the candidate on the same items.
- Agree a margin up front, for example "no more than 3 accuracy points worse".
- Run a paired bootstrap and a non-inferiority test: the candidate passes only if the whole 95% confidence interval of the difference sits above the margin.
The check is open source and needs no GPU or API key:
pip install git+https://github.com/illate/parity illate-parity check --gold test.jsonl \ --current current.preds.jsonl --candidate candidate.preds.jsonl \ --margin 0.03 --out out/parity
It returns PASS, FAIL or INCONCLUSIVE and writes a report. On our public Banking77 run (3,080 test messages, 77 intents), a fine-tuned 4B open model scored 94.0% against 89.3% for the baseline: full results.
Switch without a bad day
Same interface
Serve the new model behind an OpenAI-compatible endpoint, so the change in your code is the base URL and model name.
→ one-line changeShadow, then stage
Replay real traffic first, then route 1%, 5%, 25% and 100%, comparing answers as you go. Keep the old model as fallback while it still runs.
→ no surprise in productionDefer when unsure
Let the model return "defer" on low confidence and send those cases to a bigger model or a person. On Banking77, a 99% target answered 80% of messages at 99.1%.
→ accuracy you choose
If your model stops on 23 Oct
- This week: inventory, and export your training data and a labelled test set.
- Next week: test the options above against your test set with a parity check.
- The week after: shadow traffic, then a staged switch, with the old model as fallback.
- Before 23 Oct: 100% on the replacement, old model only as a reference.
Want us to run the check for you?
Send a few hundred labelled examples. We test the options on your data and give you a straight answer within a week, free, even if the answer is "a prompt on a newer API model is enough".