ILLATE
← Engineering notes

Your OpenAI fine-tune is being retired. What now?

Some fine-tuned models stop answering on 23 Oct 2026, and no new fine-tuning jobs run after 6 Jan 2027. This is the checklist we use to move a fine-tuned task without losing accuracy, including when the right answer is not to use us.

6 Oct 202610-minute readOpen-source parity check
01 Dates

What is changing

23 Oct 2026

Some fine-tuned models stop

Fine-tuned models built on retired base models stop answering requests. Calls to them return an error, so the feature built on them breaks, not degrades.

6 Jan 2027

No new training

No new fine-tuning jobs. Models that keep running can't be retrained to fix a regression, add a label or follow a change in your data.

Dates from OpenAI's published deprecation notices. Your exact models and dates are on the deprecations page of the OpenAI docs; check it before planning.

02 Inventory

Find every fine-tune you depend on

Fine-tuned model IDs start with ft:, so they are easy to find. Look in three places:

  • Your code and config: search the repository for ft:, including environment files and feature flags.
  • Your usage dashboard: list every model that served traffic in the last 30 days, with call volume.
  • Your team: background jobs and internal tools often call a fine-tune nobody remembers.

For each one, write down what it does (classify, route, extract, tag), how many calls it takes a month, how fast it must answer, and whether you still have the training data and a labelled test set.

03 Options

Three ways to replace it

OptionGood whenThe catch
A newer API model with a good promptThe task is easy, volume is low, and a current small model already gets it rightYou rent again: the next retirement starts the clock over. Test it; don't assume.
Fine-tuning on another hosted platformYou want to keep the fine-tuning workflow and someone else to run itStill a platform's schedule and pricing. Check what happens to your weights if you leave.
A small open model you ownThe task is narrow and repeated, and you want it to stop moving under youSomeone has to train, test and serve it. That is the work we do.

Try the cheap option first. If a current small API model with a prompt passes your test set, use it. A parity check is how you find out, whichever option wins.

04 Proof

Prove the replacement is as good, before you switch

"It looked fine on a few examples" is how migrations go wrong. The test we use:

  1. Hold out a labelled test set the new model never trained on. A few hundred items is the minimum; over 1,000 gives tight intervals.
  2. Run the current model and the candidate on the same items.
  3. Agree a margin up front, for example "no more than 3 accuracy points worse".
  4. Run a paired bootstrap and a non-inferiority test: the candidate passes only if the whole 95% confidence interval of the difference sits above the margin.

The check is open source and needs no GPU or API key:

pip install git+https://github.com/illate/parity

illate-parity check --gold test.jsonl \
  --current current.preds.jsonl --candidate candidate.preds.jsonl \
  --margin 0.03 --out out/parity

It returns PASS, FAIL or INCONCLUSIVE and writes a report. On our public Banking77 run (3,080 test messages, 77 intents), a fine-tuned 4B open model scored 94.0% against 89.3% for the baseline: full results.

05 Switch

Switch without a bad day

  1. Same interface

    Serve the new model behind an OpenAI-compatible endpoint, so the change in your code is the base URL and model name.

    → one-line change
  2. Shadow, then stage

    Replay real traffic first, then route 1%, 5%, 25% and 100%, comparing answers as you go. Keep the old model as fallback while it still runs.

    → no surprise in production
  3. Defer when unsure

    Let the model return "defer" on low confidence and send those cases to a bigger model or a person. On Banking77, a 99% target answered 80% of messages at 99.1%.

    → accuracy you choose
06 Timeline

If your model stops on 23 Oct

Want us to run the check for you?

Send a few hundred labelled examples. We test the options on your data and give you a straight answer within a week, free, even if the answer is "a prompt on a newer API model is enough".