ILLATE
For teams running fine-tuned OpenAI models

OpenAI is shutting down fine-tuning. Keep your model.

We move your fine-tuned model to an open model you own. You see it match on your own data before you pay anything, and switching takes one line of code.

parity-report.md3,080 held-out items
PASS · within the agreed 3-point margin
ModelAccuracy95% CI
ft:gpt-4.1-mini:acme::triageretiring–
qwen3-4b + your LoRA94.0%93.1–94.8
free baseline89.3%88.1–90.3

# Banking77 support triage · 77 categories · 0 invalid answers

Why now

The deadline is OpenAI's, not ours

23 Oct 2026Some models stop

Fine-tuned models built on retired base models stop answering requests.

6 Jan 2027No new training

OpenAI accepts no new fine-tuning jobs, so you can't retrain to fix a regression.

With ILLATENothing expires

Open weights can't be deprecated. Your model runs for as long as you want it to.

Dates are from OpenAI's published deprecation notices. Check your own models on the deprecations page of the OpenAI docs.

Benchmark

A 4B open model, fine-tuned, beats the alternatives

Banking77 stands in for a client: real bank-support messages sorted into 77 categories, with the same 3,080 held-out test messages for every model.

ModelAccuracy95% CINotes
Qwen3-4B + LoRA, served by vLLM
93.9%
93.1–94.741 prompt tokens per call · 0 invalid answers
TF-IDF + logistic regression
89.3%
88.1–90.3Free floor, runs on a CPU
Same Qwen3-4B, no fine-tuning
63.1%
61.4–64.7437 prompt tokens per call · 44 invalid answers
Gain over free+4.7 pts

paired 95% CI +3.7 to +5.8

Latency p950.55 s

at 87 requests/s on one NVIDIA L4

GPU cost$2.57

per million calls at full use

Training22 min

9,000 labelled examples on one L4

Method, statistics and limitations →

Demo

Ask it what a bank customer would ask

These messages come from the test set. The free baseline got each one wrong and the fine-tuned model got each one right.

Fine-tuned
Free baseline
Correct label

How it works

You see proof before you pay

  1. Send labelled examples

    A few hundred real inputs with the answers you'd accept, ideally the data you fine-tuned on. Your human-written labels are the ground truth.

  2. Get a parity report

    We fine-tune an open model and test it on a held-out set. The report says PASS, FAIL or INCONCLUSIVE against a margin agreed up front.

  3. Keep the weights

    You get the model files and an OpenAI-compatible endpoint. Host it with us or in your own cloud.

The full process, week by week →

The switch

Change one line

The endpoint speaks the OpenAI API. Your SDK, retries, logging and prompts stay as they are.

# before
- client = OpenAI()
# after
+ client = OpenAI(base_url="https://api.illate.dev/v1")

reply = client.chat.completions.create(
    model="acme-triage",
    messages=[{"role": "user", "content": ticket}],
)
How we work

Three rules we don't bend

Your labels are the truth

We measure against answers your team wrote, never against another model's opinion. We don't train on GPT, Claude or Gemini outputs.

We never hold your OpenAI key

When our model isn't sure, it says so and your own code decides what to do next. Nothing routes through your provider accounts.

We tell you when not to buy

If a cheap API model already passes on your data, the parity report says so and you owe nothing.

Find out if your model can move

Send a few hundred labelled examples. You get a parity report within a week, free.

Get a free parity check