One file, three columns
The parity check needs one CSV or JSONL file: the input your model sees, the answer a person would accept, and what your current model said. Here is the format, a script for that last column that runs with your own key, and what we do with the file.
- Rows read6
- Dropped, no label1
- Repeats, kept once1
- Labels4
- Current model agrees3 of 4
Then split 70 / 10 / 20 within each label. The 20% test slice is never trained on.
The three columns
| Column | What goes in it | |
|---|---|---|
| text | Required | The input exactly as your model receives it: the ticket body, the message, the email reply. Keep it as it was, typos included. |
| label | Required | The answer a person wrote or checked. This is what both models are scored against, so it matters more than anything else in the file. |
| current_model | Recommended | What your current model answered for the same row. Without it we can still train and test, but the verdict compares against baselines rather than against what you run today. |
Your own column names are fine; tell us which is which. Extra columns are ignored. UTF-8 CSV or one JSON object per line.
text,label,current_model "My card still hasn't arrived",card_arrival,card_arrival "The refund still isn't showing",refund_pending,extra_charge
{"text": "My card still hasn't arrived", "label": "card_arrival", "current_model": "card_arrival"}
{"text": "The refund still isn't showing", "label": "refund_pending", "current_model": "extra_charge"}How many rows
A firm verdict
The test slice holds 1,000 rows or more, enough to tell a 3-point gap from noise. This is the size we recommend.
Usually enough
Fine when the models are clearly apart. When they are close, the report may say INCONCLUSIVE and tell you how many more rows would settle it.
A first look
We can still train and show per-label results, but treat the verdict as a direction, not proof. Labels with fewer than 5 rows are trained on but not tested.
Best of all are rows your current model was not trained on, such as recent traffic that people have labelled. If you only have the fine-tuning file, send it anyway: your current model has seen those rows, which flatters it, and the report says so.
Filling in current_model, with your own key
We never ask for your API key. If you log your model's answers, join them from the logs and skip this. Otherwise this script calls your model once per row on your machine and writes the file to send.
# Runs on your machine with your key. We never see the key. import csv from openai import OpenAI client = OpenAI() # reads OPENAI_API_KEY from your environment MODEL = "ft:gpt-4.1-mini-2025-04-14:your-org::abc123" # the model production calls SYSTEM = open("system_prompt.txt").read() # the prompt production sends with open("examples.csv", newline="") as f, open("for_illate.csv", "w", newline="") as out: w = csv.DictWriter(out, fieldnames=["text", "label", "current_model"]) w.writeheader() for row in csv.DictReader(f): r = client.chat.completions.create( model=MODEL, temperature=0, messages=[{"role": "system", "content": SYSTEM}, {"role": "user", "content": row["text"]}], ) w.writerow({"text": row["text"], "label": row["label"], "current_model": r.choices[0].message.content.strip()})
For 5,000 short rows on a fine-tuned gpt-4.1-mini this costs a few dollars. Above about 20,000 rows, OpenAI's Batch API halves the price; we'll send a batch version on request. Use the same model, prompt and settings as production, so the comparison is with what your users actually get.
Remove what the task doesn't need
- Names, emails, phone numbers, addresses and account numbers, unless the task depends on them. A placeholder such as [EMAIL] keeps the sentence intact.
- Rows you are not allowed to share, such as rows from a customer whose contract forbids it.
- Labels that came from GPT, Claude or Gemini without a person checking them. Their terms restrict training other models on their outputs, so we don't.
A link, not an attachment
- Share it from Google Drive, Dropbox or as an S3 pre-signed URL, with access for poojith@illate.dev only.
- We sign a mutual NDA first if you want one, or yours.
- It goes to a storage volume used for your project alone, is used only for your model, and is deleted within 30 days or sooner on request, with a written certificate. Details: security and data.
Have the file, or not sure it fits?
Send the column names and a row count. We'll tell you before you export anything.