ILLATE
For the free check

One file, three columns

The parity check needs one CSV or JSONL file: the input your model sees, the answer a person would accept, and what your current model said. Here is the format, a script for that last column that runs with your own key, and what we do with the file.

textlabelcurrent_model
My card still hasn't arrivedcard_arrivalcard_arrivalmatch
Why was I charged twice for one order?extra_chargeextra_chargematch
How do I unlock my PIN?pin_blockedpin_blockedmatch
My card still hasn't arrivedcard_arrivalcard_arrivalrepeat
The refund still isn't showingrefund_pendingextra_chargediffers
Is there a fee to top up?top_up_feeno label
Intake
  • Rows read6
  • Dropped, no label1
  • Repeats, kept once1
  • Labels4
  • Current model agrees3 of 4

Then split 70 / 10 / 20 within each label. The 20% test slice is never trained on.

What our intake reads from your file. Illustrative rows.Empty rows dropped · repeats kept once · nothing rewritten
Format

The three columns

ColumnWhat goes in it
textRequiredThe input exactly as your model receives it: the ticket body, the message, the email reply. Keep it as it was, typos included.
labelRequiredThe answer a person wrote or checked. This is what both models are scored against, so it matters more than anything else in the file.
current_modelRecommendedWhat your current model answered for the same row. Without it we can still train and test, but the verdict compares against baselines rather than against what you run today.

Your own column names are fine; tell us which is which. Extra columns are ignored. UTF-8 CSV or one JSON object per line.

CSV
text,label,current_model
"My card still hasn't arrived",card_arrival,card_arrival
"The refund still isn't showing",refund_pending,extra_charge
JSONL
{"text": "My card still hasn't arrived", "label": "card_arrival", "current_model": "card_arrival"}
{"text": "The refund still isn't showing", "label": "refund_pending", "current_model": "extra_charge"}
Size

How many rows

5,000 or more

A firm verdict

The test slice holds 1,000 rows or more, enough to tell a 3-point gap from noise. This is the size we recommend.

1,000 to 5,000

Usually enough

Fine when the models are clearly apart. When they are close, the report may say INCONCLUSIVE and tell you how many more rows would settle it.

Under 1,000

A first look

We can still train and show per-label results, but treat the verdict as a direction, not proof. Labels with fewer than 5 rows are trained on but not tested.

Best of all are rows your current model was not trained on, such as recent traffic that people have labelled. If you only have the fine-tuning file, send it anyway: your current model has seen those rows, which flatters it, and the report says so.

Your current model's answers

Filling in current_model, with your own key

We never ask for your API key. If you log your model's answers, join them from the logs and skip this. Otherwise this script calls your model once per row on your machine and writes the file to send.

add_current_model.py
# Runs on your machine with your key. We never see the key.
import csv
from openai import OpenAI

client = OpenAI()  # reads OPENAI_API_KEY from your environment
MODEL = "ft:gpt-4.1-mini-2025-04-14:your-org::abc123"  # the model production calls
SYSTEM = open("system_prompt.txt").read()  # the prompt production sends

with open("examples.csv", newline="") as f, open("for_illate.csv", "w", newline="") as out:
    w = csv.DictWriter(out, fieldnames=["text", "label", "current_model"])
    w.writeheader()
    for row in csv.DictReader(f):
        r = client.chat.completions.create(
            model=MODEL,
            temperature=0,
            messages=[{"role": "system", "content": SYSTEM},
                      {"role": "user", "content": row["text"]}],
        )
        w.writerow({"text": row["text"], "label": row["label"],
                    "current_model": r.choices[0].message.content.strip()})

For 5,000 short rows on a fine-tuned gpt-4.1-mini this costs a few dollars. Above about 20,000 rows, OpenAI's Batch API halves the price; we'll send a batch version on request. Use the same model, prompt and settings as production, so the comparison is with what your users actually get.

Before you send

Remove what the task doesn't need

  • Names, emails, phone numbers, addresses and account numbers, unless the task depends on them. A placeholder such as [EMAIL] keeps the sentence intact.
  • Rows you are not allowed to share, such as rows from a customer whose contract forbids it.
  • Labels that came from GPT, Claude or Gemini without a person checking them. Their terms restrict training other models on their outputs, so we don't.
How to send it

A link, not an attachment

  • Share it from Google Drive, Dropbox or as an S3 pre-signed URL, with access for poojith@illate.dev only.
  • We sign a mutual NDA first if you want one, or yours.
  • It goes to a storage volume used for your project alone, is used only for your model, and is deleted within 30 days or sooner on request, with a written certificate. Details: security and data.

Have the file, or not sure it fits?

Send the column names and a row count. We'll tell you before you export anything.

Book a parity check