The endpoint API
Every migrated model runs on a dedicated endpoint that speaks the OpenAI chat completions format. If your code calls OpenAI today, it calls ILLATE after changing the base URL and key.
Overview
Each client gets one hostname and one or more models behind it. Endpoints are issued at the end of a migration, once the parity report says PASS and you sign off. There is no public endpoint and no self-serve sign-up.
- Base URL
https://<client>.api.illate.dev/v1- Format
- OpenAI chat completions, request and response
- Endpoints
POST /chat/completions,GET /models- Auth
- Bearer key, one per environment
- Self-hosting
- The same stack deploys into your AWS, GCP or Azure account with your own base URL
Authentication
Send your key in the Authorization header. Keys are issued per environment (for example staging and production), can be rotated on request, and only work on your own hostname.
Authorization: Bearer $ILLATE_KEY
We never ask for your OpenAI, Anthropic or Google keys. Calls that need a frontier model stay in your code, with your key.
Chat completions
Send the same messages you send your current fine-tuned model. The prompt your model was trained with is fixed at migration, so most clients send only the user message.
POST /v1/chat/completions { "model": "acme-triage", "messages": [ {"role": "user", "content": "I ordered a card but it has not arrived."} ] }
The response has the standard shape. message.content is always one label from your schema, or defer.
{
"id": "chatcmpl-9f2c…",
"object": "chat.completion",
"model": "acme-triage",
"choices": [{
"index": 0,
"message": {"role": "assistant", "content": "card_arrival"},
"finish_reason": "stop"
}],
"usage": {"prompt_tokens": 41, "completion_tokens": 4, "total_tokens": 45}
}
Supported request fields are model, messages, max_tokens, temperature, seed, logprobs, top_logprobs and user. Other OpenAI fields behave as they do in vLLM's OpenAI-compatible server.
Labels and defer
Decoding is constrained to your label set, so the model can't return a label that doesn't exist. In our Banking77 benchmark that meant 0 invalid answers in 3,080, against 44 for the same model without fine-tuning or constraints.
When the model's probability for its best label falls below a threshold, the endpoint returns defer instead of guessing. We set the threshold on your development set during migration, to a defer rate you choose (often 5 to 10%). Your code then takes its existing path:
r = client.chat.completions.create(model="acme-triage", messages=msgs) label = r.choices[0].message.content if label == "defer": label = classify_with_current_provider(ticket) # your code, your key
Pass logprobs: true if you want the token probabilities behind each answer for your own monitoring.
Models and versions
GET /v1/models lists the models on your endpoint. Each retrain creates a new dated version, such as acme-triage@2026-11-02. The plain name points at the version you approved, and it only moves after a new PASS report and your written sign-off. Pin a dated version if you want to control the switch yourself.
No forced deprecations. A version you rely on keeps running until you ask us to retire it. The weights are yours either way.
Errors
Errors use OpenAI's JSON shape, so existing retry and logging code handles them unchanged.
{"error": {"type": "rate_limit_error", "code": "rate_limited",
"message": "Above the agreed concurrency for this endpoint."}}
| Status | Meaning | What to do |
|---|---|---|
| 400 | Malformed request, or input longer than the model's limit | Fix the request; don't retry |
| 401 | Missing, wrong or rotated key | Check the key for this environment |
| 404 | Unknown model or version | Call GET /v1/models |
| 429 | Above the agreed concurrency | Retry with backoff, or ask us to add a replica |
| 5xx | Server error or a replica restarting | Retry with backoff, or fall back as for defer |
Limits
Limits are set per endpoint from your measured peak, and written into the hosting agreement.
- Concurrency
- Up to 32 requests in flight per replica, where our L4 benchmark peaked at 86.6 req/s. Higher peaks get more replicas, not a deeper queue.
- Latency
- p95 of 546 ms at 32 in flight on one L4 for a 41-token prompt, measured in-region. Add your network round trip.
- Input length
- 2,048 tokens per request by default; raised on request
- Warm capacity
- Production endpoints keep at least one warm replica, so there is no cold start
- Timeout
- Requests are cut off after 30 seconds
Numbers in this reference come from the Banking77 benchmark. Your endpoint's numbers are measured on your traffic before cutover and included in your report.