Engineering notes
What we measured, and how
Short write-ups from our own runs. Each one gives the numbers, how they were produced and where they stop being true.
5 Oct 2026
→
5 Oct 2026
→
5 Oct 2026
→
A 4B open model at 94% on 77-way support triage
Banking77 end to end: data split, LoRA training, paired-bootstrap parity against a free baseline, serving latency on one L4, and the limits of the result.
Zero invalid labels: constrained decoding for classification
Our first smoke test invented 68 labels in 200 answers. A token trie over the label set took that to zero, and set the concurrency limit we serve at.
When a dedicated GPU beats the API bill, and when it doesn't
Break-even for short classification calls is about 1.8M calls a month on one L4. Below that, the reason to own a model is control, not price.
Want these numbers for your model?
The free parity check produces the same report on your data: accuracy with intervals, per-class results, latency and cost.