Eval contracts and trace corpora
Once our eval defines what good means for a workflow, we are the arbiter of it. Switching vendors means rebuilding the definition of quality first.
Custom AI models trained from your work.
One model per job: a sales model, a CX model, an invoice-extraction model, a ticket-triage model, a document-review model.
Used by teams at


One agent loop turns a single user action into dozens of model calls, and they are the same few tasks over and over: classify this, route that, pull those fields. Each one is billed at the price of a model that could have written the workflow instead of running it.
classify_intent
one task, priced at the frontier rate
The same million calls, served by a model trained on this one task $200
Unit price is a blended frontier serving basis; the specialist figure is the same call at $0.20 per million tokens. The call counts are a scale, not a customer's bill.
Output tokens bill at three to five times input. Cutting what the model generates is the highest-leverage cost line you have, and it needs no training run to collect.
You are paying for volume, not intelligence. The hard part was solved months ago. What is left is the same shape of call, over and over, at a price set for reasoning the task no longer needs.
Four steps, and the fourth feeds the first. Every request your replaced workflow serves becomes training data for the next pass, which is why this is the only endpoint you own whose quality goes up between vendor releases.
production traces
feed the next pass
Step 01
Send two months of inference spend and the traces behind it. You get back a map of which workflows actually repeat, what each one costs you, and which are bounded enough to replace.
Request the route auditStep 02
Before anything is optimized, an eval contract fixes what a correct answer is for your workflow, with n counts and standard deviations. Skip it and every number that follows is decoration.
See the ladderStep 03
Only if the eval says you have to. We climb rung by rung and stop at the cheapest one that clears, which is why most bounded workflows never reach a training run at all.
Read the benchmarksStep 04
Weights you can download or an endpoint we host, your call. From then on it learns from its own production traces, and those traces are what the next pass captures.
What compoundsRouting picks the model. Training builds one when no existing model will do. Compression cuts what every call carries. They compound, and none of them touch your integration.
Share of one workload
Each request goes to the cheapest model that clears the eval for that task. The mix moves as the task’s own model gets better.
Illustrative shape, not a customer’s bill. Your split depends on your tasks.
Teachers → your model
Held-out score
Within six points of the frontier at 1/40th the cost per solved task.
29 held-out tasks, 4 trials each. Full accounting on the case study.
Same task, same eval
Let me think through this step by step. First, I need to consider each of the
available fields and determine which ones are relevant. {"intent":"refund","priority":2}
Reasoning the task no longer needs is the cheapest thing to remove, and it needs no training run. Score held at 0.9598 against 0.9560.
30-task operations holdout. Measured eval tokens, not an estimate.
Your integration is written once. Behind it, the same call gets routed down, then trained on its own traces, and the price per solved task falls while the score holds.
invoice_extraction
Worked example · one workload, three points in its life
First week
“Extract the invoice total and due date.”
After the eval contract
“Extract the invoice total and due date.”
After training
“Extract the invoice total and due date.”
The request text is identical in all three columns; that is the point. Figures are a worked example, not a customer’s account. The floor under them is measured: on τ³-bench a trained 27B reached $0.55 per solved task against Opus 5’s $23.22.
Your code, all three times
resp = client.chat.completions.create(
model=MODEL,
messages=messages,
extra_body={"task": "invoice_extraction"},
)
Byte-identical across all three. MODEL is one config value, changed with you and reversible. No redeploy, no SDK change, no migration.
What moved behind it
| Model | frontier | small open | yours |
|---|---|---|---|
| Score | 0.91 | 0.92 | 0.94 |
| $ / solve | $2.40 | $0.31 | $0.09 |
Each response says which model served it, so a quality change is traceable to a route change instead of guessed at.
invoice_extraction
Worked example · version history
| Version | Eval | $/solve | Traces | Outcome |
|---|---|---|---|---|
| v1 | 0.910 | $2.40 | 12,400 | Baseline |
| v2 | +0.006 | −34% | 18,900 | Shipped |
| v3 | +0.002 | −11% | 24,100 | Shipped |
| c4 | −0.011 | −29% | 26,700 | Rejected, missed the eval |
| v4 | +0.009 | −18% | 31,500 | Shipped |
Cheaper is necessary and not sufficient. c4 was 29% cheaper and did not ship, because it lost 0.011 on the contract. That row is the whole product: a candidate only reaches your traffic when it clears the eval and you accept it.
Every figure below is a measured run, published with its n count, its standard deviation and its cost accounting. So are the runs that lost.
A bigger blind GEPA budget lost, landing at 0.252 ±0.144. We published that too. The failures are why the rest of the numbers are believable.
Read the benchmarksTeams treat optimization as a menu. It is a ladder, and you only climb when the eval proves the rung below did not clear. Most bounded workflows stop by rung three.
Live software, days not weeks
Training, only once the eval demands it
Competitors start at rung five. That is the only rung they sell.
Enterprises get isolation. We get a compounding map of what works, without ever holding their data.
Once our eval defines what good means for a workflow, we are the arbiter of it. Switching vendors means rebuilding the definition of quality first.
Every engagement teaches which rung clears which task shape. That map compounds across customers regardless of what any single contract permits, because it contains no customer content.
A competitor arriving later starts cold against a model that has been improving on production traces for a year.
They optimize the model. We optimize the decision of whether to touch the model at all.
| Capability | Applied Compute | Osmosis | TrainLoop | Protégé |
|---|---|---|---|---|
| Learns from production traces | Yes | Yes | Not claimed | Yes |
| Continuous or automatic retraining | Yes | Hourly | Research only | Yes |
| Eval-gated ladder that can conclude “don’t train” | Not claimed | Not claimed | Not claimed | Yes |
| Ships trainingless wins as an outcome | Not claimed | Not claimed | Not claimed | Yes |
| Offline, warehouse-scale batch | Not claimed | Not claimed | Not claimed | Yes |
“Not claimed” means the capability is absent from public materials, not that the vendor is incapable of it. The category is real and well funded: Applied Compute has raised $100M, Osmosis $7M, TrainLoop went through YC W25. All three sell training.
Tushar Jain, founder
Applied Scientist II at Amazon from 2021 to 2025, working on active learning to accelerate LLM training and large-scale optimization. ML Scientist then ML Engineer II at Verisk before that, shipping applied ML into regulated enterprise.
Seven people, already shipping
Founding engineering, ML engineering, software, product design and business operations. The router and the prompt harness are live; supervised fine-tuning and reinforcement learning are in beta.
How we publish
Every result on this site carries its n count, its standard deviation and its cost accounting. When a run loses, it goes on the benchmarks page with the runs that won.
The ladder is productized. Rungs two, three and four are live software, and every engagement returns reusable rungs to the route library, so forward-deployed hours per engagement fall as coverage rises. The self-learning endpoint is the terminal product.
They assembled it to fill their own capacity, because post-training is a feature that sells compute. We stay neutral across providers, we hand back the weights, and we lead with the eval rather than the cluster. Different buyer, different incentive.
It does not have to. The cross-customer asset is the route library, a map of which rung clears which task shape, and it contains zero customer content. Per-tenant learning stays isolated and is sold as the feature.
There is a human approval gate on by default. No checkpoint promotes without sign-off in regulated tenants, and every retrain is versioned, diffable and rollback-able with a full audit trail. Silent self-redeployment is an audit finding, not a feature.
The weights and the eval contract are yours to keep and to run anywhere. Regulated data stays inside your warehouse throughout; on the 39,962-row labeling run we only needed the problem shape, meaning the domain, the labels, a few examples and the eval contract.
A cheaper model reaches your traffic only after it clears the eval contract you defined, and only after you accept it. If nothing clears, nothing moves and you keep paying what you pay today. That is the outcome we publish alongside the wins.
We are in private preview, working with a small group of design partners at a time: capture the traces, write the eval, and ship the first replacement model alongside the people who already know the work.
The fit is a production workflow where cost or latency is starting to hurt, and the people who know what good looks like are not on an ML team.
Book a callSend two months of inference spend and what currently counts as a correct answer. We do not need your data to start.
Prefer email? contact@protege.sh