Proof is the product.

Automated, provable evaluation.

Deciding, and proving, that a cheaper model or configuration is no worse, on real traffic, without a human grading every prompt. That problem is the program: the goal we are building toward, not a capability we ship today. There is no other agenda.

A program, not a labEval gate not running · proof pending

Everyone can route to a cheaper model. Almost no one can prove the answers didn't get worse.

That proof, automated, calibrated, and cheap enough to run continuously, is the bottleneck of the entire AI-cost field. It is why every savings figure on this site reads measured, never verified, and what must exist before that word is ever earned. This program is our answer, and the reason Recovea is a research institution in the making, not a dashboard.

THE RESEARCH QUESTION

Can an automated gate decide, with calibrated confidence, that a cheaper model or configuration is no worse than the one it replaces, on a route's own live traffic?

  • Real traffic, not benchmarks. A route's own prompts are the only distribution that matters; leaderboard scores don't transfer.
  • No human grading at scale. A person reading every prompt is a demo, not a system. The gate has to earn trust mechanically.
  • The proof must cost less than what it certifies. An evaluation bill that eats the savings proves nothing worth having.
  • Uncertainty resolves to "not proven." When the gate cannot tell, it says no. It never rounds up to a claim.

What's public today

Three commitments, stated plainly: the first you can check now; the other two say exactly when they become checkable.

OPEN METHODOLOGY

Every number links to how it's computed

No figure appears on this site without its method one click away. The only latency figure we publish is ~1ms p50 / 2ms p95 local gateway overhead (measured July 2026, non-production), labeled and dated exactly like that, because that is exactly what we know: a benchmark of the gateway hop, taken off production, in July 2026.

Live today
THE WEEKLY NUMBER

One measured observation a week

The program's output cadence: a single citable measurement (cost per successful output, an effective rate), honestly scoped, limits stated first. The list opens with the launch; the first number ships when it is real, not on a marketing calendar.

Starting soon
THE EVAL GATE

Shadow first; deciding when calibrated

The gate will run beside live traffic and record the verdicts it would have issued, applying none of them. Nothing runs beside live traffic today: the verdict store exists and is constrained so no verdict can be applied, no service writes to it, and the request tee that would feed it is attached only in tests. It starts deciding when its error rate is a measured number, not an assumption. We publish when it's real.

not running · proof pending

Where the gate stands

The state of the program's core instrument, in the open: status, not press release.

your appyour providerthe request path: untouchedeval gate: not runningverdicts recorded: none yet · applied: never
01Next

Shadow

The gate watches real routes and records what it would have decided. Zero verdicts applied, zero effect on traffic. Not started: nothing writes a verdict today, and the tee that would feed the gate is attached only in tests, so the deployed gateway observes nothing.

02proof pending

Deciding

Shadow verdicts are scored against outcomes until the gate's error rate is a measured number, not an assumption. Only then does it gate a live route, with the operator's sign-off, starting where it has proven itself.

03proof pending

Published

When gated decisions hold up on real traffic, measured becomes verified, and gain-share, off today, finally gets a reason to exist. We publish the calibration data alongside the claim. Not before.

The words we put next to a number

Every figure the product shows carries a basis. These are the four labels and what each one commits us to. They are definitions, not positioning.

Measured
Arithmetic over numbers the provider itself reported, priced against a version-pinned reference price list. Nothing is inferred: a count the provider did not report is stored as null, and a model absent from the price list is stored with a null cost rather than a guess. Measured is the default basis of every figure the product shows today, and every ledger row carries it as a field.
Modeled
A measured observation projected past the window it was observed in, or through a simulator. A modeled figure is always labelled modeled where it appears, and the measurement underneath it stays measured: the CLI scan band projects a month from a shorter export, and the counterfactual replays fingerprints of your own traffic through lever simulators. A projection is not an outcome, and we never round one up into one.
Applied
Something actually changed your traffic. Nothing on this site is applied: no optimization lever is activated in any deployed environment, so the counterfactual payload carries applied: false and the realized figure always equals the baseline.
Verified
Reserved, and unused. The word waits on the evaluation gate running on real traffic with a measured error rate. No savings figure on this site carries it, the platform basis vocabulary has no variant that could hold it, and gain-share stays off until it does. Where you see the word here, it is describing what would have to be true first.

Costs are computed against a frozen reference price list, today rpl-2026-08-05, whose version is pinned on every receipt and carried in every export header. A published version never changes: a rate correction ships as a new version, never as an edit of an old one, so a row priced last month re-derives to the same integer next year.

What is deliberately not claimed

A research page with nothing inflated.

  • No papers we haven't written. Nothing here cites a publication, because there are none yet. When there are, they will be linked with their data.
  • No lab we haven't staffed. This page says program because that is what exists. The word lab waits until it is true.
  • No results we haven't gated. Until the gate is calibrated on real traffic, no savings figure on this site is called verified. Every one reads measured.
  • Gain-share stays off · proof pending. We do not take a share of savings we cannot yet prove.

The program keeps a notebook.

Research notes, not a marketing blog: every post carries a number, its method one click away, its limits stated. Takes without numbers don't ship.

The weekly number arrives by email first. The subscription form is in the footer of this page.