Sample
- What this report covers
- What was measured
- The month at a glance
- Where the dollars went: by model
- Where the dollars went: by API key
- Where the dollars went: by route
- How concentrated the spend is
- The heaviest day
- What your caps did
- Cost per successful output
- Re-derivable without trusting us
- What this report does not say
- The two evenings, read by hand
- Waste found
- The ranked moves
What this report covers
This report covers 76,747 metered requests on one workspace across 31 days, 2026-08-01 to 2026-08-31. It is assembled from the receipts the gateway wrote at the time each request was served.
What was measured
$3,642.58 was metered across 76,747 requests in this window.
Every row names the published price list it was metered on. 4 lists priced this window; for the models in it their rates are identical, so the version stamp moved and the arithmetic did not.
| Price list | Requests |
|---|---|
| rpl-2026-07-15 | 6,153 |
| rpl-2026-08-05 | 51,582 |
| rpl-2026-08-22 | 12,530 |
| rpl-2026-08-29 | 6,482 |
The month at a glance
The heaviest day is 2026-08-20 at $314.43, against a mean day of $117.50. This cut is assembled from the same metered rows as the sections below it.
| Band | Metered | Share |
|---|---|---|
| Cache read | $76.55 | 2.1% |
| Uncached input | $3,038.88 | 83.4% |
| Output | $527.15 | 14.5% |
| Window | $3,642.58 | 100.0% |
Where the dollars went: by model
| Model | Provider | Requests | Metered | Share |
|---|---|---|---|---|
| claude-sonnet-4-5 | Anthropic | 16,151 | $1,529.90 | 42.0% |
| gpt-4.1 | OpenAI | 13,102 | $1,219.23 | 33.5% |
| claude-haiku-4-5 | Anthropic | 22,046 | $554.17 | 15.2% |
| gpt-4.1-mini | OpenAI | 9,653 | $174.23 | 4.8% |
| claude-opus-4-5 | Anthropic | 801 | $88.88 | 2.4% |
| gpt-4.1-nano | OpenAI | 14,994 | $76.17 | 2.1% |
| Window | 76,747 | $3,642.58 | 100.0% | |
Where the dollars went: by API key
| Key | Rhythm | Requests | Metered | Share |
|---|---|---|---|---|
| key_support | 31,972 | $1,535.10 | 42.1% | |
| key_batch | 13,102 | $1,219.23 | 33.5% | |
| key_agent | 6,225 | $548.97 | 15.1% | |
| key_dev | 10,454 | $263.11 | 7.2% | |
| key_staging | 14,994 | $76.17 | 2.1% |
Where the dollars went: by route
| Route | Requests | Metered | Share |
|---|---|---|---|
| /v1/messages | 38,998 | $2,172.95 | 59.7% |
| /v1/chat/completions | 28,096 | $1,295.39 | 35.6% |
| /v1/responses | 9,653 | $174.23 | 4.8% |
How concentrated the spend is
key_support carries 42.1% of the window on its own, and 2 keys carry half of it between them. The heaviest route is /v1/messages at 59.7%.
The heaviest day
2026-08-20 metered $314.43 across 3,717 requests. Its heaviest hour is the one an armed cap ended.
What your caps did
1,850 requests were refused by an armed cap and 451 by a run's own declared budget. That is spend that did not happen.
- Agent runner daily
- $200.00
- $199.98
- 2026-08-20 22:50:20
Cost per successful output
Across the window, $0.05 was metered for every response that succeeded. 4,853 requests answered something other than 2xx.
| Key | Answered | Did not | Per answer | Metered on failures |
|---|---|---|---|---|
| key_batch | 13,102 | 0 | $0.09 | $0.00 |
| key_agent | 6,225 | 0 | $0.09 | $0.00 |
| key_support | 31,972 | 0 | $0.05 | $0.00 |
| key_dev | 5,601 | 4,853 | $0.05 | $84.39 |
| key_staging | 14,994 | 0 | $0.01 | $0.00 |
Re-derivable without trusting us
Your rows are hash-chained: each receipt carries the hash of the row before it, and the window re-derives from the export under the published recipe.
This sample is synthetic and says so, and it carries BOTH mechanisms. Regenerate from seed 20,260,831 and every figure in this document recomputes; hash the priced rows under the published recipe and the head is ed6c8a5cc933… over 75,037 rows.
The chain covers the rows the meter priced. 1,710 receipts here report no usage at all, so they carry no cost — and a chain row has no field for an absent one. They are receipts in this document and they are not rows in that chain.
What this report does not say
This report is descriptive: it says where the dollars went, not what to do about them.
The models, the keys and the routes above are three cuts of the same dollars. They do not add together.
Nothing here is a projection. Where two figures sit side by side, the subtraction is the reader’s.
The ranked moves are not in this half of the document. They are written by hand, from this evidence.
The two evenings, read by hand
On 2026-08-20 the agent runner entered a loop at 19:40. It metered $199.98 that day against a $200.00 daily envelope, and at 22:50 the next request would have crossed it. The gateway refused that request and the 1,850 behind it, before a provider was contacted, until the loop was stopped at 23:51.
Two days earlier, on 2026-08-18, a provider wave hit key_dev between 13:00 and 16:30. A client retrying up to 5 times sent 7,868 requests; 3,143 of them returned an upstream error after the model had already generated tokens, and $84.39 was metered on requests whose caller got nothing. 1,710 more were refused at the provider's door and cost nothing — their receipts carry no cost at all rather than a zero.
No cap was armed on that key. Both evenings are in this document; neither is projected.
| Request | At | Model | Status | Tokens | Cost | Priced on |
|---|---|---|---|---|---|---|
| req_f8d9ac51d601f620 | 13:00:00 | gpt-4.1-mini | 429 | 0 in · 0 out | not reported | rpl-2026-08-05 |
| req_6ddcd3b9b4b598d4 | 13:00:01 | gpt-4.1-mini | 429 | 0 in · 0 out | not reported | rpl-2026-08-05 |
| req_3a80cdfc796cce9f | 13:00:04 | gpt-4.1-mini | 200 | 37,841 in · 997 out | $0.02 | rpl-2026-08-05 |
| req_f69aba8a9fe680a0 | 13:00:05 | gpt-4.1-mini | 503 | 61,576 in · 424 out | $0.03 | rpl-2026-08-05 |
| Request | At | Model | Status | Tokens | Cost | Priced on |
|---|---|---|---|---|---|---|
| req_9e84efc953bec0b6 | 19:40:00 | claude-sonnet-4-5 | 200 | 37,954 in · 1,811 out | $0.14 | rpl-2026-08-05 |
| req_5f60d73db18c661c | 22:03:49 | claude-sonnet-4-5 | 200 | 38,662 in · 1,455 out | $0.14 | rpl-2026-08-05 |
| req_3a59f84e62a621be | 22:50:18 | claude-sonnet-4-5 | 200 | 45,173 in · 1,364 out | $0.08 | rpl-2026-08-05 |
Waste found
The retry storm: $84.39 metered on requests that returned an error.
The cache regression: cache-read share on the busiest key fell from 70.4% to 19.8% after a deploy and returned three weeks later. The window's own share is 32.7%.
The idle weekend: $62.35 across 2026-08-15 and 2026-08-16 on key_staging, 8,928 requests, while every other key was silent.
The model mix: 801 changelog-summary calls ran on claude-opus-4-5 at $0.11 per request. 185 of the same errand ran on gpt-4.1-mini at $0.01.
| Model | changelog-summary calls | Metered | Per request |
|---|---|---|---|
| claude-opus-4-5 | 801 | $88.88 | $0.11 |
| gpt-4.1-mini | 185 | $1.55 | $0.01 |
The ranked moves
Arm a daily envelope on every key that can loop
One key had one and stopped inside its own evening. One did not and ran to the end of the wave.
Cap the retries before the retries cap you
A client that retries a large prompt five times bills the whole prompt five times. The cost of a failure is set by how much context is resent, not by how long the answer was.
Watch the cache-read share like a metric, not a feeling
A deploy reordered a prompt prefix and nothing broke, nothing alerted, and the receipts got more expensive for three weeks.
Turn off what nobody reads
A debug loop ran a whole weekend on a key nobody watches, almost entirely input tokens, with the answers thrown away.
Right-size the errand, not the model
One recurring errand runs on a frontier model out of habit while the same errand runs on a small one in the same month. Both per-request figures are printed; the division is yours.