AI Spend Assessment: findings

The full written deliverable, over one workspace’s month.

Back to the sample report

Sample — synthetic month, regenerable from the published generator

The deliverable, in full

Sample

Workspace
A mid-size workspace (synthetic sample)
Window
2026-08-01 to 2026-08-31
Requests
76,747
Metered
$3,642.58

Contents

  1. What this report coversAssembled
  2. What was measuredAssembled
  3. The month at a glanceAssembled
  4. Where the dollars went: by modelAssembled
  5. Where the dollars went: by API keyAssembled
  6. Where the dollars went: by routeAssembled
  7. How concentrated the spend isAssembled
  8. The heaviest dayAssembled
  9. What your caps didAssembled
  10. Cost per successful outputAssembled
  11. Re-derivable without trusting usAssembled
  12. What this report does not sayAssembled
  13. The two evenings, read by handWritten by hand
  14. Waste foundWritten by hand
  15. The ranked movesWritten by hand

What this report covers

This report covers 76,747 metered requests on one workspace across 31 days, 2026-08-01 to 2026-08-31. It is assembled from the receipts the gateway wrote at the time each request was served.

What was measured

$3,642.58 was metered across 76,747 requests in this window.

Every row names the published price list it was metered on. 4 lists priced this window; for the models in it their rates are identical, so the version stamp moved and the arithmetic did not.

Price listRequests
rpl-2026-07-156,153
rpl-2026-08-0551,582
rpl-2026-08-2212,530
rpl-2026-08-296,482

The month at a glance

The heaviest day is 2026-08-20 at $314.43, against a mean day of $117.50. This cut is assembled from the same metered rows as the sections below it.

The month at a glanceDaily metered spend across 31 days, split into cache read, uncached input, output. The heaviest day is 2026-08-20 at $314.43; an armed cap ended that evening and the rest of the night is refusals, which cost nothing and so are not on this chart.Cache readUncached inputOutput$0$100$200$300$400161116$314212631cap held 22:50 · 1,850 refused
Fig 1. Bands are cache read, uncached input, output, on one scale. Refusals cost nothing and are not on it.
The three bands, over the window
BandMeteredShare
Cache read$76.552.1%
Uncached input$3,038.8883.4%
Output$527.1514.5%
Window$3,642.58100.0%

Where the dollars went: by model

Where the dollars went, by modelSix models, ranked by metered spend. The bar length is the row's share of the heaviest row; the arrow is the second half of the window against the first (2026-08-01 to 2026-08-15 against the rest).SECOND HALF VS FIRSTclaude-sonnet-4-5$1,529.9016,151 requests · 42.0%▲ $215.10gpt-4.1$1,219.2313,102 requests · 33.5%▲ $274.32claude-haiku-4-5$554.1722,046 requests · 15.2%▲ $32.22gpt-4.1-mini$174.239,653 requests · 4.8%▲ $166.46claude-opus-4-5$88.88801 requests · 2.4%▲ $0.68gpt-4.1-nano$76.1714,994 requests · 2.1%▲ $6.29
Fig 2. Bar length is the row against the heaviest row.
ModelProviderRequestsMeteredShare
claude-sonnet-4-5Anthropic16,151$1,529.9042.0%
gpt-4.1OpenAI13,102$1,219.2333.5%
claude-haiku-4-5Anthropic22,046$554.1715.2%
gpt-4.1-miniOpenAI9,653$174.234.8%
claude-opus-4-5Anthropic801$88.882.4%
gpt-4.1-nanoOpenAI14,994$76.172.1%
Window76,747$3,642.58100.0%

Where the dollars went: by API key

Where the dollars went, by API keyFive keys, ranked by metered spend, each holding the colour it holds everywhere else in this document.SECOND HALF VS FIRSTkey_support$1,535.1031,972 requests · 42.1%▲ $93.25key_batch$1,219.2313,102 requests · 33.5%▲ $274.32key_agent$548.976,225 requests · 15.1%▲ $154.08key_dev$263.1110,454 requests · 7.2%▲ $167.14key_staging$76.1714,994 requests · 2.1%▲ $6.29
Fig 3. Colour follows the key, never its rank.
Share of spend, by API keyFive keys over $3,642.58 of metered spend. Two of them carry half the window between them.42.1%33.5%15.1%7.2%$3,642.58METERED
Fig 4. Share of the window, by key.
KeyRhythmRequestsMeteredShare
key_support
key_support: metered spend, day by day
31,972$1,535.1042.1%
key_batch
key_batch: metered spend, day by day
13,102$1,219.2333.5%
key_agent
key_agent: metered spend, day by day
6,225$548.9715.1%
key_dev
key_dev: metered spend, day by day
10,454$263.117.2%
key_staging
key_staging: metered spend, day by day
14,994$76.172.1%

Where the dollars went: by route

Where the dollars went, by routeThree routes, ranked by metered spend. Models and routes are two cuts of the same dollars; they do not add together.SECOND HALF VS FIRST/v1/messages$2,172.9538,998 requests · 59.7%▲ $248.01/v1/chat/completions$1,295.3928,096 requests · 35.6%▲ $280.61/v1/responses$174.239,653 requests · 4.8%▲ $166.46
Fig 5. Routes and models are two cuts of the same dollars.
RouteRequestsMeteredShare
/v1/messages38,998$2,172.9559.7%
/v1/chat/completions28,096$1,295.3935.6%
/v1/responses9,653$174.234.8%

How concentrated the spend is

key_support carries 42.1% of the window on its own, and 2 keys carry half of it between them. The heaviest route is /v1/messages at 59.7%.

The heaviest day

2026-08-20 metered $314.43 across 3,717 requests. Its heaviest hour is the one an armed cap ended.

2026-08-20, hour by hourThe heaviest day of the window, cut into 24 UTC hours. Spend climbs through the evening and stops inside the 23:00 hour, where the key's daily envelope refused everything behind it.Cache readUncached inputOutput$0$50$100$1500003060912151821$102
Fig 6. 2026-08-20, in UTC hours, on the same three bands.

What your caps did

1,850 requests were refused by an armed cap and 451 by a run's own declared budget. That is spend that did not happen.

Budget
Agent runner daily
Envelope
$200.00
Metered when it refused
$199.98
First refusal
2026-08-20 22:50:20

Cost per successful output

Across the window, $0.05 was metered for every response that succeeded. 4,853 requests answered something other than 2xx.

Cost per successful output, by keyMetered spend divided by 2xx responses. One key carries 3,143 requests that returned an error after the provider had already generated tokens.BILLED ON FAILURESkey_batch$0.0913,102 answered · 0 did notkey_agent$0.096,225 answered · 0 did notkey_support$0.0531,972 answered · 0 did notkey_dev$0.055,601 answered · 4,853 did not▲ $84.39key_staging$0.0114,994 answered · 0 did not
Fig 7. Everything metered, divided by what answered.
KeyAnsweredDid notPer answerMetered on failures
key_batch13,1020$0.09$0.00
key_agent6,2250$0.09$0.00
key_support31,9720$0.05$0.00
key_dev5,6014,853$0.05$84.39
key_staging14,9940$0.01$0.00

Re-derivable without trusting us

Your rows are hash-chained: each receipt carries the hash of the row before it, and the window re-derives from the export under the published recipe.

This sample is synthetic and says so, and it carries BOTH mechanisms. Regenerate from seed 20,260,831 and every figure in this document recomputes; hash the priced rows under the published recipe and the head is ed6c8a5cc933… over 75,037 rows.

The chain covers the rows the meter priced. 1,710 receipts here report no usage at all, so they carry no cost — and a chain row has no field for an absent one. They are receipts in this document and they are not rows in that chain.

What this report does not say

This report is descriptive: it says where the dollars went, not what to do about them.

The models, the keys and the routes above are three cuts of the same dollars. They do not add together.

Nothing here is a projection. Where two figures sit side by side, the subtraction is the reader’s.

The ranked moves are not in this half of the document. They are written by hand, from this evidence.

Written by hand, from the evidence above

The two evenings, read by hand

On 2026-08-20 the agent runner entered a loop at 19:40. It metered $199.98 that day against a $200.00 daily envelope, and at 22:50 the next request would have crossed it. The gateway refused that request and the 1,850 behind it, before a provider was contacted, until the loop was stopped at 23:51.

Two days earlier, on 2026-08-18, a provider wave hit key_dev between 13:00 and 16:30. A client retrying up to 5 times sent 7,868 requests; 3,143 of them returned an upstream error after the model had already generated tokens, and $84.39 was metered on requests whose caller got nothing. 1,710 more were refused at the provider's door and cost nothing — their receipts carry no cost at all rather than a zero.

No cap was armed on that key. Both evenings are in this document; neither is projected.

The two evenings, one point per requestEvery sampled request on 2026-08-18 and 2026-08-20, placed by the UTC hour it was made and by what it cost, coloured by the key that was billed. The afternoon nothing was armed for sits left; the evening a cap ended sits right and stops.key_supportkey_batchkey_agent$0.00$0.10$0.20$0.3000:0004:0008:0012:0016:0020:0024:00
Fig 8. One point per sampled request on 2026-08-18 and 2026-08-20.
Receipts from 2026-08-18
RequestAtModelStatusTokensCostPriced on
req_f8d9ac51d601f62013:00:00gpt-4.1-mini4290 in · 0 outnot reportedrpl-2026-08-05
req_6ddcd3b9b4b598d413:00:01gpt-4.1-mini4290 in · 0 outnot reportedrpl-2026-08-05
req_3a80cdfc796cce9f13:00:04gpt-4.1-mini20037,841 in · 997 out$0.02rpl-2026-08-05
req_f69aba8a9fe680a013:00:05gpt-4.1-mini50361,576 in · 424 out$0.03rpl-2026-08-05
Receipts from 2026-08-20
RequestAtModelStatusTokensCostPriced on
req_9e84efc953bec0b619:40:00claude-sonnet-4-520037,954 in · 1,811 out$0.14rpl-2026-08-05
req_5f60d73db18c661c22:03:49claude-sonnet-4-520038,662 in · 1,455 out$0.14rpl-2026-08-05
req_3a59f84e62a621be22:50:18claude-sonnet-4-520045,173 in · 1,364 out$0.08rpl-2026-08-05

Waste found

The retry storm: $84.39 metered on requests that returned an error.

The cache regression: cache-read share on the busiest key fell from 70.4% to 19.8% after a deploy and returned three weeks later. The window's own share is 32.7%.

The idle weekend: $62.35 across 2026-08-15 and 2026-08-16 on key_staging, 8,928 requests, while every other key was silent.

The model mix: 801 changelog-summary calls ran on claude-opus-4-5 at $0.11 per request. 185 of the same errand ran on gpt-4.1-mini at $0.01.

Cache-read share on the support keyProvider cache-read tokens as a share of all input tokens on the busiest key, day by day. The share collapses on 2026-08-06 and returns three weeks later.08-06 · 24.8%
Fig 9. Cache-read share, day by day.
The same errand, two models
Modelchangelog-summary callsMeteredPer request
claude-opus-4-5801$88.88$0.11
gpt-4.1-mini185$1.55$0.01

The ranked moves

Impact is modeled and net of quality, in your own dollars.

  1. 1

    Arm a daily envelope on every key that can loop

    One key had one and stopped inside its own evening. One did not and ran to the end of the wave.

    Confirm it
    §9 and §13, and the 1,850 refusals behind the 22:50 line in Fig 1.
  2. 2

    Cap the retries before the retries cap you

    A client that retries a large prompt five times bills the whole prompt five times. The cost of a failure is set by how much context is resent, not by how long the answer was.

    Confirm it
    §10, where the same key is the only one carrying failures.
  3. 3

    Watch the cache-read share like a metric, not a feeling

    A deploy reordered a prompt prefix and nothing broke, nothing alerted, and the receipts got more expensive for three weeks.

    Confirm it
    §14 and the cache series beside it.
  4. 4

    Turn off what nobody reads

    A debug loop ran a whole weekend on a key nobody watches, almost entirely input tokens, with the answers thrown away.

    Confirm it
    §14, and the two flat bars in Fig 1 at 2026-08-15 and 2026-08-16.
  5. 5

    Right-size the errand, not the model

    One recurring errand runs on a frontier model out of habit while the same errand runs on a small one in the same month. Both per-request figures are printed; the division is yours.

    Confirm it
    §14, and the by-model table in §4.

Sample figures throughout. A customer’s document carries their own.