Claude Benchmarks
Suite cost (USD)
- sonnet-5 (high)
- fable-5 (high)
- haiku-4-5
- opus-4-8 (high)
- opus-5 (high)
- fable-5-1 (high)
Data table
| Claude Code CLI version | sonnet-5 (high) | fable-5 (high) | haiku-4-5 | opus-4-8 (high) | sonnet-5 (medium) | opus-5 (high) | opus-5 (low) | opus-5 (max) | opus-5 (medium) | opus-5 (xhigh) | fable-5 (medium) | opus-4-8 (medium) | fable-5 (max) | fable-5 (xhigh) | opus-4-8 (max) | opus-4-8 (xhigh) | sonnet-5 (max) | sonnet-5 (xhigh) | fable-5-1 (high) | fable-5-1 (medium) | fable-5-1 (max) | fable-5-1 (xhigh) |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 2.1.196 | 0.39959 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| 2.1.202 | 0.307441 | 1.585473 | 0.138659 | 0.787706 | 0.292861 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| 2.1.204 | 0.43591 | 1.589345 | 0.136938 | 0.7473 | 0.423905 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| 2.1.205 | 0.401936 | 1.412947 | 0.137007 | 0.68471 | 0.399699 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| 2.1.206 | 0.415858 | 1.452328 | 0.134095 | 0.711829 | 0.411794 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| 2.1.207 | 0.313161 | 1.169142 | 0.151735 | 0.494431 | 0.292131 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| 2.1.208 | 0.301529 | 1.052804 | 0.137217 | 0.473321 | 0.271329 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| 2.1.210 | 0.282338 | 1.052505 | 0.135879 | 0.478824 | 0.261952 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| 2.1.211 | 0.304312 | 1.042501 | 0.135169 | 0.441384 | 0.257755 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| 2.1.212 | 0.304967 | 1.037562 | 0.131874 | 0.449024 | 0.266073 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| 2.1.214 | 0.298263 | 1.00327 | 0.140986 | 0.470604 | 0.270236 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| 2.1.215 | 0.289674 | 0.98978 | 0.129231 | 0.403363 | 0.265379 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| 2.1.216 | 0.269861 | 0.932445 | 0.143514 | 0.419973 | 0.250677 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| 2.1.217 | 0.29384 | 0.970618 | 0.138043 | 0.405417 | 0.256946 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| 2.1.218 | 0.284298 | 0.993526 | 0.128458 | 0.402994 | 0.265344 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| 2.1.219 | 0.286448 | 1.040209 | 0.132845 | 0.408153 | 0.261634 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| 2.1.220 | 0.276864 | 0.987417 | 0.125554 | 0.40121 | 0.264421 | 0.657846 | 0.432675 | 0.815055 | 0.566123 | 0.762884 | — | — | — | — | — | — | — | — | — | — | — | — |
| 2.1.221 | 0.286049 | 0.977372 | 0.131999 | 0.401626 | — | 0.65512 | — | 0.873209 | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| 2.1.222 | 0.29028 | 1.098188 | 0.131929 | 0.443637 | 0.268545 | 0.646637 | — | — | 0.56267 | — | 0.892297 | 0.380577 | — | — | — | — | — | — | — | — | — | — |
| 2.1.223 | 0.296099 | 0.997915 | 0.136063 | 0.419563 | 0.26388 | 0.654669 | — | — | 0.563948 | — | 1.019459 | 0.417344 | — | — | — | — | — | — | — | — | — | — |
| 2.1.224 | 0.292122 | 1.085281 | 0.131863 | 0.441856 | 0.264969 | 0.662316 | — | — | 0.574012 | — | 0.931341 | 0.397382 | — | — | — | — | — | — | — | — | — | — |
| 2.1.225 | 0.29478 | 1.084727 | 0.136727 | 0.424706 | 0.26217 | 0.68277 | — | — | 0.582747 | — | 0.945328 | 0.392409 | — | — | — | — | — | — | — | — | — | — |
| 2.1.226 | 0.298991 | 1.059551 | 0.143038 | 0.435342 | 0.262172 | 0.688851 | — | — | 0.593155 | — | 0.924563 | 0.423615 | — | — | — | — | — | — | — | — | — | — |
| 2.1.227 | 0.296432 | 1.074569 | 0.132082 | 0.474679 | 0.260069 | 0.72558 | — | — | 0.567964 | — | 0.992387 | 0.43579 | — | — | — | — | — | — | — | — | — | — |
| 2.1.228 | 0.309403 | 1.06973 | 0.133659 | 0.422149 | 0.253376 | 0.679035 | — | — | 0.573721 | — | 0.912215 | 0.413112 | — | — | — | — | — | — | — | — | — | — |
| 2.1.229 | 0.30224 | 1.069366 | 0.136153 | 0.452609 | 0.259779 | 0.683558 | — | — | 0.564781 | — | 0.909137 | 0.424274 | — | — | — | — | — | — | — | — | — | — |
| 2.1.231 | 0.301637 | 1.046028 | 0.134929 | 0.442898 | 0.261313 | 0.709022 | — | — | 0.574353 | — | 0.946331 | 0.407458 | — | — | — | — | — | — | — | — | — | — |
| 2.1.232 | 0.296741 | 1.046169 | 0.126988 | 0.448157 | 0.263367 | 0.636067 | — | — | 0.593776 | — | 0.934624 | 0.429926 | — | — | — | — | — | — | — | — | — | — |
| 2.1.233 | 0.299494 | 1.039348 | 0.131054 | 0.397076 | 0.263568 | 0.660987 | — | — | 0.608725 | — | 0.893037 | 0.423469 | — | — | — | — | — | — | — | — | — | — |
| 2.1.234 | 0.283871 | 1.049789 | 0.123664 | 0.415698 | 0.267799 | 0.627984 | — | — | 0.55549 | — | 0.884193 | 0.415903 | — | — | — | — | — | — | — | — | — | — |
| 2.1.235 | 0.28956 | 1.027151 | 0.128207 | 0.450877 | 0.251189 | 0.621802 | — | — | 0.533325 | — | 0.882008 | 0.436997 | — | — | — | — | — | — | — | — | — | — |
| 2.1.236 | 0.287515 | 0.957125 | 0.132953 | 0.454512 | 0.273248 | 0.631873 | — | — | 0.516582 | — | 0.923976 | 0.445696 | — | — | — | — | — | — | — | — | — | — |
| 2.1.237 | 0.295928 | 1.062191 | 0.131669 | 0.429356 | 0.254227 | 0.624628 | — | — | 0.512943 | — | 0.923061 | 0.42662 | — | — | — | — | — | — | — | — | — | — |
| 2.1.238 | 0.273972 | 0.959072 | 0.135126 | 0.418491 | 0.26626 | 0.620326 | — | — | 0.505544 | — | 0.902759 | 0.421965 | — | — | — | — | — | — | — | — | — | — |
| 2.1.239 | 0.277027 | 1.005888 | 0.132268 | 0.44007 | 0.256925 | 0.668432 | — | — | 0.521478 | — | 0.863074 | 0.412674 | — | — | — | — | — | — | — | — | — | — |
| 2.1.240 | 0.296605 | 1.050034 | 0.129466 | 0.464848 | — | 0.627648 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| 2.1.241 | 0.284886 | 1.019615 | 0.134739 | 0.432911 | 0.264402 | 0.634396 | — | — | 0.495076 | — | 0.89114 | 0.440323 | — | — | — | — | — | — | — | — | — | — |
| 2.1.243 | 0.281347 | 1.065704 | 0.128777 | 0.439055 | 0.258768 | 0.639663 | — | — | 0.510795 | — | 0.895002 | 0.418189 | — | — | — | — | — | — | — | — | — | — |
| 2.1.245 | 0.277618 | 1.078632 | 0.127699 | 0.437461 | 0.275164 | 0.62811 | — | — | 0.506187 | — | 0.938824 | 0.4392 | — | — | — | — | — | — | — | — | — | — |
| 2.1.246 | 0.292935 | 1.026344 | 0.128329 | 0.430404 | 0.262555 | 0.598752 | — | — | 0.512927 | — | 0.90136 | 0.445008 | — | — | — | — | — | — | — | — | — | — |
| 2.1.247 | 0.274649 | 1.069361 | 0.12853 | 0.452896 | 0.262064 | 0.612011 | — | 0.881819 | 0.509641 | 0.759401 | 0.944405 | 0.431513 | 1.468829 | 2.045082 | 0.710167 | 0.549591 | 0.427245 | 0.319708 | — | — | — | — |
| 2.1.248 | 0.241969 | 0.850789 | 0.116072 | 0.372271 | 0.221059 | 0.537779 | — | 0.807233 | 0.445608 | 0.674777 | 0.786564 | 0.379949 | 2.842755 | 1.903706 | 0.662301 | 0.457867 | 0.387419 | 0.280323 | — | — | — | — |
| 2.1.250 | 0.243362 | 0.945936 | 0.119991 | 0.372947 | 0.22632 | 0.55282 | — | 0.828557 | 0.448003 | 0.696182 | 0.760813 | 0.354301 | 2.829469 | 1.871419 | 0.650926 | 0.45787 | 0.354036 | 0.281948 | — | — | — | — |
| 2.1.251 | 0.250543 | 0.870722 | 0.118075 | 0.385571 | 0.235105 | 0.53371 | — | 0.787098 | 0.415421 | 0.72368 | 0.784472 | 0.359907 | 3.017697 | 1.857015 | 0.670861 | 0.433931 | 0.36842 | 0.282411 | — | — | — | — |
| 2.1.252 | 0.25062 | 0.891468 | 0.12186 | 0.372919 | 0.236097 | 0.527568 | — | — | 0.405826 | — | 0.775454 | 0.361293 | — | — | — | — | — | — | — | — | — | — |
| 2.1.257 | 0.25659 | 0.87196 | 0.114944 | 0.380457 | 0.225249 | 0.525918 | — | — | 0.439004 | — | 0.738726 | 0.363476 | — | — | — | — | — | — | — | — | — | — |
| 2.1.258 | 0.247104 | 0.819051 | 0.11879 | 0.383242 | — | 0.535461 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
| 2.1.259 | 0.248202 | 0.839904 | 0.118003 | 0.386357 | 0.233196 | 0.526803 | — | — | 0.430426 | — | 0.747331 | 0.360691 | — | — | — | — | — | — | — | — | — | — |
| 2.1.260 | 0.248078 | 0.836241 | 0.117441 | 0.367197 | 0.223846 | 0.554445 | — | — | 0.415243 | — | 0.787331 | 0.365427 | — | — | — | — | — | — | — | — | — | — |
| 2.1.261 | 0.242446 | 0.867681 | 0.1213 | 0.37043 | 0.222379 | 0.554465 | — | — | 0.44107 | — | 0.74602 | 0.368507 | — | — | — | — | — | — | 0.870402 | 0.748292 | — | — |
| 2.1.263 | 0.246518 | 0.928523 | 0.12043 | 0.383042 | 0.229081 | 0.580929 | — | — | 0.412345 | — | 0.717638 | 0.366587 | — | — | — | — | — | — | 0.881951 | 0.737321 | — | — |
| 2.1.265 | 0.251823 | 0.880232 | 0.121911 | 0.386437 | 0.227746 | 0.601206 | — | — | 0.429159 | — | 0.806786 | 0.345241 | — | — | — | — | — | — | 0.840734 | 0.751287 | — | — |
| 2.1.266 | 0.251599 | 0.825064 | 0.118225 | 0.380074 | 0.226531 | 0.582992 | — | — | 0.439395 | — | 0.790282 | 0.358274 | — | — | — | — | — | — | 0.821263 | 0.755636 | — | — |
| 2.1.267 | 0.260124 | 0.913793 | 0.123804 | 0.380495 | 0.223534 | 0.584541 | — | — | 0.463294 | — | 0.756674 | 0.359009 | — | — | — | — | — | — | 0.848672 | 0.755992 | — | — |
| 2.1.268 | 0.247574 | 0.847196 | 0.120601 | 0.402512 | 0.233332 | 0.572696 | — | — | 0.467918 | — | 0.74891 | 0.371418 | — | — | — | — | — | — | 0.88616 | 0.76212 | — | — |
| 2.1.269 | 0.247532 | 0.883161 | 0.125496 | 0.37619 | 0.225884 | 0.595832 | — | — | 0.46909 | — | 0.792319 | 0.367078 | — | — | — | — | — | — | 0.854351 | 0.764239 | — | — |
| 2.1.270 | 0.248077 | 0.872274 | 0.120672 | 0.36345 | 0.217774 | 0.606586 | — | 0.788327 | 0.449836 | 0.69398 | 0.779055 | 0.383992 | 1.302987 | 0.970925 | 0.612335 | 0.461792 | 0.359038 | 0.280361 | 0.885484 | 0.760937 | 2.413816 | 1.371101 |
| 2.1.271 | 0.248705 | 0.875075 | 0.117043 | 0.379788 | 0.228741 | 0.607335 | — | — | 0.468559 | — | 0.78133 | 0.364044 | — | — | — | — | — | — | 0.920169 | 0.755311 | — | — |
Data table
| Metric | fable-5 (high) | fable-5-1 (high) | haiku-4-5 | opus-4-8 (high) | opus-5 (high) | sonnet-5 (high) |
|---|---|---|---|---|---|---|
| Cost per session (USD)mean | 0.104402 | 0.10984 | 0.013765 | 0.046933 | 0.073275 | 0.029335 |
| Cost per session (USD)median | 0.111082 | 0.116529 | 0.01547 | 0.050216 | 0.065865 | 0.031911 |
| Output tokens | 501 | 580 | 704 | 456 | 570 | 394 |
| Input tokens | 6 | 66 | 26 | 6 | 8 | 8 |
| Cache-read tokens | 50080 | 51481 | 62021 | 45809 | 65947.5 | 108334 |
| Cache-write tokens | 568 | 962 | 988 | 722 | 1110 | 830 |
| Input + cache tokens | 50615 | 52256 | 62690 | 46285 | 66913 | 111003 |
| API calls | 3 | 3 | 3 | 3 | 4 | 4 |
| Tool calls | 2 | 2 | 2 | 2 | 3 | 3 |
| Thinking blocks | 1 | 1 | 3 | 0 | 1 | 1 |
| Latency (s, wall-clock) | 10.001 | 9.39 | 9.198 | 8.519 | 9.7695 | 7.932 |
| Suite cost (USD) | No published column carried Suite cost (USD) for this cycle. | |||||
Data table
| Metric | fable-5 (medium) | fable-5-1 (medium) | opus-4-8 (medium) | opus-5 (medium) | sonnet-5 (medium) |
|---|---|---|---|---|---|
| Cost per session (USD)mean | 0.098681 | 0.094154 | 0.046869 | 0.060352 | 0.027328 |
| Cost per session (USD)median | 0.100608 | 0.100156 | 0.052755 | 0.05741 | 0.031323 |
| Output tokens | 447 | 448 | 459 | 455 | 356.5 |
| Input tokens | 6 | 66 | 6 | 6 | 8 |
| Cache-read tokens | 49728 | 48031 | 45809.5 | 49203 | 106887 |
| Cache-write tokens | 259 | 812 | 432 | 1210 | 779 |
| Input + cache tokens | 50436 | 52167 | 46263.5 | 49768 | 110992.5 |
| API calls | 3 | 3 | 3 | 3 | 4 |
| Tool calls | 2 | 2 | 2 | 2 | 3 |
| Thinking blocks | 1 | 0 | 0 | 0 | 1 |
| Latency (s, wall-clock) | 9.109 | 8.1805 | 9.445 | 7.854 | 6.9565 |
| Suite cost (USD) | No published column carried Suite cost (USD) for this cycle. | ||||
Claude Code CLI 2.1.271 Analysis
Measured 2026-09-15 00:46 UTC from this cycle’s benchmark runs, written up from those measurements alone, then checked by an independent reviewer. Where the wording and the tables disagree, the tables are correct.
Verdict
At the high reasoning setting, one benchmark task costs, cheapest to most expensive:
- haiku-4-5 $0.0138 per session
- sonnet-5 $0.0293 per session
- opus-4-8 $0.0469 per session
- opus-5 $0.0733 per session
- fable-5 $0.1044 per session
- fable-5-1 $0.1098 per session
At the medium setting:
- sonnet-5 $0.0273 per session
- opus-4-8 $0.0469 per session
- opus-5 $0.0604 per session
- fable-5-1 $0.0942 per session
- fable-5 $0.0987 per session
Across the last 10 cycles at the high setting (9 cycles for fable-5-1):
- opus-5 drifted up, staying inside $0.0662 to $0.0733
- fable-5 showed no clear direction, staying inside $0.1018 to $0.1139
- fable-5-1 showed no clear direction, staying inside $0.1027 to $0.1098
- haiku-4-5 showed no clear direction, staying inside $0.0138 to $0.0148
- opus-4-8 showed no clear direction, staying inside $0.0457 to $0.0496
- sonnet-5 showed no clear direction, staying inside $0.0290 to $0.0314
Across the last 10 cycles at the medium setting (9 cycles for fable-5-1):
- fable-5 drifted up, staying inside $0.0917 to $0.0987
- opus-4-8 drifted up, staying inside $0.0435 to $0.0469
- opus-5 drifted up, staying inside $0.0502 to $0.0604
- fable-5-1 showed no clear direction, staying inside $0.0926 to $0.0960
- sonnet-5 showed no clear direction, staying inside $0.0265 to $0.0282
This cycle moved from CLI version 2.1.270 to 2.1.271, and the families with a like-for-like comparison moved as follows:
- haiku fell 4.0% at the high setting
- sonnet fell 1.4% at the high setting
- sonnet rose 2.3% at the medium setting
In all three cases the individual runs behind those averages overlap heavily, so the moves do not stand out from normal run-to-run variation.
One more thing worth knowing: the spread between families is far larger than any of this cycle’s moves. At the high setting, haiku-4-5 costs 87.5% less per task than fable-5-1, and that gap is large enough to stand clear of run-to-run variation.
Which cost figure is which – setting high
The same sessions summarized three ways. They disagree on purpose, and none is the corrected version of another.
- Mean $/session: multiplies out to real spend, which makes it the budgeting number, and it is what the Cost Trends chart plots.
- Median $/session: shrugs off an unusually cheap or expensive run, but it does not sum to real spend.
- Suite cost: the price of one full pass over every prompt in the catalog, the per-prompt medians added up, so it answers what a whole run costs rather than what one session costs.
| family | models compared | mean $/sess | mean delta % | median $/sess | median delta % | suite $ | suite delta % | prompts |
|---|---|---|---|---|---|---|---|---|
| haiku | haiku-4-5 | 0.0143 -> 0.0138 | -4.0% | 0.0144 -> 0.0155 | +7.5% | 0.1207 -> 0.1170 | -3.0% | 9 |
| sonnet | sonnet-5 | 0.0297 -> 0.0293 | -1.4% | 0.0327 -> 0.0319 | -2.4% | 0.2481 -> 0.2487 | +0.3% | 9 |
Suite cost uses only the prompts both cycles ran, so the two sides always add up the same catalog.
Which cost figure is which – setting medium
The same sessions summarized three ways. They disagree on purpose, and none is the corrected version of another.
- Mean $/session: multiplies out to real spend, which makes it the budgeting number, and it is what the Cost Trends chart plots.
- Median $/session: shrugs off an unusually cheap or expensive run, but it does not sum to real spend.
- Suite cost: the price of one full pass over every prompt in the catalog, the per-prompt medians added up, so it answers what a whole run costs rather than what one session costs.
| family | models compared | mean $/sess | mean delta % | median $/sess | median delta % | suite $ | suite delta % | prompts |
|---|---|---|---|---|---|---|---|---|
| sonnet | sonnet-5 | 0.0267 -> 0.0273 | +2.3% | 0.0289 -> 0.0313 | +8.2% | 0.2178 -> 0.2287 | +5.0% | 9 |
Suite cost uses only the prompts both cycles ran, so the two sides always add up the same catalog.
What these terms mean
New to this report? Here is the vocabulary used below, in plain words. Skip ahead if you already know it.
- Session (also “task”): one complete run of one prompt by one model, start to finish. Every cost in this report is per session.
- Family: a model line, such as opus or sonnet. One family can cover more than one version, for example opus-4-8 and opus-5.
- Cache, warm and cold: the tool reuses recent context instead of sending it again, which is far cheaper. A session that gets to reuse it is “warm”; one that has to send everything fresh is “cold”. Whether a given session lands warm or cold is partly luck, and that luck moves the cost.
- Cache luck (or cache-hit luck): the cost swing caused by how many sessions happened to land warm rather than cold, rather than by any real change in the tool or the model.
- Blended cost: cost measured across every session, warm and cold alike. This is the real money figure, and it is what the cost feed shows. “Blended” says which sessions are counted, not how they are averaged; see “Which cost figure is which” above for the averaging.
- Warm-only cost: cost measured across warm sessions only. Removing the cold ones removes most of the luck, so this is the fair test of whether cost actually changed.
- Pooled: all sessions for a family added together into a single number, rather than split out per prompt.
- Mean cost per session: the ordinary average across every session. It multiplies out to real spend, so it is the budgeting figure and the one this report leads with.
- Median cost per session: the middle session once they are lined up by cost. It ignores an unusually cheap or expensive run, but it does not add up to real spend.
- Suite cost: the price of one full pass over every prompt in the catalog, found by taking each prompt’s median and adding those up. It is a total for the whole catalog, not an average per session, so it is a larger number than either figure above and is not comparable to them.
- Cell: one family and one prompt combination, for example haiku on the rename-refactor prompt. Cells hold few sessions, so a single cell is noisy.
- n: how many sessions sit behind a number. A larger n is more trustworthy.
- p-value, and “significant”: the same prompt run twice costs different amounts, so a p-value asks whether two sets of runs can be told apart at all. Line up every run from both cycles: if one cycle’s runs sit consistently higher, p is near 0 and the difference is called “significant”; if the runs are mixed together, p is near 1 and the two sets look like the same thing measured twice. Precisely, it is the chance of seeing a gap at least this lopsided if the two cycles were truly identical. A high p is not weak evidence of a change, it is no evidence either way.
- API-equivalent cost: dollars calculated at published API rates. It is not necessarily dollars billed, because a subscription plan may already cover the usage.
Pooled blended comparison
These are the headline figures: the average real spend on one task, counting every session in the cycle, including the unlucky ones that started with a cold cache. Each setting is compared against its own previous run at the same setting. Only families that ran in both cycles at a given setting can be compared.
High reasoning setting, nightly-2.1.270 to nightly-2.1.271:
| family | models compared | n (prev -> cur) | prev $/sess | cur $/sess | delta % | Mann-Whitney p |
|---|---|---|---|---|---|---|
| haiku | haiku-4-5 | 53 -> 53 | 0.0143 | 0.0138 | -4.0% | 0.726 |
| sonnet | sonnet-5 | 53 -> 53 | 0.0297 | 0.0293 | -1.4% | 0.922 |
Medium reasoning setting, nightly-medium-2.1.270 to nightly-medium-2.1.271:
| family | models compared | n (prev -> cur) | prev $/sess | cur $/sess | delta % | Mann-Whitney p |
|---|---|---|---|---|---|---|
| sonnet | sonnet-5 | 53 -> 54 | 0.0267 | 0.0273 | +2.3% | 0.616 |
Pooled warm comparison (significance control)
Whether a cost move is believable is a separate question from how big it looks. This second view keeps only the sessions that ran with a warm cache, which removes the random luck of a cold start, and compares the individual run costs from each cycle against each other. It is the check to trust when asking whether cost actually changed.
High reasoning setting:
| family | models compared | n (prev -> cur) | prev $/sess | cur $/sess | delta % | Mann-Whitney p |
|---|---|---|---|---|---|---|
| haiku | haiku-4-5 | 45 -> 45 | 0.0132 | 0.0129 | -2.7% | 0.806 |
| sonnet | sonnet-5 | 45 -> 45 | 0.0276 | 0.0271 | -1.8% | 0.904 |
Medium reasoning setting:
| family | models compared | n (prev -> cur) | prev $/sess | cur $/sess | delta % | Mann-Whitney p |
|---|---|---|---|---|---|---|
| sonnet | sonnet-5 | 45 -> 45 | 0.0239 | 0.0250 | +4.7% | 0.529 |
Every one of these comparisons sits far above the 0.05 line. That does not mean nothing changed: it means this measurement cannot tell these moves apart from the ordinary variation between runs.
What the reasoning setting costs
The same release was measured twice, once at the high reasoning setting and once at medium. Comparing the two pooled sets of session costs for each model gives the following, with medium measured against high and a negative percentage meaning medium is the cheaper setting.
| model | high mean $ | medium mean $ | change | p | n |
|---|---|---|---|---|---|
fable-5 |
0.1044 | 0.0987 | -5.5% | 0.275 | 53/55 |
fable-5-1 |
0.1098 | 0.0942 | -14.3% | 0.141 | 53/54 |
opus-4-8 |
0.0469 | 0.0469 | -0.1% | 0.711 | 53/56 |
opus-5 |
0.0733 | 0.0604 | -17.6% | 0.165 | 56/61 |
sonnet-5 |
0.0293 | 0.0273 | -6.8% | 0.271 | 53/54 |
Medium is the cheaper setting in every row, but no row clears the 0.05 line:
- opus-5 is 17.6% cheaper at medium, and the difference cannot be told apart from run-to-run variation
- fable-5-1 is 14.3% cheaper at medium, and the difference cannot be told apart from run-to-run variation
- sonnet-5 is 6.8% cheaper at medium, and the difference cannot be told apart from run-to-run variation
- fable-5 is 5.5% cheaper at medium, and the difference cannot be told apart from run-to-run variation
- opus-4-8 is 0.1% cheaper at medium, and the difference cannot be told apart from run-to-run variation
haiku-4-5 was measured at the high setting only, so it has no cross-setting figure.
Notable cells (secondary)
At the high setting, no individual prompt moved beyond the noise threshold. At the medium setting, two sonnet-5 prompts moved more than the rest:
fix-off-by-one, which fixes a one-line off-by-one bug in a small Python function, rose 30.8% on warm runs- that prompt ran 6 times in each cycle, at p=0.251
rename-refactor, which renames a function everywhere it appears, in its definition, its export, and every call site across several files, rose 22.4% on warm runs- that prompt ran 6 times in each cycle, at p=0.465
| family - prompt | warm cost delta % | n (prev/cur) | cost p | blended delta % |
|---|---|---|---|---|
sonnet - fix-off-by-one |
+30.8% | 6/6 | 0.251 | +13.8% |
sonnet - rename-refactor |
+22.4% | 6/6 | 0.465 | +25.1% |
Both sit above the 0.05 line. They also come from nine prompt comparisons run at once. With that many tests, roughly one apparent hit per twenty cells is expected by chance alone. Treat them as something to watch next cycle, not as a result.
Caveats
- Two cost views are reported. The headline is the real per-task spend across every session, including cold-start cache misses, and it is what the dashboard shows. The warm-only view strips out that random cache luck and is the control for whether a change is real. A headline move with no matching warm-only signal may simply be cache luck rather than a genuine change.
- The two cycles being compared ran different CLI versions (2.1.270 then 2.1.271), which is the point of the benchmark: any pooled difference measures the effect of the CLI version. It should not be read as a statement about the models themselves without a separate test that holds the CLI version fixed.
- Individual prompt results come from nine simultaneous comparisons. With that many tests, roughly one apparent hit per twenty cells is expected purely by chance, so a single standout cell is not a finding on its own.
- Token capture for Haiku is known to be incomplete, so its cost and token figures should be read as lower bounds.
- Timings are wall-clock, including the time the model spends thinking, and costs are calculated as API-equivalent amounts rather than dollars actually billed.
Verification appendix
- Verifier verdict: PASS after 3 round(s).
- A first draft was rejected and revised; the objections were resolved.