Claude Benchmarks

Models
Effort

6 lines shown

Suite cost (USD)

  • sonnet-5 (high)
  • fable-5 (high)
  • haiku-4-5
  • opus-4-8 (high)
  • opus-5 (high)
  • fable-5-1 (high)
Suite cost (USD)
Data table
Claude Code CLI versionsonnet-5 (high)fable-5 (high)haiku-4-5opus-4-8 (high)sonnet-5 (medium)opus-5 (high)opus-5 (low)opus-5 (max)opus-5 (medium)opus-5 (xhigh)fable-5 (medium)opus-4-8 (medium)fable-5 (max)fable-5 (xhigh)opus-4-8 (max)opus-4-8 (xhigh)sonnet-5 (max)sonnet-5 (xhigh)fable-5-1 (high)fable-5-1 (medium)fable-5-1 (max)fable-5-1 (xhigh)
2.1.1960.39959
2.1.2020.3074411.5854730.1386590.7877060.292861
2.1.2040.435911.5893450.1369380.74730.423905
2.1.2050.4019361.4129470.1370070.684710.399699
2.1.2060.4158581.4523280.1340950.7118290.411794
2.1.2070.3131611.1691420.1517350.4944310.292131
2.1.2080.3015291.0528040.1372170.4733210.271329
2.1.2100.2823381.0525050.1358790.4788240.261952
2.1.2110.3043121.0425010.1351690.4413840.257755
2.1.2120.3049671.0375620.1318740.4490240.266073
2.1.2140.2982631.003270.1409860.4706040.270236
2.1.2150.2896740.989780.1292310.4033630.265379
2.1.2160.2698610.9324450.1435140.4199730.250677
2.1.2170.293840.9706180.1380430.4054170.256946
2.1.2180.2842980.9935260.1284580.4029940.265344
2.1.2190.2864481.0402090.1328450.4081530.261634
2.1.2200.2768640.9874170.1255540.401210.2644210.6578460.4326750.8150550.5661230.762884
2.1.2210.2860490.9773720.1319990.4016260.655120.873209
2.1.2220.290281.0981880.1319290.4436370.2685450.6466370.562670.8922970.380577
2.1.2230.2960990.9979150.1360630.4195630.263880.6546690.5639481.0194590.417344
2.1.2240.2921221.0852810.1318630.4418560.2649690.6623160.5740120.9313410.397382
2.1.2250.294781.0847270.1367270.4247060.262170.682770.5827470.9453280.392409
2.1.2260.2989911.0595510.1430380.4353420.2621720.6888510.5931550.9245630.423615
2.1.2270.2964321.0745690.1320820.4746790.2600690.725580.5679640.9923870.43579
2.1.2280.3094031.069730.1336590.4221490.2533760.6790350.5737210.9122150.413112
2.1.2290.302241.0693660.1361530.4526090.2597790.6835580.5647810.9091370.424274
2.1.2310.3016371.0460280.1349290.4428980.2613130.7090220.5743530.9463310.407458
2.1.2320.2967411.0461690.1269880.4481570.2633670.6360670.5937760.9346240.429926
2.1.2330.2994941.0393480.1310540.3970760.2635680.6609870.6087250.8930370.423469
2.1.2340.2838711.0497890.1236640.4156980.2677990.6279840.555490.8841930.415903
2.1.2350.289561.0271510.1282070.4508770.2511890.6218020.5333250.8820080.436997
2.1.2360.2875150.9571250.1329530.4545120.2732480.6318730.5165820.9239760.445696
2.1.2370.2959281.0621910.1316690.4293560.2542270.6246280.5129430.9230610.42662
2.1.2380.2739720.9590720.1351260.4184910.266260.6203260.5055440.9027590.421965
2.1.2390.2770271.0058880.1322680.440070.2569250.6684320.5214780.8630740.412674
2.1.2400.2966051.0500340.1294660.4648480.627648
2.1.2410.2848861.0196150.1347390.4329110.2644020.6343960.4950760.891140.440323
2.1.2430.2813471.0657040.1287770.4390550.2587680.6396630.5107950.8950020.418189
2.1.2450.2776181.0786320.1276990.4374610.2751640.628110.5061870.9388240.4392
2.1.2460.2929351.0263440.1283290.4304040.2625550.5987520.5129270.901360.445008
2.1.2470.2746491.0693610.128530.4528960.2620640.6120110.8818190.5096410.7594010.9444050.4315131.4688292.0450820.7101670.5495910.4272450.319708
2.1.2480.2419690.8507890.1160720.3722710.2210590.5377790.8072330.4456080.6747770.7865640.3799492.8427551.9037060.6623010.4578670.3874190.280323
2.1.2500.2433620.9459360.1199910.3729470.226320.552820.8285570.4480030.6961820.7608130.3543012.8294691.8714190.6509260.457870.3540360.281948
2.1.2510.2505430.8707220.1180750.3855710.2351050.533710.7870980.4154210.723680.7844720.3599073.0176971.8570150.6708610.4339310.368420.282411
2.1.2520.250620.8914680.121860.3729190.2360970.5275680.4058260.7754540.361293
2.1.2570.256590.871960.1149440.3804570.2252490.5259180.4390040.7387260.363476
2.1.2580.2471040.8190510.118790.3832420.535461
2.1.2590.2482020.8399040.1180030.3863570.2331960.5268030.4304260.7473310.360691
2.1.2600.2480780.8362410.1174410.3671970.2238460.5544450.4152430.7873310.365427
2.1.2610.2424460.8676810.12130.370430.2223790.5544650.441070.746020.3685070.8704020.748292
2.1.2630.2465180.9285230.120430.3830420.2290810.5809290.4123450.7176380.3665870.8819510.737321
2.1.2650.2518230.8802320.1219110.3864370.2277460.6012060.4291590.8067860.3452410.8407340.751287
2.1.2660.2515990.8250640.1182250.3800740.2265310.5829920.4393950.7902820.3582740.8212630.755636
2.1.2670.2601240.9137930.1238040.3804950.2235340.5845410.4632940.7566740.3590090.8486720.755992
2.1.2680.2475740.8471960.1206010.4025120.2333320.5726960.4679180.748910.3714180.886160.76212
2.1.2690.2475320.8831610.1254960.376190.2258840.5958320.469090.7923190.3670780.8543510.764239
2.1.2700.2480770.8722740.1206720.363450.2177740.6065860.7883270.4498360.693980.7790550.3839921.3029870.9709250.6123350.4617920.3590380.2803610.8854840.7609372.4138161.371101
2.1.2710.2487050.8750750.1170430.3797880.2287410.6073350.4685590.781330.3640440.9201690.755311
highnightly-2.1.271

Data table
Metricfable-5 (high)fable-5-1 (high)haiku-4-5opus-4-8 (high)opus-5 (high)sonnet-5 (high)
Cost per session (USD)mean0.1044020.109840.0137650.0469330.0732750.029335
Cost per session (USD)median0.1110820.1165290.015470.0502160.0658650.031911
Output tokens501580704456570394
Input tokens66626688
Cache-read tokens5008051481620214580965947.5108334
Cache-write tokens5689629887221110830
Input + cache tokens5061552256626904628566913111003
API calls333344
Tool calls222233
Thinking blocks113011
Latency (s, wall-clock)10.0019.399.1988.5199.76957.932
Suite cost (USD)No published column carried Suite cost (USD) for this cycle.
mediumnightly-medium-2.1.271

Data table
Metricfable-5 (medium)fable-5-1 (medium)opus-4-8 (medium)opus-5 (medium)sonnet-5 (medium)
Cost per session (USD)mean0.0986810.0941540.0468690.0603520.027328
Cost per session (USD)median0.1006080.1001560.0527550.057410.031323
Output tokens447448459455356.5
Input tokens666668
Cache-read tokens497284803145809.549203106887
Cache-write tokens2598124321210779
Input + cache tokens504365216746263.549768110992.5
API calls33334
Tool calls22223
Thinking blocks10001
Latency (s, wall-clock)9.1098.18059.4457.8546.9565
Suite cost (USD)No published column carried Suite cost (USD) for this cycle.

Claude Code CLI 2.1.271 Analysis

Measured 2026-09-15 00:46 UTC from this cycle’s benchmark runs, written up from those measurements alone, then checked by an independent reviewer. Where the wording and the tables disagree, the tables are correct.

Verdict

At the high reasoning setting, one benchmark task costs, cheapest to most expensive:

  • haiku-4-5 $0.0138 per session
  • sonnet-5 $0.0293 per session
  • opus-4-8 $0.0469 per session
  • opus-5 $0.0733 per session
  • fable-5 $0.1044 per session
  • fable-5-1 $0.1098 per session

At the medium setting:

  • sonnet-5 $0.0273 per session
  • opus-4-8 $0.0469 per session
  • opus-5 $0.0604 per session
  • fable-5-1 $0.0942 per session
  • fable-5 $0.0987 per session

Across the last 10 cycles at the high setting (9 cycles for fable-5-1):

  • opus-5 drifted up, staying inside $0.0662 to $0.0733
  • fable-5 showed no clear direction, staying inside $0.1018 to $0.1139
  • fable-5-1 showed no clear direction, staying inside $0.1027 to $0.1098
  • haiku-4-5 showed no clear direction, staying inside $0.0138 to $0.0148
  • opus-4-8 showed no clear direction, staying inside $0.0457 to $0.0496
  • sonnet-5 showed no clear direction, staying inside $0.0290 to $0.0314

Across the last 10 cycles at the medium setting (9 cycles for fable-5-1):

  • fable-5 drifted up, staying inside $0.0917 to $0.0987
  • opus-4-8 drifted up, staying inside $0.0435 to $0.0469
  • opus-5 drifted up, staying inside $0.0502 to $0.0604
  • fable-5-1 showed no clear direction, staying inside $0.0926 to $0.0960
  • sonnet-5 showed no clear direction, staying inside $0.0265 to $0.0282

This cycle moved from CLI version 2.1.270 to 2.1.271, and the families with a like-for-like comparison moved as follows:

  • haiku fell 4.0% at the high setting
  • sonnet fell 1.4% at the high setting
  • sonnet rose 2.3% at the medium setting

In all three cases the individual runs behind those averages overlap heavily, so the moves do not stand out from normal run-to-run variation.

One more thing worth knowing: the spread between families is far larger than any of this cycle’s moves. At the high setting, haiku-4-5 costs 87.5% less per task than fable-5-1, and that gap is large enough to stand clear of run-to-run variation.

Which cost figure is which – setting high

The same sessions summarized three ways. They disagree on purpose, and none is the corrected version of another.

  • Mean $/session: multiplies out to real spend, which makes it the budgeting number, and it is what the Cost Trends chart plots.
  • Median $/session: shrugs off an unusually cheap or expensive run, but it does not sum to real spend.
  • Suite cost: the price of one full pass over every prompt in the catalog, the per-prompt medians added up, so it answers what a whole run costs rather than what one session costs.
family models compared mean $/sess mean delta % median $/sess median delta % suite $ suite delta % prompts
haiku haiku-4-5 0.0143 -> 0.0138 -4.0% 0.0144 -> 0.0155 +7.5% 0.1207 -> 0.1170 -3.0% 9
sonnet sonnet-5 0.0297 -> 0.0293 -1.4% 0.0327 -> 0.0319 -2.4% 0.2481 -> 0.2487 +0.3% 9

Suite cost uses only the prompts both cycles ran, so the two sides always add up the same catalog.

Which cost figure is which – setting medium

The same sessions summarized three ways. They disagree on purpose, and none is the corrected version of another.

  • Mean $/session: multiplies out to real spend, which makes it the budgeting number, and it is what the Cost Trends chart plots.
  • Median $/session: shrugs off an unusually cheap or expensive run, but it does not sum to real spend.
  • Suite cost: the price of one full pass over every prompt in the catalog, the per-prompt medians added up, so it answers what a whole run costs rather than what one session costs.
family models compared mean $/sess mean delta % median $/sess median delta % suite $ suite delta % prompts
sonnet sonnet-5 0.0267 -> 0.0273 +2.3% 0.0289 -> 0.0313 +8.2% 0.2178 -> 0.2287 +5.0% 9

Suite cost uses only the prompts both cycles ran, so the two sides always add up the same catalog.

What these terms mean

New to this report? Here is the vocabulary used below, in plain words. Skip ahead if you already know it.

  • Session (also “task”): one complete run of one prompt by one model, start to finish. Every cost in this report is per session.
  • Family: a model line, such as opus or sonnet. One family can cover more than one version, for example opus-4-8 and opus-5.
  • Cache, warm and cold: the tool reuses recent context instead of sending it again, which is far cheaper. A session that gets to reuse it is “warm”; one that has to send everything fresh is “cold”. Whether a given session lands warm or cold is partly luck, and that luck moves the cost.
  • Cache luck (or cache-hit luck): the cost swing caused by how many sessions happened to land warm rather than cold, rather than by any real change in the tool or the model.
  • Blended cost: cost measured across every session, warm and cold alike. This is the real money figure, and it is what the cost feed shows. “Blended” says which sessions are counted, not how they are averaged; see “Which cost figure is which” above for the averaging.
  • Warm-only cost: cost measured across warm sessions only. Removing the cold ones removes most of the luck, so this is the fair test of whether cost actually changed.
  • Pooled: all sessions for a family added together into a single number, rather than split out per prompt.
  • Mean cost per session: the ordinary average across every session. It multiplies out to real spend, so it is the budgeting figure and the one this report leads with.
  • Median cost per session: the middle session once they are lined up by cost. It ignores an unusually cheap or expensive run, but it does not add up to real spend.
  • Suite cost: the price of one full pass over every prompt in the catalog, found by taking each prompt’s median and adding those up. It is a total for the whole catalog, not an average per session, so it is a larger number than either figure above and is not comparable to them.
  • Cell: one family and one prompt combination, for example haiku on the rename-refactor prompt. Cells hold few sessions, so a single cell is noisy.
  • n: how many sessions sit behind a number. A larger n is more trustworthy.
  • p-value, and “significant”: the same prompt run twice costs different amounts, so a p-value asks whether two sets of runs can be told apart at all. Line up every run from both cycles: if one cycle’s runs sit consistently higher, p is near 0 and the difference is called “significant”; if the runs are mixed together, p is near 1 and the two sets look like the same thing measured twice. Precisely, it is the chance of seeing a gap at least this lopsided if the two cycles were truly identical. A high p is not weak evidence of a change, it is no evidence either way.
  • API-equivalent cost: dollars calculated at published API rates. It is not necessarily dollars billed, because a subscription plan may already cover the usage.

Pooled blended comparison

These are the headline figures: the average real spend on one task, counting every session in the cycle, including the unlucky ones that started with a cold cache. Each setting is compared against its own previous run at the same setting. Only families that ran in both cycles at a given setting can be compared.

High reasoning setting, nightly-2.1.270 to nightly-2.1.271:

family models compared n (prev -> cur) prev $/sess cur $/sess delta % Mann-Whitney p
haiku haiku-4-5 53 -> 53 0.0143 0.0138 -4.0% 0.726
sonnet sonnet-5 53 -> 53 0.0297 0.0293 -1.4% 0.922

Medium reasoning setting, nightly-medium-2.1.270 to nightly-medium-2.1.271:

family models compared n (prev -> cur) prev $/sess cur $/sess delta % Mann-Whitney p
sonnet sonnet-5 53 -> 54 0.0267 0.0273 +2.3% 0.616

Pooled warm comparison (significance control)

Whether a cost move is believable is a separate question from how big it looks. This second view keeps only the sessions that ran with a warm cache, which removes the random luck of a cold start, and compares the individual run costs from each cycle against each other. It is the check to trust when asking whether cost actually changed.

High reasoning setting:

family models compared n (prev -> cur) prev $/sess cur $/sess delta % Mann-Whitney p
haiku haiku-4-5 45 -> 45 0.0132 0.0129 -2.7% 0.806
sonnet sonnet-5 45 -> 45 0.0276 0.0271 -1.8% 0.904

Medium reasoning setting:

family models compared n (prev -> cur) prev $/sess cur $/sess delta % Mann-Whitney p
sonnet sonnet-5 45 -> 45 0.0239 0.0250 +4.7% 0.529

Every one of these comparisons sits far above the 0.05 line. That does not mean nothing changed: it means this measurement cannot tell these moves apart from the ordinary variation between runs.

What the reasoning setting costs

The same release was measured twice, once at the high reasoning setting and once at medium. Comparing the two pooled sets of session costs for each model gives the following, with medium measured against high and a negative percentage meaning medium is the cheaper setting.

model high mean $ medium mean $ change p n
fable-5 0.1044 0.0987 -5.5% 0.275 53/55
fable-5-1 0.1098 0.0942 -14.3% 0.141 53/54
opus-4-8 0.0469 0.0469 -0.1% 0.711 53/56
opus-5 0.0733 0.0604 -17.6% 0.165 56/61
sonnet-5 0.0293 0.0273 -6.8% 0.271 53/54

Medium is the cheaper setting in every row, but no row clears the 0.05 line:

  • opus-5 is 17.6% cheaper at medium, and the difference cannot be told apart from run-to-run variation
  • fable-5-1 is 14.3% cheaper at medium, and the difference cannot be told apart from run-to-run variation
  • sonnet-5 is 6.8% cheaper at medium, and the difference cannot be told apart from run-to-run variation
  • fable-5 is 5.5% cheaper at medium, and the difference cannot be told apart from run-to-run variation
  • opus-4-8 is 0.1% cheaper at medium, and the difference cannot be told apart from run-to-run variation

haiku-4-5 was measured at the high setting only, so it has no cross-setting figure.

Notable cells (secondary)

At the high setting, no individual prompt moved beyond the noise threshold. At the medium setting, two sonnet-5 prompts moved more than the rest:

  • fix-off-by-one, which fixes a one-line off-by-one bug in a small Python function, rose 30.8% on warm runs
  • that prompt ran 6 times in each cycle, at p=0.251
  • rename-refactor, which renames a function everywhere it appears, in its definition, its export, and every call site across several files, rose 22.4% on warm runs
  • that prompt ran 6 times in each cycle, at p=0.465
family - prompt warm cost delta % n (prev/cur) cost p blended delta %
sonnet - fix-off-by-one +30.8% 6/6 0.251 +13.8%
sonnet - rename-refactor +22.4% 6/6 0.465 +25.1%

Both sit above the 0.05 line. They also come from nine prompt comparisons run at once. With that many tests, roughly one apparent hit per twenty cells is expected by chance alone. Treat them as something to watch next cycle, not as a result.

Caveats

  • Two cost views are reported. The headline is the real per-task spend across every session, including cold-start cache misses, and it is what the dashboard shows. The warm-only view strips out that random cache luck and is the control for whether a change is real. A headline move with no matching warm-only signal may simply be cache luck rather than a genuine change.
  • The two cycles being compared ran different CLI versions (2.1.270 then 2.1.271), which is the point of the benchmark: any pooled difference measures the effect of the CLI version. It should not be read as a statement about the models themselves without a separate test that holds the CLI version fixed.
  • Individual prompt results come from nine simultaneous comparisons. With that many tests, roughly one apparent hit per twenty cells is expected purely by chance, so a single standout cell is not a finding on its own.
  • Token capture for Haiku is known to be incomplete, so its cost and token figures should be read as lower bounds.
  • Timings are wall-clock, including the time the model spends thinking, and costs are calculated as API-equivalent amounts rather than dollars actually billed.

Verification appendix

  • Verifier verdict: PASS after 3 round(s).
  • A first draft was rejected and revised; the objections were resolved.