Highlights
Model Analytics Highlights
Distilled 2026-09-15 00:51 UTC from the published section write-ups alone, then checked by an independent reviewer. Each section’s own page carries the detail and the tables behind these figures.
What this shows
The Model Analytics area tracks what it costs to get a unit of work done by an AI coding tool, and whether that cost is moving. Each section runs its own fixed catalog of prompts repeatedly against one command line tool, prices every run at the provider’s published API rates, and compares the current release against the one before it. Open the section for the tool you care about; each one carries its own figures and its own limits.
Claude Benchmarks
This section measures what one benchmark task costs per session on the Claude Code CLI, at a high and a medium reasoning setting, priced at published API rates. At the high reasoning setting one task costs, cheapest to most expensive:
- haiku-4-5 $0.0138 per session
- sonnet-5 $0.0293 per session
- opus-4-8 $0.0469 per session
- opus-5 $0.0733 per session
- fable-5 $0.1044 per session
- fable-5-1 $0.1098 per session
At the medium setting:
- sonnet-5 $0.0273 per session
- opus-4-8 $0.0469 per session
- opus-5 $0.0604 per session
- fable-5-1 $0.0942 per session
- fable-5 $0.0987 per session
Across the last 10 cycles at the high setting (9 cycles for fable-5-1), the section reports each family separately:
- opus-5 drifted up, staying inside $0.0662 to $0.0733
- fable-5 showed no clear direction, staying inside $0.1018 to $0.1139
- fable-5-1 showed no clear direction, staying inside $0.1027 to $0.1098
- haiku-4-5 showed no clear direction, staying inside $0.0138 to $0.0148
- opus-4-8 showed no clear direction, staying inside $0.0457 to $0.0496
- sonnet-5 showed no clear direction, staying inside $0.0290 to $0.0314
Across the last 10 cycles at the medium setting (9 cycles for fable-5-1):
- fable-5 drifted up, staying inside $0.0917 to $0.0987
- opus-4-8 drifted up, staying inside $0.0435 to $0.0469
- opus-5 drifted up, staying inside $0.0502 to $0.0604
- fable-5-1 showed no clear direction, staying inside $0.0926 to $0.0960
- sonnet-5 showed no clear direction, staying inside $0.0265 to $0.0282
This cycle moved from CLI version 2.1.270 to 2.1.271, and the families with a like-for-like comparison moved as follows:
- haiku fell 4.0% at the high setting
- sonnet fell 1.4% at the high setting
- sonnet rose 2.3% at the medium setting
In all three cases the individual runs behind those averages overlap heavily, so the moves do not stand out from normal run-to-run variation. The section adds that the spread between families is far larger than any of this cycle’s moves: at the high setting haiku-4-5 costs 87.5% less per task than fable-5-1, a gap large enough to stand clear of run-to-run variation. Its caveats note that token capture for Haiku is incomplete, so its figures are lower bounds, and that a pooled difference between the two cycles measures the effect of the CLI version, so it should not be read as a statement about the models themselves without a separate test that holds the CLI version fixed.
Codex Benchmarks
This section measures what one task costs per session on the Codex CLI, reported separately at high and medium reasoning and priced at OpenAI’s published list rates. At high reasoning one task costs:
- luna $0.0032 per session
- terra $0.0271 per session
- sol $0.0557 per session
At medium reasoning one task costs:
- luna $0.0030 per session
- terra $0.0233 per session
- sol $0.0523 per session
Across the last 10 cycles at high reasoning, the section reports each model separately:
- sol drifted down, inside $0.0528 to $0.0834
- luna showed no clear direction, inside $0.0031 to $0.0036
- terra showed no clear direction, inside $0.0249 to $0.0272
Across the last 10 cycles at medium reasoning:
- luna showed no clear direction, inside $0.0026 to $0.0031
- sol showed no clear direction, inside $0.0490 to $0.0804
- terra showed no clear direction, inside $0.0225 to $0.0247
Against the previous CLI version, headline spend per session moved this way:
- luna at high reasoning rose 2.1%
- sol at high reasoning rose 5.6%
- terra at high reasoning rose 8.5%
- luna at medium reasoning rose 5.8%
- sol at medium reasoning rose 1.9%
- terra at medium reasoning fell 2.3%
Comparing the individual session costs of the two cycles, none of these six moves can be told apart from ordinary run-to-run variation, and that test is not controlled for cache state, because this study cannot tell which sessions started with a warm cache, which makes it weaker evidence than the same test run on warm sessions alone. The section also reports that sol is the dearest model at both settings, with the other two sitting well under it:
- luna at high reasoning costs 94.2% less than sol
- terra at high reasoning costs 51.4% less than sol
- luna at medium reasoning costs 94.2% less than sol
- terra at medium reasoning costs 55.4% less than sol
It stresses that high and medium reasoning are not interchangeable, and that its prompts are different prompts from the Claude benchmark’s, so the two sets of results should not be read as a like for like comparison.
Verification appendix
- Verifier verdict: PASS after 4 round(s).
- A first draft was rejected and revised; the objections were resolved.