Codex Benchmarks
Suite cost (USD)
- sol (high)
- terra (high)
- luna (high)
Data table
| Codex CLI version | sol (medium) | terra (medium) | luna (medium) | sol (high) | terra (high) | luna (high) |
|---|---|---|---|---|---|---|
| 0.145.0 | 0.600989 | 0.210022 | 0.022581 | 0.655028 | 0.217918 | 0.023138 |
| 0.146.0 | 0.679805 | 0.224578 | 0.025403 | 0.655652 | 0.234645 | 0.025814 |
| 0.146.1 | 0.630408 | 0.220089 | 0.023329 | 0.66662 | 0.246495 | 0.0262 |
| 0.147.0 | 0.597284 | 0.219957 | 0.022157 | 0.639826 | 0.235122 | 0.026121 |
| 0.148.0 | 0.690976 | 0.227794 | 0.029381 | 0.791735 | 0.245905 | 0.031325 |
| 0.149.0 | 0.66847 | 0.223742 | 0.027309 | 0.689255 | 0.242489 | 0.030541 |
| 0.149.1 | 0.725717 | 0.224333 | 0.02747 | 0.728893 | 0.238071 | 0.032493 |
| 0.150.1 | 0.594659 | 0.215502 | 0.02433 | 0.617141 | 0.237537 | 0.029076 |
| 0.151.0 | 0.65808 | 0.20997 | 0.024788 | 0.660714 | 0.229257 | 0.028269 |
| 0.152.0 | 0.734977 | 0.213057 | 0.024017 | 0.51986 | 0.23388 | 0.028935 |
| 0.152.1 | 0.473833 | 0.208794 | 0.024961 | 0.528608 | 0.234578 | 0.029412 |
| 0.153.0 | 0.424839 | 0.212104 | 0.024865 | 0.481455 | 0.244619 | 0.02883 |
| 0.153.1 | 0.450947 | 0.210881 | 0.025044 | 0.500448 | 0.223639 | 0.027293 |
| 0.153.2 | 0.473928 | 0.217671 | 0.027764 | 0.515561 | 0.222857 | 0.02944 |
| 0.153.3 | 0.461096 | 0.182929 | — | — | — | — |
| 0.153.4 | 0.439336 | 0.202371 | 0.025556 | 0.471782 | 0.217674 | 0.028716 |
| 0.154.0 | 0.45782 | 0.20759 | 0.027473 | 0.495271 | 0.233061 | 0.029482 |
Data table
| Metric | luna (high) | sol (high) | terra (high) |
|---|---|---|---|
| Cost per session (USD)mean | 0.003205 | 0.055728 | 0.027072 |
| Cost per session (USD)median | 0.003023 | 0.052631 | 0.028072 |
| Output tokens | 551 | 459.5 | 416.5 |
| Input tokens | 5675 | 3689.5 | 4353.5 |
| Cache-read tokens | 47104 | 57408 | 55808 |
| Cache-write tokens | 0 | 0 | 0 |
| Input + cache tokens | 52779 | 61097.5 | 60161.5 |
| API calls | 1 | 1 | 1 |
| Tool calls | 2 | 2 | 2 |
| Latency (s, wall-clock) | 17.6045 | 27.041 | 17.0865 |
| Reasoning tokens | 51.5 | 18.5 | 14 |
| Suite cost (USD) | No published column carried Suite cost (USD) for this cycle. | ||
Data table
| Metric | luna (medium) | sol (medium) | terra (medium) |
|---|---|---|---|
| Cost per session (USD)mean | 0.003027 | 0.052276 | 0.023337 |
| Cost per session (USD)median | 0.002642 | 0.054674 | 0.0244 |
| Output tokens | 420.5 | 402 | 405.5 |
| Input tokens | 5073.5 | 3968 | 3583.5 |
| Cache-read tokens | 37120 | 54464 | 56320 |
| Cache-write tokens | 0 | 0 | 0 |
| Input + cache tokens | 42193.5 | 58432 | 59903.5 |
| API calls | 1 | 1 | 1 |
| Tool calls | 2 | 2 | 2 |
| Latency (s, wall-clock) | 17.158 | 25.136 | 16.7085 |
| Reasoning tokens | 31 | 15.5 | 11 |
| Suite cost (USD) | No published column carried Suite cost (USD) for this cycle. | ||
Codex CLI 0.154.0 Analysis
Measured 2026-09-10 14:46 UTC from the benchmark runs this release recorded, written up from those measurements alone, then checked by an independent reviewer. Where the wording and the tables disagree, the tables are correct.
Verdict
At high reasoning, one task costs:
- luna: $0.0032 per session
- terra: $0.0271 per session
- sol: $0.0557 per session
At medium reasoning, one task costs:
- luna: $0.0030 per session
- terra: $0.0233 per session
- sol: $0.0523 per session
Across the last 10 cycles at high reasoning:
- luna: no clear direction, inside $0.0031 to $0.0036
- sol: drifted down, inside $0.0528 to $0.0834
- terra: no clear direction, inside $0.0249 to $0.0272
Across the last 10 cycles at medium reasoning:
- luna: no clear direction, inside $0.0026 to $0.0031
- sol: no clear direction, inside $0.0490 to $0.0804
- terra: no clear direction, inside $0.0225 to $0.0247
Against the previous CLI version, headline spend per session moved this way:
- luna at high reasoning rose 2.1%
- sol at high reasoning rose 5.6%
- terra at high reasoning rose 8.5%
- luna at medium reasoning rose 5.8%
- sol at medium reasoning rose 1.9%
- terra at medium reasoning fell 2.3%
Comparing the individual session costs of the two cycles, none of these six moves can be told apart from ordinary run-to-run variation; and that test is not controlled for cache state, because this study cannot tell which sessions started with a warm cache, which makes it weaker evidence than the same test run on warm sessions alone.
One more thing worth knowing: sol is the dearest model at both settings, and the other two sit well under it:
- luna at high reasoning costs 94.2% less than sol
- terra at high reasoning costs 51.4% less than sol
- luna at medium reasoning costs 94.2% less than sol
- terra at medium reasoning costs 55.4% less than sol
Which cost figure is which
The same sessions summarized three ways. They disagree on purpose, and none is the corrected version of another.
- Mean $/session: multiplies out to real spend, which makes it the budgeting number, and it is what the cost chart plots.
- Median $/session: shrugs off an unusually cheap or expensive run, but it does not sum to real spend.
- Suite cost: the price of one full pass over every prompt in the catalog, the per-prompt medians added up, so it answers what a whole run costs rather than what one session costs.
What these terms mean
New to this report? Here is the vocabulary used below, in plain words. Skip ahead if you already know it.
- Session (also “task”): one complete run of one prompt by one model, start to finish. Every cost in this report is per session.
- Reasoning setting: how much internal reasoning the model is asked to spend before answering. This benchmark runs each setting separately and reports them separately, because the same model at a different setting is a different cost.
- Mean cost per session: the ordinary average across every session. It multiplies out to real spend, so it is the budgeting figure and the one this report leads with.
- Median cost per session: the middle session once they are lined up by cost. It ignores an unusually cheap or expensive run, but it does not add up to real spend.
- Suite cost: the price of one full pass over every prompt in the catalog, found by taking each prompt’s median and adding those up. It is a total for the whole catalog, not an average per session, so it is a larger number than either figure above and is not comparable to them.
- Cache, and cache-read tokens: the tool reuses recent context instead of sending it again, which is far cheaper than sending it fresh. Most of what a session sends is reused this way, so how much of a session lands in the cache moves its cost.
- n: how many sessions sit behind a number. A larger n is more trustworthy.
- p-value, and “significant”: the same prompt run twice costs different amounts, so a p-value asks whether two sets of runs can be told apart at all. Line up every run from both cycles: if one cycle’s runs sit consistently higher, p is near 0 and the difference is called “significant”; if the runs are mixed together, p is near 1 and the two sets look like the same thing measured twice. A high p is not weak evidence of a change, it is no evidence either way.
- API-equivalent cost: dollars calculated at published API rates. It is not necessarily dollars billed, because a subscription plan may already cover the usage.
What a task costs
The table below is this release’s measured spend, one row per model and reasoning setting, with 90 sessions behind each row. The headline figure is the mean dollars per session.
| model | reasoning | n | mean $/sess | median $/sess | suite $ |
|---|---|---|---|---|---|
| luna | high | 90 | 0.0032 | 0.0030 | 0.0295 |
| sol | high | 90 | 0.0557 | 0.0526 | 0.4953 |
| terra | high | 90 | 0.0271 | 0.0281 | 0.2331 |
| luna | medium | 90 | 0.0030 | 0.0026 | 0.0275 |
| sol | medium | 90 | 0.0523 | 0.0547 | 0.4578 |
| terra | medium | 90 | 0.0233 | 0.0244 | 0.2076 |
Compared with the previous version
The version before this one in the published order is 0.153.4. Each row compares the same model at the same reasoning setting across the two cycles, with 90 sessions on each side.
| model | reasoning | prev CLI | mean $/sess | mean delta % | median $/sess | median delta % | suite $ | suite delta % | n (prev -> cur) | Mann-Whitney p |
|---|---|---|---|---|---|---|---|---|---|---|
| luna | high | 0.153.4 | 0.0031 -> 0.0032 | +2.1% | 0.0028 -> 0.0030 | +8.7% | 0.0287 -> 0.0295 | +2.7% | 90 -> 90 | 0.591 |
| sol | high | 0.153.4 | 0.0528 -> 0.0557 | +5.6% | 0.0542 -> 0.0526 | -3.0% | 0.4718 -> 0.4953 | +5.0% | 90 -> 90 | 0.589 |
| terra | high | 0.153.4 | 0.0249 -> 0.0271 | +8.5% | 0.0263 -> 0.0281 | +6.9% | 0.2177 -> 0.2331 | +7.1% | 90 -> 90 | 0.317 |
| luna | medium | 0.153.4 | 0.0029 -> 0.0030 | +5.8% | 0.0026 -> 0.0026 | +0.6% | 0.0256 -> 0.0275 | +7.5% | 90 -> 90 | 0.544 |
| sol | medium | 0.153.4 | 0.0513 -> 0.0523 | +1.9% | 0.0507 -> 0.0547 | +7.8% | 0.4393 -> 0.4578 | +4.2% | 90 -> 90 | 0.702 |
| terra | medium | 0.153.4 | 0.0239 -> 0.0233 | -2.3% | 0.0238 -> 0.0244 | +2.5% | 0.2024 -> 0.2076 | +2.6% | 90 -> 90 | 0.832 |
Every p-value here sits above the 0.05 line, the closest being terra at high reasoning (p=0.317). Comparing the individual session costs of the two cycles, none of these moves can be told apart from ordinary run-to-run variation. That is not the same as saying nothing changed: something may well have, and this measurement cannot see it. Note also that the direction of a model’s mean and its median do not always agree, as with sol at high reasoning, where the mean rose and the median fell.
What the reasoning setting costs
Asking for high reasoning instead of medium costs more for all three models. The table gives the increase for each.
| model | medium $/sess | high $/sess | high vs medium % |
|---|---|---|---|
| luna | 0.0030 | 0.0032 | +5.9% |
| sol | 0.0523 | 0.0557 | +6.6% |
| terra | 0.0233 | 0.0271 | +16.0% |
Caveats
- These dollar figures are API-equivalent costs worked out at published list prices, not money that was actually billed. If you use a subscription plan, the same usage may already be covered by it.
- Models are priced at their standard published rates. Where a promotional rate is running today, it is deliberately not applied, because repricing the whole history at a temporary rate would make every past cycle look cheaper than it was when measured.
- Whether a move between versions is bigger than ordinary run-to-run variation is judged by comparing the individual session costs of the two cycles. That comparison is not controlled for cache state: it includes every session, because this study cannot tell which ones started with a warm cache. Cache luck therefore sits on both sides of every comparison, so treat the result as weaker evidence than a test that had removed it.
- The two cycles being compared ran different CLI versions, so any move captures the effect of the version change bundled with whatever else differed on the day. It is not a statement about the models themselves.
- High and medium reasoning are reported separately and are not interchangeable. The same model at a different setting is a different cost, and a figure from one setting should never be read as standing in for the other.
- Reported timings are wall-clock and include the time the model spends on its own reasoning.
- These runs use their own fixed set of prompts. The study is built the same way as the Claude benchmark, but the prompts are different prompts, so the two sets of results should not be read as a like for like comparison.
Verification appendix
- Verifier verdict: PASS after 4 round(s).
- A first draft was rejected and revised; the objections were resolved.