CodexvsClaude Code

Compare coding benchmark results, model API rates, context windows, and subscription costs for Codex and Claude Code.

Updated October 1, 2026

Model benchmarks

Terminal-Bench 4.0 tests Codex + GPT-6 Astra and Claude Code + Fable 5.1 on the same 330 trials. Scores reflect both the model and harness, at five effort levels.

Codex + GPT-6 AstraClaude Code + Fable 5.1
Terminal-Bench 4.0 resolution rate against run cost. The data is also in the table below.0%25%50%75%100%$0$2k$4k$6klowmaxlowmax
Each line is one harness running one model, so it measures the pair together, not the harness or the model alone. The chart plots first-attempt resolution rate on a full 0–100% scale against the leaderboard’s total cost for the 330-trial run, at five effort levels. Whiskers show the margin of error.[36]
  • At xhigh effort, both pairs solve 57.9% on the first attempt; their margins of error overlap.[36]
  • Codex + GPT-6 Astra costs $2,351 for the xhigh run; Claude Code + Fable 5.1 costs $4,872.[36]
  • With five attempts, Claude Code + Fable 5.1 solves 78.8% against 71.2% at max effort.[36]
Terminal-Bench 4.0 results by effort level
EffortCodexGPT-6 AstraClaude CodeFable 5.1
First tryPass@5Run costFirst tryPass@5Run cost
low50.6% ±2.863.6%$1,55743.3% ±3.665.2%$2,359
medium54.2% ±2.766.7%$1,91553.9% ±3.471.2%$2,833
high57.9% ±3.071.2%$2,26954.5% ±3.475.8%$3,985
xhigh57.9% ±2.769.7%$2,35157.9% ±3.475.8%$4,872
max58.2% ±2.871.2%$3,26757.9% ±3.878.8%$6,244
  • Neither pair is the everyday configuration. Claude Code defaults to Opus 5.5, OpenAI recommends GPT-6.1 Sol for Codex, and neither has a leaderboard entry yet.[24][2][36]
  • Anthropic’s launch posts report higher Terminal-Bench 4.0 scores for Opus 5.5 (66.4% at xhigh effort) and Sonnet 5.5 (70.6%). Those are Anthropic’s own runs, not leaderboard entries. Anthropic lists no Terminal-Bench 4.0 score for GPT-6 Sol, and quotes GPT-6 Astra at 57.9% as reported by OpenAI.[37][38]
  • Terminal tasks aren’t your codebase. OpenAI itself warns against drawing frontier conclusions from SWE-bench Verified, so we don’t cite it.[39]

DeepSWE v1.1

DeepSWE tests 28 models on 113 tasks using one mini-swe-agent harness. Each point shows a model’s highest-scoring effort setting, including seven models unchecked by default on the source site—not Codex versus Claude Code.[40]

OpenAIAnthropicOther models
DeepSWE v1.1 resolution rate against average cost per task for 28 models. Values also appear in the table below.0%25%50%75%100%$0$5$10$15$20$25$30
Each point shows first-attempt resolution against average cost per task on a 0–100% scale. Whiskers show the leaderboard’s margin of error. The table lists every model and its effort setting.[40]
DeepSWE v1.1 results for all 28 models at their highest-scoring effort settings
ModelEffort
GPT-6 Astraxhigh74% ±3%$4.43
Gemini 3.8 Flashhigh74% ±1%$2.36
Claude Opus 5max74% ±4%$11.84
GPT-5.6 Solmax73% ±3%$6.46
Claude Fable 5xhigh70% ±3%$13.41
GPT-5.6 Terramax70% ±3%$3.96
GLM-5.3max69% ±3%$3.99
Kimi K3max69% ±5%$4.65
Grok 4.6medium67% ±2%$3.45
GPT-5.6 Lunamax67% ±4%$0.61
GPT-5.5xhigh67% ±6%$7.23
Gemini 3.7 Flashmedium65% ±3%$2.03
GLM-5.3 Flashmax63% ±4%$0.24
DeepSeek V4 Promax63% ±6%$1.67
Claude Opus 4.8max59% ±2%$13.22
Qwen 3.8 Maxxhigh57% ±3%$3.73
Muse Spark 1.2xhigh55% ±2%$3.70
Claude Sonnet 5max54% ±4%$26.40
Grok 4.5high54% ±2%$2.42
DeepSeek V4 Flashmax53% ±4%$0.46
Muse Spark 1.1xhigh53% ±3%$2.36
GPT-5.4xhigh52% ±2%$5.65
Gemini 3.6 Flashhigh47% ±4%$2.21
GLM-5.2max44% ±2%$3.92
Gemini 3.5 Flashhigh36% ±4%$3.45
Kimi K2.7 Codedefault31% ±1%$2.82
Claude Sonnet 4.6high30% ±4%$5.52
Gemini 3.1 Pro Previewhigh12% ±1%$2.14

Leaderboard snapshot: September 22, 2026. These effort settings are not necessarily the models’ defaults.[40]

Model and subscription pricing

Usage estimates and multipliers are relative to each vendor's own plan, not a common token allowance. Real cost per completed task depends on model and harness token use.

Individual

$0

Codex

Free[1]

GPT-6 Luna at standard speed in the desktop app, subject to rollout.

Claude Code

Claude Free doesn’t include Claude Code.

$8

Codex

Go[1]

GPT-6 Luna at standard speed in the desktop app, subject to rollout.

Claude Code

No listed plan at this price

$20

Codex

Plus[1]

~15–160 GPT-6.1 Sol local messages / 5h

Estimate, not a fixed limit. Shared with ChatGPT Work; weekly limits may apply.

Claude Code

Pro[20][28]

Standard Pro allowance

Shared with Claude chat; $17/mo billed annually.

$100

Codex

Pro $100[1]

No five-hour limit

Weekly limits may apply; no fixed token allowance published.

Claude Code

Max 5x[21]

5× Pro per-session allowance

Five-hour resets; weekly limit applies.

$200

Codex

Pro $200[1]

Higher usage; no fixed multiplier published

No five-hour limit; weekly limits may apply.

Claude Code

Max 20x[21]

20× Pro per-session allowance

Five-hour resets; weekly limit applies.

$500

Codex

Pro $500[1]

No five-hour limit

Adds GPT-6 Astra Ultrafast; weekly limits may apply.

Claude Code

No listed plan at this price

Teams, per seat

$20Annual billing; $25 monthly

Codex

Business[1]

At least 2 users. Includes SAML SSO, MFA, larger cloud VMs, and no training on business data by default.

Claude Code

Team standard[20]

2 to 150 seats. Includes Claude Code, SSO, central billing, and no training on your content by default.

$100Annual billing; $125 monthly

Codex

No listed plan at this price

Claude Code

Team premium[20]

5× standard seat usage

You can mix seat types.

Custom

Codex

Enterprise & Edu[1]

Contact sales. Adds SCIM, EKM, RBAC, Compliance API audit logs, and retention and residency controls.

Claude Code

Enterprise[20]

$20 per seat billed annually, plus usage at API rates. Adds SCIM, audit logs, IP allowlisting, and a HIPAA-ready option.

What each vendor publishes about usage

These numbers measure different things, one a message allowance and the other an observed spend, so don’t compare them directly. They’re starting points for planning usage, not a cost-per-task comparison.

Codex

15–160

local messages per 5 hours on Plus with GPT-6.1 Sol[1]

This is OpenAI’s estimate, and GPT-6 Luna gets 350–3,000. Cloud tasks can use more of the allowance than local messages, and weekly limits may also apply.

Claude Code

~$13

per developer per active day[22]

This is Anthropic’s average across enterprise deployments: $150–250 per developer per month, and under $30 per active day for 90% of users.

Prices are in US dollars per million tokens. API use is billed separately from subscriptions; token rates alone do not predict the cost of completing a task.

ModelContextInput → outputLong prompts
GPT-6 Luna[18]$0.10$0.01$0.501.05MOver 272K input: 2× input, 1.5× output
Claude Haiku 4.5[34][35]$1.00$0.10$5.00200KNo surcharge
Claude Sonnet 5.5[34][35]$2.00$0.20$10.001MStandard rate across the full window
GPT-6.1 Sol[17][2]$2.00$0.10$10.001.05MOver 272K input: 2× input, 1.5× output
Claude Opus 5.5[34][24][35]$4.00$0.20$20.001MStandard rate across the full window
Claude Fable 5.1[34][24][35]$10.00$0.25$50.001MStandard rate across the full window
GPT-6 Astra[19]$10.00$1.00$50.001.05MOver 272K input: 2× input, 1.5× output

Estimate a month

Illustrative API cost at the same token counts, not cost per task.

Lowest estimate: GPT-6 Luna at $4.30

  1. GPT-6 Luna$4.30
  2. Claude Haiku 4.5$43.00
  3. GPT-6.1 Sol$78.00
  4. Claude Sonnet 5.5$86.00
  5. Claude Opus 5.5$156.00
  6. Claude Fable 5.1$370.00
  7. GPT-6 Astra$430.00

This assumes no single request goes over 272K input tokens. OpenAI bills any request that does at 2× input and 1.5× output for that whole request. A monthly total can’t show whether that happened, so check your per-request prompt sizes. Anthropic’s 1M-context models don’t charge extra for long prompts.

These are list prices at standard speed. The estimate leaves out cache writes (OpenAI: 1.25× input; Anthropic: 1.25× for five-minute or 2× for one-hour caching), batch discounts, regional pricing, and tax. Agents can use different token counts for the same task.

Recommendations by task

Building web interfaces

Claude Code

#1 on Arena WebDev

Opus 5.5 max scored 1,818 against GPT-6 Astra max at 1,789 in Arena’s September 30 human-preference ranking. This measures frontend model output, not Claude Code versus Codex as harnesses.[41]

Terminal-heavy engineering

Codex

57.9% solved at less than half the run cost

Codex + GPT-6 Astra and Claude Code + Fable 5.1 both solved 57.9% on the first attempt at xhigh effort. Their full Terminal-Bench runs cost $2,351 and $4,872 respectively. Neither is the everyday model.[36][2][24]

Lower-cost general tasks

Codex

$0.72 vs $5.98 per index task

At max effort, Artificial Analysis scores GPT-6.1 Sol 52 and Opus 5.5 58 on its ten-evaluation Intelligence Index. Average API cost per index task was $0.72 and $5.98 respectively. These are models in its test setup, not cost per PR.[42]

Features and integrations

Both agents edit files, run commands, and support MCP, skills, hooks, subagents, and plugins.

SupportedPartial or with caveatsNot supported

Openness and inference

Open-source agent[3][4][25]

CodexCLI, SDK, app server

Claude CodeProprietary

Amazon Bedrock[14][33]

CodexBuilt-in provider

Claude CodeSupported

Azure OpenAI / Foundry[13][15][33]

CodexLocal configured provider

Claude CodeSupported

Google Cloud[12][33]

CodexNo direct setup documented

Claude CodeSupported

Instruction files[11][32]

CodexAGENTS.md

Claude CodeCLAUDE.md; AGENTS.md fallback

Safety defaults

Local sandbox[5][26]

CodexOn by default; no network

Claude CodeOpt-in; can fail open

Default approvals[5][27]

CodexAsk for network or outside writes

Claude CodeClassifier reviews actions

Windows sandbox[6][26]

CodexNative

Claude CodeWSL2 only

Hosted and automated work

Cloud tasks[7][1][28][23]

CodexChatGPT plan required

Claude CodePro, Max, Team; self-hosting for teams

Scheduled and triggered runs[8][29]

CodexApp schedules; cloud triggers

Claude CodeCloud routines (preview)

Managed PR review[9][1][30][31]

CodexIncluded from Plus

Claude CodeTeam / Enterprise preview; $15–25/review

Plans and billing

Free access[1][20]

CodexLimited free access

Claude CodeStarts at $20/mo

Enterprise usage[1][20]

CodexCredits or Plus-level limits

Claude Code$20/seat + API usage

How to compare both on your codebase

You can compare both on the same repository, with shared instructions and acceptance criteria. Here’s how we’d run a fair comparison on your own code.

  1. 01

    Share one instruction file

    Put your conventions in AGENTS.md. Codex reads it natively, and Claude Code reads it when there’s no CLAUDE.md, or alongside one if you turn that setting on.[11][32]

  2. 02

    Pick real tasks

    Choose three to five closed issues with known-good fixes, such as a bug, a small feature, and a refactor. Give both tools the same acceptance criteria.

  3. 03

    Start from the same state

    Run each tool from the same commit, with the same dependencies and a clean environment, so neither benefits from leftover state.

  4. 04

    Compare the pull requests

    Judge the diffs, the test results, and how long review took. Record what each run used from your plan, or what it cost in API tokens.

Common questions

Sources and methodology

We use vendor documentation and pricing, plus Terminal-Bench, DeepSWE, Arena, and Artificial Analysis results. We checked these sources on October 1, 2026. Prices and plans change often, so confirm with the vendor before you buy.

The recommendations are our interpretation of published measurements, not an overall score for either tool. Model-only tests don't establish which harness will perform better on your repository.

Replicas publishes this comparison and supports both harnesses. If you spot an error, email founders@replicas.dev and we'll fix it.

  1. 1Codex pricing and plan features · OpenAI
  2. 2Codex models · OpenAI
  3. 3Codex open-source components · OpenAI
  4. 4openai/codex repository and license · GitHub
  5. 5Codex agent approvals and security · OpenAI
  6. 6Codex Windows sandbox · OpenAI
  7. 7Codex cloud environments · OpenAI
  8. 8Codex scheduled tasks · OpenAI
  9. 9Codex code review in GitHub · OpenAI
  10. 10Use Codex in Linear · OpenAI
  11. 11Codex AGENTS.md instructions · OpenAI
  12. 12Codex configuration reference · OpenAI
  13. 13Codex Azure model provider configuration · OpenAI
  14. 14Codex with Amazon Bedrock · OpenAI
  15. 15Codex with Azure OpenAI in Microsoft Foundry · Microsoft
  16. 16OpenAI API pricing · OpenAI
  17. 17GPT-6.1 Sol model specifications · OpenAI
  18. 18GPT-6 Luna model specifications · OpenAI
  19. 19GPT-6 Astra model specifications · OpenAI
  20. 20Claude plans and pricing · Anthropic
  21. 21Claude Max plan tiers and pricing · Anthropic
  22. 22Manage Claude Code costs · Anthropic
  23. 23Claude Code platforms and integrations · Anthropic
  24. 24Claude Code model configuration · Anthropic
  25. 25anthropics/claude-code license · GitHub
  26. 26Claude Code sandboxing · Anthropic
  27. 27Claude Code permission modes · Anthropic
  28. 28Use Claude Code in the cloud · Anthropic
  29. 29Claude Code routines · Anthropic
  30. 30Claude Code Review · Anthropic
  31. 31Claude Code GitHub Actions · Anthropic
  32. 32Claude Code memory and CLAUDE.md · Anthropic
  33. 33Claude Code on third-party platforms · Anthropic
  34. 34Claude API pricing · Anthropic
  35. 35Claude models overview · Anthropic
  36. 36Terminal-Bench 4.0 leaderboard · Terminal-Bench
  37. 37Introducing Claude Opus 5.5 · Anthropic
  38. 38Introducing Claude Sonnet 5.5 · Anthropic
  39. 39Separating signal from noise in coding evaluations · OpenAI
  40. 40DeepSWE v1.1 leaderboard · Datacurve
  41. 41Code Arena WebDev leaderboard · Arena
  42. 42GPT-6.1 Sol versus Claude Opus 5.5 model comparison · Artificial Analysis