CodexvsClaude Code
Model benchmarks
Terminal-Bench 4.0 tests Codex + GPT-6 Astra and Claude Code + Fable 5.1 on the same 330 trials. Scores reflect both the model and harness, at five effort levels.
| Effort | CodexGPT-6 Astra | Claude CodeFable 5.1 | ||||
|---|---|---|---|---|---|---|
| First try | Pass@5 | Run cost | First try | Pass@5 | Run cost | |
| low | 50.6% ±2.8 | 63.6% | $1,557 | 43.3% ±3.6 | 65.2% | $2,359 |
| medium | 54.2% ±2.7 | 66.7% | $1,915 | 53.9% ±3.4 | 71.2% | $2,833 |
| high | 57.9% ±3.0 | 71.2% | $2,269 | 54.5% ±3.4 | 75.8% | $3,985 |
| xhigh | 57.9% ±2.7 | 69.7% | $2,351 | 57.9% ±3.4 | 75.8% | $4,872 |
| max | 58.2% ±2.8 | 71.2% | $3,267 | 57.9% ±3.8 | 78.8% | $6,244 |
- Neither pair is the everyday configuration. Claude Code defaults to Opus 5.5, OpenAI recommends GPT-6.1 Sol for Codex, and neither has a leaderboard entry yet.[24][2][36]
- Anthropic’s launch posts report higher Terminal-Bench 4.0 scores for Opus 5.5 (66.4% at xhigh effort) and Sonnet 5.5 (70.6%). Those are Anthropic’s own runs, not leaderboard entries. Anthropic lists no Terminal-Bench 4.0 score for GPT-6 Sol, and quotes GPT-6 Astra at 57.9% as reported by OpenAI.[37][38]
- Terminal tasks aren’t your codebase. OpenAI itself warns against drawing frontier conclusions from SWE-bench Verified, so we don’t cite it.[39]
DeepSWE v1.1
DeepSWE tests 28 models on 113 tasks using one mini-swe-agent harness. Each point shows a model’s highest-scoring effort setting, including seven models unchecked by default on the source site—not Codex versus Claude Code.[40]
| Model | Effort | ||
|---|---|---|---|
| GPT-6 Astra | xhigh | 74% ±3% | $4.43 |
| high | 74% ±1% | $2.36 | |
| Claude Opus 5 | max | 74% ±4% | $11.84 |
| GPT-5.6 Sol | max | 73% ±3% | $6.46 |
| Claude Fable 5 | xhigh | 70% ±3% | $13.41 |
| GPT-5.6 Terra | max | 70% ±3% | $3.96 |
| max | 69% ±3% | $3.99 | |
| max | 69% ±5% | $4.65 | |
| medium | 67% ±2% | $3.45 | |
| GPT-5.6 Luna | max | 67% ±4% | $0.61 |
| GPT-5.5 | xhigh | 67% ±6% | $7.23 |
| medium | 65% ±3% | $2.03 | |
| max | 63% ±4% | $0.24 | |
| max | 63% ±6% | $1.67 | |
| Claude Opus 4.8 | max | 59% ±2% | $13.22 |
| xhigh | 57% ±3% | $3.73 | |
| xhigh | 55% ±2% | $3.70 | |
| Claude Sonnet 5 | max | 54% ±4% | $26.40 |
| high | 54% ±2% | $2.42 | |
| max | 53% ±4% | $0.46 | |
| xhigh | 53% ±3% | $2.36 | |
| GPT-5.4 | xhigh | 52% ±2% | $5.65 |
| high | 47% ±4% | $2.21 | |
| max | 44% ±2% | $3.92 | |
| high | 36% ±4% | $3.45 | |
| default | 31% ±1% | $2.82 | |
| Claude Sonnet 4.6 | high | 30% ±4% | $5.52 |
| high | 12% ±1% | $2.14 |
Leaderboard snapshot: September 22, 2026. These effort settings are not necessarily the models’ defaults.[40]
Model and subscription pricing
Usage estimates and multipliers are relative to each vendor's own plan, not a common token allowance. Real cost per completed task depends on model and harness token use.
Individual
$0
Claude Free doesn’t include Claude Code.
$8
No listed plan at this price
$20
Plus[1]
~15–160 GPT-6.1 Sol local messages / 5h
Estimate, not a fixed limit. Shared with ChatGPT Work; weekly limits may apply.
$100
$200
Pro $200[1]
Higher usage; no fixed multiplier published
No five-hour limit; weekly limits may apply.
$500
No listed plan at this price
Teams, per seat
$20Annual billing; $25 monthly
Business[1]
At least 2 users. Includes SAML SSO, MFA, larger cloud VMs, and no training on business data by default.
Team standard[20]
2 to 150 seats. Includes Claude Code, SSO, central billing, and no training on your content by default.
$100Annual billing; $125 monthly
No listed plan at this price
What each vendor publishes about usage
These numbers measure different things, one a message allowance and the other an observed spend, so don’t compare them directly. They’re starting points for planning usage, not a cost-per-task comparison.
15–160
local messages per 5 hours on Plus with GPT-6.1 Sol[1]
This is OpenAI’s estimate, and GPT-6 Luna gets 350–3,000. Cloud tasks can use more of the allowance than local messages, and weekly limits may also apply.
~$13
per developer per active day[22]
This is Anthropic’s average across enterprise deployments: $150–250 per developer per month, and under $30 per active day for 90% of users.
Prices are in US dollars per million tokens. API use is billed separately from subscriptions; token rates alone do not predict the cost of completing a task.
| Model | Context | Input → output | Long prompts | |||
|---|---|---|---|---|---|---|
| GPT-6 Luna[18] | $0.10 | $0.01 | $0.50 | 1.05M | Over 272K input: 2× input, 1.5× output | |
| Claude Haiku 4.5[34][35] | $1.00 | $0.10 | $5.00 | 200K | No surcharge | |
| Claude Sonnet 5.5[34][35] | $2.00 | $0.20 | $10.00 | 1M | Standard rate across the full window | |
| GPT-6.1 Sol[17][2] | $2.00 | $0.10 | $10.00 | 1.05M | Over 272K input: 2× input, 1.5× output | |
| Claude Opus 5.5[34][24][35] | $4.00 | $0.20 | $20.00 | 1M | Standard rate across the full window | |
| Claude Fable 5.1[34][24][35] | $10.00 | $0.25 | $50.00 | 1M | Standard rate across the full window | |
| GPT-6 Astra[19] | $10.00 | $1.00 | $50.00 | 1.05M | Over 272K input: 2× input, 1.5× output |
Lowest estimate: GPT-6 Luna at $4.30
- GPT-6 Luna$4.30
- Claude Haiku 4.5$43.00
- GPT-6.1 Sol$78.00
- Claude Sonnet 5.5$86.00
- Claude Opus 5.5$156.00
- Claude Fable 5.1$370.00
- GPT-6 Astra$430.00
This assumes no single request goes over 272K input tokens. OpenAI bills any request that does at 2× input and 1.5× output for that whole request. A monthly total can’t show whether that happened, so check your per-request prompt sizes. Anthropic’s 1M-context models don’t charge extra for long prompts.
These are list prices at standard speed. The estimate leaves out cache writes (OpenAI: 1.25× input; Anthropic: 1.25× for five-minute or 2× for one-hour caching), batch discounts, regional pricing, and tax. Agents can use different token counts for the same task.
Recommendations by task
Building web interfaces
Claude Code#1 on Arena WebDev
Opus 5.5 max scored 1,818 against GPT-6 Astra max at 1,789 in Arena’s September 30 human-preference ranking. This measures frontend model output, not Claude Code versus Codex as harnesses.[41]
Terminal-heavy engineering
CodexLower-cost general tasks
Codex$0.72 vs $5.98 per index task
At max effort, Artificial Analysis scores GPT-6.1 Sol 52 and Opus 5.5 58 on its ten-evaluation Intelligence Index. Average API cost per index task was $0.72 and $5.98 respectively. These are models in its test setup, not cost per PR.[42]
Features and integrations
Both agents edit files, run commands, and support MCP, skills, hooks, subagents, and plugins.
Openness and inference
Safety defaults
Hosted and automated work
CodexChatGPT plan required
Claude CodePro, Max, Team; self-hosting for teams
Plans and billing
How to compare both on your codebase
You can compare both on the same repository, with shared instructions and acceptance criteria. Here’s how we’d run a fair comparison on your own code.
- 01
Share one instruction file
Put your conventions in AGENTS.md. Codex reads it natively, and Claude Code reads it when there’s no CLAUDE.md, or alongside one if you turn that setting on.[11][32]
- 02
Pick real tasks
Choose three to five closed issues with known-good fixes, such as a bug, a small feature, and a refactor. Give both tools the same acceptance criteria.
- 03
Start from the same state
Run each tool from the same commit, with the same dependencies and a clean environment, so neither benefits from leftover state.
- 04
Compare the pull requests
Judge the diffs, the test results, and how long review took. Record what each run used from your plan, or what it cost in API tokens.
Common questions
Sources and methodology
We use vendor documentation and pricing, plus Terminal-Bench, DeepSWE, Arena, and Artificial Analysis results. We checked these sources on October 1, 2026. Prices and plans change often, so confirm with the vendor before you buy.
The recommendations are our interpretation of published measurements, not an overall score for either tool. Model-only tests don't establish which harness will perform better on your repository.
Replicas publishes this comparison and supports both harnesses. If you spot an error, email founders@replicas.dev and we'll fix it.
- 1Codex pricing and plan features · OpenAI
- 2Codex models · OpenAI
- 3Codex open-source components · OpenAI
- 4openai/codex repository and license · GitHub
- 5Codex agent approvals and security · OpenAI
- 6Codex Windows sandbox · OpenAI
- 7Codex cloud environments · OpenAI
- 8Codex scheduled tasks · OpenAI
- 9Codex code review in GitHub · OpenAI
- 10Use Codex in Linear · OpenAI
- 11Codex AGENTS.md instructions · OpenAI
- 12Codex configuration reference · OpenAI
- 13Codex Azure model provider configuration · OpenAI
- 14Codex with Amazon Bedrock · OpenAI
- 15Codex with Azure OpenAI in Microsoft Foundry · Microsoft
- 16OpenAI API pricing · OpenAI
- 17GPT-6.1 Sol model specifications · OpenAI
- 18GPT-6 Luna model specifications · OpenAI
- 19GPT-6 Astra model specifications · OpenAI
- 20Claude plans and pricing · Anthropic
- 21Claude Max plan tiers and pricing · Anthropic
- 22Manage Claude Code costs · Anthropic
- 23Claude Code platforms and integrations · Anthropic
- 24Claude Code model configuration · Anthropic
- 25anthropics/claude-code license · GitHub
- 26Claude Code sandboxing · Anthropic
- 27Claude Code permission modes · Anthropic
- 28Use Claude Code in the cloud · Anthropic
- 29Claude Code routines · Anthropic
- 30Claude Code Review · Anthropic
- 31Claude Code GitHub Actions · Anthropic
- 32Claude Code memory and CLAUDE.md · Anthropic
- 33Claude Code on third-party platforms · Anthropic
- 34Claude API pricing · Anthropic
- 35Claude models overview · Anthropic
- 36Terminal-Bench 4.0 leaderboard · Terminal-Bench
- 37Introducing Claude Opus 5.5 · Anthropic
- 38Introducing Claude Sonnet 5.5 · Anthropic
- 39Separating signal from noise in coding evaluations · OpenAI
- 40DeepSWE v1.1 leaderboard · Datacurve
- 41Code Arena WebDev leaderboard · Arena
- 42GPT-6.1 Sol versus Claude Opus 5.5 model comparison · Artificial Analysis