# Codex vs Claude Code: benchmarks, model pricing, and plans

Compare Codex and Claude Code benchmark results, model API rates including cached input, context windows, subscription prices and usage limits. Reviewed 2026-10-01. Published by Replicas, which runs both tools.

- Canonical: https://replicas.dev/resources/codex-vs-claude-code

## Quick comparison

Compare coding benchmark results, model API rates, context windows, and subscription costs for Codex and Claude Code.

- **Codex:** Lower measured API cost in both the model and harness tests shown here. [42](https://artificialanalysis.ai/models/comparisons/gpt-6-1-sol-vs-claude-opus-5-5)[36](https://www.tbench.ai/leaderboard)

- **Claude Code:** Higher model score; the tested harness pair tied on first tries at a higher cost. [42](https://artificialanalysis.ai/models/comparisons/gpt-6-1-sol-vs-claude-opus-5-5)[36](https://www.tbench.ai/leaderboard)

## Model benchmarks

Terminal-Bench 4.0 tests Codex + GPT-6 Astra and Claude Code + Fable 5.1 on the same 330 trials. Scores reflect both the model and harness, at five effort levels. [36](https://www.tbench.ai/leaderboard)

| Effort | Codex + GPT-6 Astra first try | Pass@5 | Run cost | Claude Code + Fable 5.1 first try | Pass@5 | Run cost |
| --- | --- | --- | --- | --- | --- | --- |
| low | 50.6% ±2.8 | 63.6% | $1,557 | 43.3% ±3.6 | 65.2% | $2,359 |
| medium | 54.2% ±2.7 | 66.7% | $1,915 | 53.9% ±3.4 | 71.2% | $2,833 |
| high | 57.9% ±3 | 71.2% | $2,269 | 54.5% ±3.4 | 75.8% | $3,985 |
| xhigh | 57.9% ±2.7 | 69.7% | $2,351 | 57.9% ±3.4 | 75.8% | $4,872 |
| max | 58.2% ±2.8 | 71.2% | $3,267 | 57.9% ±3.8 | 78.8% | $6,244 |

- At xhigh effort, both pairs solve 57.9% on the first attempt; their margins of error overlap. [36](https://www.tbench.ai/leaderboard)
- Codex + GPT-6 Astra costs $2,351 for the xhigh run; Claude Code + Fable 5.1 costs $4,872. [36](https://www.tbench.ai/leaderboard)
- With five attempts, Claude Code + Fable 5.1 solves 78.8% against 71.2% at max effort. [36](https://www.tbench.ai/leaderboard)

- Neither pair is the everyday configuration. Claude Code defaults to Opus 5.5, OpenAI recommends GPT-6.1 Sol for Codex, and neither has a leaderboard entry yet. [24](https://code.claude.com/docs/en/model-config)[2](https://learn.chatgpt.com/docs/models)[36](https://www.tbench.ai/leaderboard)
- Anthropic’s launch posts report higher Terminal-Bench 4.0 scores for Opus 5.5 (66.4% at xhigh effort) and Sonnet 5.5 (70.6%). Those are Anthropic’s own runs, not leaderboard entries. Anthropic lists no Terminal-Bench 4.0 score for GPT-6 Sol, and quotes GPT-6 Astra at 57.9% as reported by OpenAI. [37](https://www.anthropic.com/news/claude-opus-5-5)[38](https://www.anthropic.com/claude-sonnet-5-5)
- Terminal tasks aren’t your codebase. OpenAI itself warns against drawing frontier conclusions from SWE-bench Verified, so we don’t cite it. [39](https://openai.com/index/separating-signal-from-noise-coding-evaluations/)

### DeepSWE v1.1

DeepSWE tests 28 models on 113 tasks using one mini-swe-agent harness. Each point shows a model’s highest-scoring effort setting, including seven models unchecked by default on the source site—not Codex versus Claude Code. [40](https://deepswe.datacurve.ai/)

| Model | Effort | Solved | Avg cost / task |
| --- | --- | --- | --- |
| GPT-6 Astra | xhigh | 74% ±3% | $4.43 |
| Gemini 3.8 Flash | high | 74% ±1% | $2.36 |
| Claude Opus 5 | max | 74% ±4% | $11.84 |
| GPT-5.6 Sol | max | 73% ±3% | $6.46 |
| Claude Fable 5 | xhigh | 70% ±3% | $13.41 |
| GPT-5.6 Terra | max | 70% ±3% | $3.96 |
| GLM-5.3 | max | 69% ±3% | $3.99 |
| Kimi K3 | max | 69% ±5% | $4.65 |
| Grok 4.6 | medium | 67% ±2% | $3.45 |
| GPT-5.6 Luna | max | 67% ±4% | $0.61 |
| GPT-5.5 | xhigh | 67% ±6% | $7.23 |
| Gemini 3.7 Flash | medium | 65% ±3% | $2.03 |
| GLM-5.3 Flash | max | 63% ±4% | $0.24 |
| DeepSeek V4 Pro | max | 63% ±6% | $1.67 |
| Claude Opus 4.8 | max | 59% ±2% | $13.22 |
| Qwen 3.8 Max | xhigh | 57% ±3% | $3.73 |
| Muse Spark 1.2 | xhigh | 55% ±2% | $3.70 |
| Claude Sonnet 5 | max | 54% ±4% | $26.40 |
| Grok 4.5 | high | 54% ±2% | $2.42 |
| DeepSeek V4 Flash | max | 53% ±4% | $0.46 |
| Muse Spark 1.1 | xhigh | 53% ±3% | $2.36 |
| GPT-5.4 | xhigh | 52% ±2% | $5.65 |
| Gemini 3.6 Flash | high | 47% ±4% | $2.21 |
| GLM-5.2 | max | 44% ±2% | $3.92 |
| Gemini 3.5 Flash | high | 36% ±4% | $3.45 |
| Kimi K2.7 Code | default | 31% ±1% | $2.82 |
| Claude Sonnet 4.6 | high | 30% ±4% | $5.52 |
| Gemini 3.1 Pro Preview | high | 12% ±1% | $2.14 |

Leaderboard snapshot: September 22, 2026. These effort settings are not necessarily the models’ defaults. [40](https://deepswe.datacurve.ai/)

## API pricing (USD per million tokens)

Standard direct-API prices. Text and image input → text output. Cache writes, long prompts, provider discounts, and differences in tokens per task are not included.

| Model | Tool | Input | Cache read | Output | Context | Input → output | Long prompts |
| --- | --- | --- | --- | --- | --- | --- | --- |
| GPT-6 Luna [18](https://developers.openai.com/api/docs/models/gpt-6-luna) | Codex | $0.1 | $0.01 | $0.5 | 1.05M | Text + image → text | Over 272K input: 2× input, 1.5× output |
| Claude Haiku 4.5 [34](https://platform.claude.com/docs/en/about-claude/pricing)[35](https://platform.claude.com/docs/en/about-claude/models/overview) | Claude Code | $1 | $0.1 | $5 | 200K | Text + image → text | No surcharge |
| Claude Sonnet 5.5 [34](https://platform.claude.com/docs/en/about-claude/pricing)[35](https://platform.claude.com/docs/en/about-claude/models/overview) | Claude Code | $2 | $0.2 | $10 | 1M | Text + image → text | Standard rate across the full window |
| GPT-6.1 Sol [17](https://developers.openai.com/api/docs/models/gpt-6.1-sol)[2](https://learn.chatgpt.com/docs/models) | Codex | $2 | $0.1 | $10 | 1.05M | Text + image → text | Over 272K input: 2× input, 1.5× output |
| Claude Opus 5.5 [34](https://platform.claude.com/docs/en/about-claude/pricing)[24](https://code.claude.com/docs/en/model-config)[35](https://platform.claude.com/docs/en/about-claude/models/overview) | Claude Code | $4 | $0.2 | $20 | 1M | Text + image → text | Standard rate across the full window |
| Claude Fable 5.1 [34](https://platform.claude.com/docs/en/about-claude/pricing)[24](https://code.claude.com/docs/en/model-config)[35](https://platform.claude.com/docs/en/about-claude/models/overview) | Claude Code | $10 | $0.25 | $50 | 1M | Text + image → text | Standard rate across the full window |
| GPT-6 Astra [19](https://developers.openai.com/api/docs/models/gpt-6-astra) | Codex | $10 | $1 | $50 | 1.05M | Text + image → text | Over 272K input: 2× input, 1.5× output |

## Subscription pricing

Plan allowances are vendor-specific and shared with their chat products. A 5× Claude plan is 5× Claude Pro per session, not 5× Codex. API prices do not describe included subscription usage.

| Plan type | Monthly | Codex | Claude Code |
| --- | --- | --- | --- |
| Individual | $0 | Free [1](https://learn.chatgpt.com/docs/pricing): GPT-6 Luna at standard speed in the desktop app, subject to rollout. | Claude Free doesn’t include Claude Code. |
| Individual | $8 | Go [1](https://learn.chatgpt.com/docs/pricing): GPT-6 Luna at standard speed in the desktop app, subject to rollout. | No listed plan at this price |
| Individual | $20 | Plus [1](https://learn.chatgpt.com/docs/pricing): ~15–160 GPT-6.1 Sol local messages / 5h. Estimate, not a fixed limit. Shared with ChatGPT Work; weekly limits may apply. | Pro [20](https://claude.com/pricing)[28](https://code.claude.com/docs/en/claude-code-on-the-web): Standard Pro allowance. Shared with Claude chat; $17/mo billed annually. |
| Individual | $100 | Pro $100 [1](https://learn.chatgpt.com/docs/pricing): No five-hour limit. Weekly limits may apply; no fixed token allowance published. | Max 5x [21](https://support.claude.com/en/articles/11049741-what-is-the-max-plan): 5× Pro per-session allowance. Five-hour resets; weekly limit applies. |
| Individual | $200 | Pro $200 [1](https://learn.chatgpt.com/docs/pricing): Higher usage; no fixed multiplier published. No five-hour limit; weekly limits may apply. | Max 20x [21](https://support.claude.com/en/articles/11049741-what-is-the-max-plan): 20× Pro per-session allowance. Five-hour resets; weekly limit applies. |
| Individual | $500 | Pro $500 [1](https://learn.chatgpt.com/docs/pricing): No five-hour limit. Adds GPT-6 Astra Ultrafast; weekly limits may apply. | No listed plan at this price |
| Teams, per seat | $20 (Annual billing; $25 monthly) | Business [1](https://learn.chatgpt.com/docs/pricing): At least 2 users. Includes SAML SSO, MFA, larger cloud VMs, and no training on business data by default. | Team standard [20](https://claude.com/pricing): 2 to 150 seats. Includes Claude Code, SSO, central billing, and no training on your content by default. |
| Teams, per seat | $100 (Annual billing; $125 monthly) | No listed plan at this price | Team premium [20](https://claude.com/pricing): 5× standard seat usage. You can mix seat types. |
| Teams, per seat | Custom | Enterprise & Edu [1](https://learn.chatgpt.com/docs/pricing): Contact sales. Adds SCIM, EKM, RBAC, Compliance API audit logs, and retention and residency controls. | Enterprise [20](https://claude.com/pricing): $20 per seat billed annually, plus usage at API rates. Adds SCIM, audit logs, IP allowlisting, and a HIPAA-ready option. |

## What each vendor publishes about usage

- **Codex:** 15–160 local messages per 5 hours on Plus with GPT-6.1 Sol [1](https://learn.chatgpt.com/docs/pricing). This is OpenAI’s estimate, and GPT-6 Luna gets 350–3,000. Cloud tasks can use more of the allowance than local messages, and weekly limits may also apply.
- **Claude Code:** ~$13 per developer per active day [22](https://code.claude.com/docs/en/costs). This is Anthropic’s average across enterprise deployments: $150–250 per developer per month, and under $30 per active day for 90% of users.

## Recommendations by task

- **Building web interfaces — Claude Code:** #1 on Arena WebDev. Opus 5.5 max scored 1,818 against GPT-6 Astra max at 1,789 in Arena’s September 30 human-preference ranking. This measures frontend model output, not Claude Code versus Codex as harnesses. [41](https://arena.ai/leaderboard/code/webdev)
- **Terminal-heavy engineering — Codex:** 57.9% solved at less than half the run cost. Codex + GPT-6 Astra and Claude Code + Fable 5.1 both solved 57.9% on the first attempt at xhigh effort. Their full Terminal-Bench runs cost $2,351 and $4,872 respectively. Neither is the everyday model. [36](https://www.tbench.ai/leaderboard)[2](https://learn.chatgpt.com/docs/models)[24](https://code.claude.com/docs/en/model-config)
- **Lower-cost general tasks — Codex:** $0.72 vs $5.98 per index task. At max effort, Artificial Analysis scores GPT-6.1 Sol 52 and Opus 5.5 58 on its ten-evaluation Intelligence Index. Average API cost per index task was $0.72 and $5.98 respectively. These are models in its test setup, not cost per PR. [42](https://artificialanalysis.ai/models/comparisons/gpt-6-1-sol-vs-claude-opus-5-5)

## Features: Openness and inference

| Feature | Codex | Claude Code |
| --- | --- | --- |
| Open-source agent [3](https://learn.chatgpt.com/docs/open-source)[4](https://github.com/openai/codex)[25](https://github.com/anthropics/claude-code/blob/main/LICENSE.md) | The CLI, SDK, and app server are Apache-2.0. The IDE extension and Codex cloud are not open source. | The public repository hosts plugins and issue tracking. Its license is “All rights reserved” under Anthropic’s Commercial Terms. |
| Amazon Bedrock [14](https://learn.chatgpt.com/docs/amazon-bedrock)[33](https://code.claude.com/docs/en/third-party-integrations) | Codex documents a built-in Amazon Bedrock provider. | Claude Code documents Amazon Bedrock. |
| Azure OpenAI / Foundry [13](https://learn.chatgpt.com/docs/config-file/config-advanced)[15](https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/codex?view=foundry-classic)[33](https://code.claude.com/docs/en/third-party-integrations) | Codex can use Azure OpenAI in Microsoft Foundry with a configured model provider for local inference, not hosted Codex cloud tasks. | Claude Code documents Microsoft Foundry. |
| Google Cloud [12](https://learn.chatgpt.com/docs/config-file/config-reference)[33](https://code.claude.com/docs/en/third-party-integrations) | Codex documents custom model providers, but not a direct Google Cloud Agent Platform setup. | Claude Code documents Google Cloud Agent Platform. |
| Instruction files [11](https://learn.chatgpt.com/docs/agent-configuration/agents-md)[32](https://code.claude.com/docs/en/memory) | AGENTS.md, concatenated from global to project level, with a 32 KiB default cap. The docs don’t mention CLAUDE.md. | CLAUDE.md and .claude/rules. It reads AGENTS.md natively when no CLAUDE.md exists, and a setting makes it read both. |

## Features: Safety defaults

| Feature | Codex | Claude Code |
| --- | --- | --- |
| Local sandbox [5](https://learn.chatgpt.com/docs/agent-approvals-security)[26](https://code.claude.com/docs/en/sandboxing) | On by default and enforced by the OS. Writes are limited to the workspace, the network is off, and .git is read-only. | Turned on with `/sandbox` or `sandbox.enabled`, or enforced through managed settings. If the sandbox can’t start, commands run unsandboxed unless you set it to fail instead. |
| Default approvals [5](https://learn.chatgpt.com/docs/agent-approvals-security)[27](https://code.claude.com/docs/en/permission-modes) | The Auto preset works freely inside the workspace and asks before using the network or editing outside it. | Auto mode, the default since v2.1.283, has a classifier model review each action. Deny rules apply in every mode. |
| Windows sandbox [6](https://learn.chatgpt.com/docs/windows/windows-sandbox)[26](https://code.claude.com/docs/en/sandboxing) | Has a native Windows sandbox, with elevated and unelevated modes. | Not supported on native Windows. Run Claude Code in WSL2 to use the sandbox. |

## Features: Hosted and automated work

| Feature | Codex | Claude Code |
| --- | --- | --- |
| Cloud tasks [7](https://learn.chatgpt.com/docs/cloud)[1](https://learn.chatgpt.com/docs/pricing)[28](https://code.claude.com/docs/en/claude-code-on-the-web)[23](https://code.claude.com/docs/en/platforms) | Each task gets its own cloud environment and keeps running while your computer sleeps. Needs a ChatGPT plan, since API-key sign-in has no cloud features. | Runs in Anthropic-managed VMs on Pro, Max, and Team. Team and Enterprise can use self-hosted environments. |
| Scheduled and triggered runs [8](https://learn.chatgpt.com/docs/automations)[29](https://code.claude.com/docs/en/routines) | Scheduled tasks run locally in the desktop app, which needs your computer on, or in the cloud from the web. Supported app events can trigger runs on eligible plans. | Routines, a research preview, run in the cloud on a schedule, an API call, or a GitHub event. Local desktop tasks and `/loop` are separate features. |
| Managed PR review [9](https://learn.chatgpt.com/docs/third-party/github)[1](https://learn.chatgpt.com/docs/pricing)[30](https://code.claude.com/docs/en/code-review)[31](https://code.claude.com/docs/en/github-actions) | Mention `@codex review` or turn on automatic reviews in GitHub. Included from Plus. | Code Review is a research preview on Team and Enterprise that averages $15–25 per review, billed to usage credits. GitHub Actions works on any plan. |

## Features: Plans and billing

| Feature | Codex | Claude Code |
| --- | --- | --- |
| Free access [1](https://learn.chatgpt.com/docs/pricing)[20](https://claude.com/pricing) | The Free and $8 Go plans include GPT-6 Luna in the desktop app, subject to rollout. | Not on the Free plan. Access starts with Pro at $20/mo. |
| Enterprise usage [1](https://learn.chatgpt.com/docs/pricing)[20](https://claude.com/pricing) | With flexible pricing, usage scales with credits. Without it, each seat gets Plus-level limits. | All usage is billed at API rates on top of the $20 seat fee. Admins set spend limits. |

## How to compare both on your codebase

You can compare both on the same repository, with shared instructions and acceptance criteria. Here’s how we’d run a fair comparison on your own code.

- **Share one instruction file:** Put your conventions in AGENTS.md. Codex reads it natively, and Claude Code reads it when there’s no CLAUDE.md, or alongside one if you turn that setting on. [11](https://learn.chatgpt.com/docs/agent-configuration/agents-md)[32](https://code.claude.com/docs/en/memory)
- **Pick real tasks:** Choose three to five closed issues with known-good fixes, such as a bug, a small feature, and a refactor. Give both tools the same acceptance criteria.
- **Start from the same state:** Run each tool from the same commit, with the same dependencies and a clean environment, so neither benefits from leftover state.
- **Compare the pull requests:** Judge the diffs, the test results, and how long review took. Record what each run used from your plan, or what it cost in API tokens.

Disclosure: Replicas publishes this comparison and supports both harnesses. It runs Codex and Claude Code in isolated cloud workspaces, so you can send the same task to each and review both pull requests side by side, signed in with your existing ChatGPT or Claude account or API keys. Try both on Replicas: https://app.replicas.dev/auth?mode=signup

## FAQ

### Is Codex or Claude Code better?
Neither has a demonstrated overall edge. On the public Terminal-Bench 4.0 leaderboard, the strongest Codex and Claude Code configurations both solve about 58% of tasks, with overlapping margins of error and different costs. Choose based on the plan you already pay for, the safety defaults you need, and where you want inference to run.

### Is Codex cheaper than Claude Code?
Full access starts at $20/mo for both, but Codex also has limited access on its Free and $8 Go plans. On the API, GPT-6.1 Sol and Claude Sonnet 5.5 both cost $2 per million input tokens and $10 per million output tokens. Claude Code’s default model, Opus 5.5, costs $4 and $20. What you actually pay depends on how many tokens each agent uses on your work.

### Can I use Codex or Claude Code with an API key instead of a subscription?
Yes. Both bill per token at API rates when you sign in with a key. With Codex, an API key covers the CLI, SDK, and IDE extension, but not cloud features like GitHub code review or Slack.

### Is Codex open source? Is Claude Code?
The Codex CLI, SDK, and app server are open source under Apache-2.0, but the IDE extension and Codex cloud are not. Claude Code isn’t open source: its public repository hosts plugins and issue tracking under an “All rights reserved” license.

### Which models do Codex and Claude Code use?
OpenAI recommends GPT-6.1 Sol in Codex for complex work, GPT-6 Luna for focused tasks, and GPT-6 Astra for the hardest problems, depending on your plan. Claude Code defaults to Opus 5.5 on Anthropic plans and the API, and you can switch to Sonnet 5.5, Haiku 4.5, or Fable 5.1.

### Can Claude Code read AGENTS.md?
Yes. Claude Code reads AGENTS.md natively when a repository has no CLAUDE.md, and a setting makes it read both. Codex’s documentation doesn’t mention reading CLAUDE.md.

### Do both run in the cloud?
Yes. Codex cloud tasks need a ChatGPT plan. Claude Code cloud sessions run in Anthropic-managed VMs on Pro, Max, Team, and eligible Enterprise seats. Both also support scheduled work, and Claude Code’s cloud routines are still a research preview.

## Sources

1. [Codex pricing and plan features](https://learn.chatgpt.com/docs/pricing) (OpenAI)
2. [Codex models](https://learn.chatgpt.com/docs/models) (OpenAI)
3. [Codex open-source components](https://learn.chatgpt.com/docs/open-source) (OpenAI)
4. [openai/codex repository and license](https://github.com/openai/codex) (GitHub)
5. [Codex agent approvals and security](https://learn.chatgpt.com/docs/agent-approvals-security) (OpenAI)
6. [Codex Windows sandbox](https://learn.chatgpt.com/docs/windows/windows-sandbox) (OpenAI)
7. [Codex cloud environments](https://learn.chatgpt.com/docs/cloud) (OpenAI)
8. [Codex scheduled tasks](https://learn.chatgpt.com/docs/automations) (OpenAI)
9. [Codex code review in GitHub](https://learn.chatgpt.com/docs/third-party/github) (OpenAI)
10. [Use Codex in Linear](https://learn.chatgpt.com/docs/third-party/linear) (OpenAI)
11. [Codex AGENTS.md instructions](https://learn.chatgpt.com/docs/agent-configuration/agents-md) (OpenAI)
12. [Codex configuration reference](https://learn.chatgpt.com/docs/config-file/config-reference) (OpenAI)
13. [Codex Azure model provider configuration](https://learn.chatgpt.com/docs/config-file/config-advanced) (OpenAI)
14. [Codex with Amazon Bedrock](https://learn.chatgpt.com/docs/amazon-bedrock) (OpenAI)
15. [Codex with Azure OpenAI in Microsoft Foundry](https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/codex?view=foundry-classic) (Microsoft)
16. [OpenAI API pricing](https://developers.openai.com/api/docs/pricing) (OpenAI)
17. [GPT-6.1 Sol model specifications](https://developers.openai.com/api/docs/models/gpt-6.1-sol) (OpenAI)
18. [GPT-6 Luna model specifications](https://developers.openai.com/api/docs/models/gpt-6-luna) (OpenAI)
19. [GPT-6 Astra model specifications](https://developers.openai.com/api/docs/models/gpt-6-astra) (OpenAI)
20. [Claude plans and pricing](https://claude.com/pricing) (Anthropic)
21. [Claude Max plan tiers and pricing](https://support.claude.com/en/articles/11049741-what-is-the-max-plan) (Anthropic)
22. [Manage Claude Code costs](https://code.claude.com/docs/en/costs) (Anthropic)
23. [Claude Code platforms and integrations](https://code.claude.com/docs/en/platforms) (Anthropic)
24. [Claude Code model configuration](https://code.claude.com/docs/en/model-config) (Anthropic)
25. [anthropics/claude-code license](https://github.com/anthropics/claude-code/blob/main/LICENSE.md) (GitHub)
26. [Claude Code sandboxing](https://code.claude.com/docs/en/sandboxing) (Anthropic)
27. [Claude Code permission modes](https://code.claude.com/docs/en/permission-modes) (Anthropic)
28. [Use Claude Code in the cloud](https://code.claude.com/docs/en/claude-code-on-the-web) (Anthropic)
29. [Claude Code routines](https://code.claude.com/docs/en/routines) (Anthropic)
30. [Claude Code Review](https://code.claude.com/docs/en/code-review) (Anthropic)
31. [Claude Code GitHub Actions](https://code.claude.com/docs/en/github-actions) (Anthropic)
32. [Claude Code memory and CLAUDE.md](https://code.claude.com/docs/en/memory) (Anthropic)
33. [Claude Code on third-party platforms](https://code.claude.com/docs/en/third-party-integrations) (Anthropic)
34. [Claude API pricing](https://platform.claude.com/docs/en/about-claude/pricing) (Anthropic)
35. [Claude models overview](https://platform.claude.com/docs/en/about-claude/models/overview) (Anthropic)
36. [Terminal-Bench 4.0 leaderboard](https://www.tbench.ai/leaderboard) (Terminal-Bench)
37. [Introducing Claude Opus 5.5](https://www.anthropic.com/news/claude-opus-5-5) (Anthropic)
38. [Introducing Claude Sonnet 5.5](https://www.anthropic.com/claude-sonnet-5-5) (Anthropic)
39. [Separating signal from noise in coding evaluations](https://openai.com/index/separating-signal-from-noise-coding-evaluations/) (OpenAI)
40. [DeepSWE v1.1 leaderboard](https://deepswe.datacurve.ai/) (Datacurve)
41. [Code Arena WebDev leaderboard](https://arena.ai/leaderboard/code/webdev) (Arena)
42. [GPT-6.1 Sol versus Claude Opus 5.5 model comparison](https://artificialanalysis.ai/models/comparisons/gpt-6-1-sol-vs-claude-opus-5-5) (Artificial Analysis)

## Related docs


