AI model reference · 2026
GPT-5.3 Codex
GPT-5.3-Codex is OpenAI’s most advanced agentic coding model, combining the frontier software engineering performance of GPT-5.2-Codex with the broader reasoning and professional knowledge capabilities of GPT-5.2.
- Context window
- 400,000 tokens
- Maximum output
- 128,000 tokens
- Base provider API input
- $1.75 / 1M tokens
- Base provider API rate
- Published evidence
- 8 benchmarks, 0 Arena results
Overview
About GPT-5.3 Codex
A concise catalog overview. Technical limits and published evaluation evidence are listed separately below.
GPT-5.3-Codex is OpenAI’s most advanced agentic coding model, combining the frontier software engineering performance of GPT-5.2-Codex with the broader reasoning and professional knowledge capabilities of GPT-5.2.
Published evaluations
Benchmark results
Only results mapped to this exact model and backed by a named source appear here. Different evaluation protocols are not treated as interchangeable.
| Benchmark | Score | Evaluation details | Source |
|---|---|---|---|
| BrowseCompA benchmark of difficult, verifiable information-seeking questions designed to measure an agent's ability to locate hard-to-find facts through web browsing. | 77.3% | AccuracyVersion: BrowseComp · Tools: Web browsing · Reasoning effort: xhighMethodology | OpenAIProvider reported |
| OSWorld-VerifiedA verified computer-use benchmark in which multimodal agents operate desktop applications and are graded from the resulting environment state. | 74.0% | Mean task rewardVersion: OSWorld-Verified · Reasoning effort: xhighProtocol note: Updated result using the API parameter that preserves original image resolution; supersedes OpenAI's 64.7 launch resultMethodology | OpenAIProvider reported |
| SWE-Bench ProA contamination-resistant software-engineering benchmark with long-horizon tasks across multiple programming languages. Public and private splits are distinct protocols. | 56.8% | ResolvedVersion: SWE-Bench Pro · Split: 731 public tasks · Reasoning effort: xhighMethodology | OpenAIProvider reported |
| SWE-Lancer (IC-Diamond subset)The individual-contributor Diamond subset of SWE-Lancer, which evaluates economically valuable real-world software-engineering tasks. | 81.4% | Tasks completedSplit: IC Diamond · Reasoning effort: xhighMethodology | OpenAIProvider reported |
| Terminal-Bench 2.0Version 2.0 of the benchmark for completing realistic tasks in terminal environments. Results must not be merged with Terminal-Bench 2.1. | 77.3% | Pass rateVersion: 2.0 · Reasoning effort: xhighMethodology | OpenAIProvider reported |
| Terminal-Bench 2.1Version 2.1 of the benchmark for completing realistic tasks in terminal environments. Harness, resource limits, and attempt count are part of the protocol. | 79.1% | Mean task successVersion: 2.1 · Harness: Codex CLIProtocol note: Verified owner-leaderboard run; not protocol-compatible with Meta's provider-reported bash-tool-only harnessMethodology | Terminal-BenchBenchmark owner |
| GPQA DiamondThe highest-quality subset of Graduate-Level Google-Proof Q&A, designed to test expert-level scientific reasoning in biology, physics, and chemistry. | 92.6% | AccuracyVersion: GPQA Diamond · Reasoning effort: xhighMethodology | OpenAIProvider reported |
| Cybersecurity CTFsCapture-the-flag challenges used to evaluate an agent's practical cybersecurity task performance. | 77.6% | Challenges completedReasoning effort: xhigh | OpenAIProvider reported |
Human preference
Arena results
Arena scores come from blind human preference votes. They are reported separately from task benchmarks and are not used as substitutes for missing benchmark results.
Published Arena rows for webdev are conflicting or ambiguous; no score was selected.
Technical reference
Specifications and API facts
Technical limits and reference provider prices come from the current catalog unless a reviewed external source is listed. Missing values stay marked as unavailable.
- Provider
- OpenAI
- Release date
- Feb 5, 2026
- Catalog added
- Feb 24, 2026
- Context window
- 400,000 tokens
- Maximum output
- 128,000 tokens
- Knowledge cutoff
- Not publicly disclosed
- License
- Proprietary
- Input types
- Text, Image, File
- Output types
- Text
- Reasoning
- Supported
- Provider endpoint tool parameters
- Accepts tool parameters
- Base provider API input
- $1.75 / 1M tokens
- Base provider cached input
- $0.175 / 1M tokens
- Base provider API output
- $14 / 1M tokens
These are reference provider API rates, not Writingmate checkout charges. Access in Writingmate follows the allowances of your Writingmate plan.
Provenance
Sources and update status
Source links are attached to the facts and evaluations they support. Catalog-only values are not presented as independently verified claims.
- Arena AI LeaderboardArena AI · Retrieved Aug 25, 2026
- Introducing GPT-5.3-CodexOpenAI · Retrieved Aug 9, 2026
- Introducing GPT-5.4OpenAI · Retrieved Aug 9, 2026
- Terminal-Bench 2.1 Verified LeaderboardTerminal-Bench · Retrieved Aug 9, 2026
- OpenRouter model catalogOpenRouter
Evidence last updated Aug 25, 2026.
Catalog record updated Feb 24, 2026.
Catalog-added dates describe when a model entered the catalog, not necessarily its public release date.
Try GPT-5.3 Codex in Writingmate
Use this model, then compare its response with other AI models in the same workspace.