AI model reference · 2026
Claude Sonnet 4
Claude Sonnet 4 significantly enhances the capabilities of its predecessor, Sonnet 3.7, excelling in both coding and reasoning tasks with improved precision and controllability.
- Context window
- 1,000,000 tokens
- Maximum output
- 64,000 tokens
- Base provider API input
- $3.00 / 1M tokens
- Starting provider API rate; context tiers apply
- Published evidence
- 2 benchmarks, 0 Arena results
Overview
About Claude Sonnet 4
A concise catalog overview. Technical limits and published evaluation evidence are listed separately below.
Claude Sonnet 4 significantly enhances the capabilities of its predecessor, Sonnet 3.7, excelling in both coding and reasoning tasks with improved precision and controllability.
Published evaluations
Benchmark results
Only results mapped to this exact model and backed by a named source appear here. Different evaluation protocols are not treated as interchangeable.
| Benchmark | Score | Evaluation details | Source |
|---|---|---|---|
| OSWorld-VerifiedA verified computer-use benchmark in which multimodal agents operate desktop applications and are graded from the resulting environment state. | 41.4% | Mean task rewardVersion: OSWorld-Verified v1 · Split: Up to 361 tasks, excluding eight Google Drive tasks where required · Harness: Official OSWorld unified environment · Attempts: One run · Tools: GUI interactionProtocol note: Only 100-step general-model rows are included in this normalized snapshot. Individual evaluated-task counts can differ when a task could not run.; Run date: 2025-07-27; Evaluated tasks: 361Methodology | OSWorldBenchmark owner |
| SWE-Bench VerifiedA human-validated subset of real GitHub issues used to measure whether a coding agent can produce repository patches that resolve the associated tests. | 64.9% | ResolvedVersion: SWE-Bench Verified 500 · Split: Verified, 500 tasks · Harness: mini-SWE-agent / version 1.0.0 · Attempts: One attemptProtocol note: Harness version: 1.0.0; Run date: 2025-07-26Methodology | SWE-benchBenchmark owner |
Human preference
Arena results
Arena scores come from blind human preference votes. They are reported separately from task benchmarks and are not used as substitutes for missing benchmark results.
No exact Arena match is available for this model. Similar names and provider variants are not merged automatically.
Technical reference
Specifications and API facts
Technical limits and reference provider prices come from the current catalog unless a reviewed external source is listed. Missing values stay marked as unavailable.
- Provider
- Anthropic
- Release date
- Not available in reviewed sources
- Catalog added
- May 22, 2025
- Context window
- 1,000,000 tokens
- Maximum output
- 64,000 tokens
- Knowledge cutoff
- 2025-01-31
- License
- Not available in reviewed sources
- Input types
- Image, Text, File
- Output types
- Text
- Reasoning
- Supported
- Provider endpoint tool parameters
- Accepts tool parameters
- Base provider API input
- $3.00 / 1M tokens
- Base provider cached input
- $0.30 / 1M tokens
- Base provider API output
- $15 / 1M tokens
- Provider tier at ≥ 200,000 prompt tokens
- Input $6.00 / 1M tokensOutput $22.5 / 1M tokensCached input $0.60 / 1M tokens
Provider token rates are context-dependent. The base rates apply below the listed prompt thresholds; the matching tier applies at or above each threshold.
These are reference provider API rates, not Writingmate checkout charges. Access in Writingmate follows the allowances of your Writingmate plan.
Provenance
Sources and update status
Source links are attached to the facts and evaluations they support. Catalog-only values are not presented as independently verified claims.
- OSWorld-Verified ResultsOSWorld · Retrieved Aug 9, 2026
- SWE-bench Verified Bash Only LeaderboardSWE-bench · Retrieved Aug 9, 2026
- OpenRouter model catalogOpenRouter
Evidence last updated Aug 9, 2026.
Catalog record updated May 22, 2025.
Catalog-added dates describe when a model entered the catalog, not necessarily its public release date.
Try Claude Sonnet 4 in Writingmate
Use this model, then compare its response with other AI models in the same workspace.