WritingmateWritingmate

AI model reference · 2026

GPT-5.6 Terra

OpenAI: GPT-5.6 Terra logoOpenAIModel details, evidence, and API facts

GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier.

Text generation
Image input
Reasoning
Provider endpoint accepts tool parameters
Context window
1,050,000 tokens
Maximum output
128,000 tokens
Base provider API input
$2.00 / 1M tokens
Starting provider API rate; context tiers apply
Published evidence
6 benchmarks, 0 Arena results

Overview

About GPT-5.6 Terra

A concise catalog overview. Technical limits and published evaluation evidence are listed separately below.

GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier.

text
multimodal
reasoning

Published evaluations

Benchmark results

Only results mapped to this exact model and backed by a named source appear here. Different evaluation protocols are not treated as interchangeable.

Source-backed benchmark results for GPT-5.6 Terra
BenchmarkScoreEvaluation detailsSource
AutomationBench-AAA 657-task benchmark of multi-step work across simulated SaaS applications in six business domains. The Artificial Analysis protocol reports a guardrail-aware score and is distinct from both the public Zapier split and the unrelated dynamic AutoBench framework.45.6%Guardrail-aware scoreVersion: AutomationBench-AA 2026 · Split: 657 tasks across Finance, HR, Marketing, Operations, Sales, and Support · Harness: Artificial Analysis independent implementation · Attempts: One evaluated run per task · Tools: REST APIs for simulated SaaS applications · Evaluator: Programmatic end-state grader with guardrail checks · Reasoning effort: maxMethodology Artificial AnalysisIndependently reproduced
DeepSWE 1.1A long-horizon software-engineering benchmark with 113 original tasks graded by hand-written tests.70.0%±3Pass@1Version: DeepSWE 1.1 · Split: 113 tasks across 91 repositories · Harness: mini-swe-agent · Reasoning effort: maxProtocol note: All models use the same harness. Cost, output-token, and step counts are retained per result as protocol context.; Confidence interval: ±3Methodology DataCurveBenchmark owner
Terminal-Bench 2.1Version 2.1 of the benchmark for completing realistic tasks in terminal environments. Harness, resource limits, and attempt count are part of the protocol.78.4%±1.3Mean task successVersion: Terminal-Bench 2.1, Harbor dataset revision 6 · Split: 89 revised tasks · Harness: Codex · Attempts: Five attempts per task · Reasoning effort: maxProtocol note: Every result retains its agent. Rows with different agents are not model-only comparisons.; Agent: Codex; Run date: 2026-07-11; Confidence interval: ±1.3Methodology Terminal-BenchBenchmark owner
GPQA DiamondThe highest-quality subset of Graduate-Level Google-Proof Q&A, designed to test expert-level scientific reasoning in biology, physics, and chemistry.92.5%AccuracyVersion: GPQA Diamond 198 · Split: 198 Diamond questions · Harness: Artificial Analysis independent evaluation · Tools: Not disclosed on the score page · Reasoning effort: maxMethodology Artificial AnalysisIndependently reproduced
Humanity's Last ExamA 2,500-question expert-level benchmark spanning dozens of academic fields. Tool-assisted and no-tools results are separate protocols and must not be merged.42.9%AccuracyVersion: May 2025 text-only revision · Split: 2,158 text-only questions from the 2,500-question May 2025 revision · Attempts: pass@1 · Tools: No browser or retrieval tools · Evaluator: LLM equality checker with numerical tolerance · Reasoning effort: maxMethodology Artificial AnalysisIndependently reproduced
LiveBenchA contamination-resistant benchmark refreshed on a fixed release cadence. Scores from different LiveBench releases must never be compared as the same protocol.77.9%Mean of category averagesVersion: LiveBench 2026-06-25 · Split: 2026-06-25 release, 23 tasks across seven categories · Reasoning effort: maxProtocol note: Overall is the mean of category averages. This protocol is not comparable with the 2026-01-08 release.Methodology LiveBenchBenchmark owner

Human preference

Arena results

Arena scores come from blind human preference votes. They are reported separately from task benchmarks and are not used as substitutes for missing benchmark results.

No exact Arena match is available for this model. Similar names and provider variants are not merged automatically.

Technical reference

Specifications and API facts

Technical limits and reference provider prices come from the current catalog unless a reviewed external source is listed. Missing values stay marked as unavailable.

Model
Provider
OpenAI
Release date
Not available in reviewed sources
Catalog added
Jul 9, 2026
Context window
1,050,000 tokens
Maximum output
128,000 tokens
Knowledge cutoff
2026-02-16
License
Not available in reviewed sources
Capabilities and API
Input types
File, Image, Text
Output types
Text
Reasoning
Supported
Provider endpoint tool parameters
Accepts tool parameters
Base provider API input
$2.00 / 1M tokens
Base provider cached input
$0.20 / 1M tokens
Base provider API output
$12 / 1M tokens
Provider tier at ≥ 272,000 prompt tokens
Input $4.00 / 1M tokensOutput $18 / 1M tokensCached input $0.40 / 1M tokens

Provider token rates are context-dependent. The base rates apply below the listed prompt thresholds; the matching tier applies at or above each threshold.

These are reference provider API rates, not Writingmate checkout charges. Access in Writingmate follows the allowances of your Writingmate plan.

Provenance

Sources and update status

Source links are attached to the facts and evaluations they support. Catalog-only values are not presented as independently verified claims.

Try GPT-5.6 Terra in Writingmate

Use this model, then compare its response with other AI models in the same workspace.