AI model reference · 2026
Kimi K3
Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI.
- Context window
- 1,048,576 tokens
- Maximum output
- Not available
- Base provider API input
- $3.00 / 1M tokens
- Base provider API rate
- Published evidence
- 5 benchmarks, 0 Arena results
Overview
About Kimi K3
A concise catalog overview. Technical limits and published evaluation evidence are listed separately below.
Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI.
Published evaluations
Benchmark results
Only results mapped to this exact model and backed by a named source appear here. Different evaluation protocols are not treated as interchangeable.
| Benchmark | Score | Evaluation details | Source |
|---|---|---|---|
| AutomationBench-AAA 657-task benchmark of multi-step work across simulated SaaS applications in six business domains. The Artificial Analysis protocol reports a guardrail-aware score and is distinct from both the public Zapier split and the unrelated dynamic AutoBench framework. | 52.7% | Guardrail-aware scoreVersion: AutomationBench-AA 2026 · Split: 657 tasks across Finance, HR, Marketing, Operations, Sales, and Support · Harness: Artificial Analysis independent implementation · Attempts: One evaluated run per task · Tools: REST APIs for simulated SaaS applications · Evaluator: Programmatic end-state grader with guardrail checks · Reasoning effort: maxMethodology | Artificial AnalysisIndependently reproduced |
| DeepSWE 1.1A long-horizon software-engineering benchmark with 113 original tasks graded by hand-written tests. | 69.0%±5 | Pass@1Version: DeepSWE 1.1 · Split: 113 tasks across 91 repositories · Harness: mini-swe-agent · Reasoning effort: maxProtocol note: All models use the same harness. Cost, output-token, and step counts are retained per result as protocol context.; Confidence interval: ±5Methodology | DataCurveBenchmark owner |
| GPQA DiamondThe highest-quality subset of Graduate-Level Google-Proof Q&A, designed to test expert-level scientific reasoning in biology, physics, and chemistry. | 93.5% | AccuracyVersion: GPQA Diamond 198 · Split: 198 Diamond questions · Harness: Artificial Analysis independent evaluation · Tools: Not disclosed on the score page · Reasoning effort: maxMethodology | Artificial AnalysisIndependently reproduced |
| Humanity's Last ExamA 2,500-question expert-level benchmark spanning dozens of academic fields. Tool-assisted and no-tools results are separate protocols and must not be merged. | 46.9% | AccuracyVersion: May 2025 text-only revision · Split: 2,158 text-only questions from the 2,500-question May 2025 revision · Attempts: pass@1 · Tools: No browser or retrieval tools · Evaluator: LLM equality checker with numerical tolerance · Reasoning effort: maxMethodology | Artificial AnalysisIndependently reproduced |
| LiveBenchA contamination-resistant benchmark refreshed on a fixed release cadence. Scores from different LiveBench releases must never be compared as the same protocol. | 79.2% | Mean of category averagesVersion: LiveBench 2026-06-25 · Split: 2026-06-25 release, 23 tasks across seven categoriesProtocol note: Overall is the mean of category averages. This protocol is not comparable with the 2026-01-08 release.Methodology | LiveBenchBenchmark owner |
Human preference
Arena results
Arena scores come from blind human preference votes. They are reported separately from task benchmarks and are not used as substitutes for missing benchmark results.
No exact Arena match is available for this model. Similar names and provider variants are not merged automatically.
Technical reference
Specifications and API facts
Technical limits and reference provider prices come from the current catalog unless a reviewed external source is listed. Missing values stay marked as unavailable.
- Provider
- Moonshot AI
- Release date
- Not available in reviewed sources
- Catalog added
- Jul 16, 2026
- Context window
- 1,048,576 tokens
- Maximum output
- Not available
- Knowledge cutoff
- Not available in reviewed sources
- License
- Not available in reviewed sources
- Input types
- Text, Image, Video
- Output types
- Text
- Reasoning
- Supported
- Provider endpoint tool parameters
- Accepts tool parameters
- Base provider API input
- $3.00 / 1M tokens
- Base provider cached input
- $0.30 / 1M tokens
- Base provider API output
- $15 / 1M tokens
These are reference provider API rates, not Writingmate checkout charges. Access in Writingmate follows the allowances of your Writingmate plan.
Provenance
Sources and update status
Source links are attached to the facts and evaluations they support. Catalog-only values are not presented as independently verified claims.
- AutomationBench-AA: Agentic SaaS Workflow BenchmarkArtificial Analysis · Retrieved Aug 9, 2026
- GPQA Diamond Benchmark LeaderboardArtificial Analysis · Retrieved Aug 9, 2026
- Humanity's Last Exam Benchmark LeaderboardArtificial Analysis · Retrieved Aug 9, 2026
- DeepSWE 1.1 LeaderboardDataCurve · Retrieved Aug 9, 2026
- LiveBench 2026-06-25 LeaderboardLiveBench · Retrieved Aug 9, 2026
- OpenRouter model catalogOpenRouter
Evidence last updated Aug 9, 2026.
Catalog record updated Jul 16, 2026.
Catalog-added dates describe when a model entered the catalog, not necessarily its public release date.
Try Kimi K3 in Writingmate
Use this model, then compare its response with other AI models in the same workspace.