Pricing breakdown
| Tier | Rate / 1M tokens |
|---|---|
| Standard input | $0.95 |
| Standard output | $4.00 |
| Published context window | 262,144 tokens |
| benchr planning allowance | 200,000 tokens |
How to read the two benchmark figures
SWE-bench Verified evaluates repository-level software repair, while GPQA Diamond contains difficult graduate-level science questions. Moonshot reports 80.2% and 90.5% respectively in the Kimi K2.6 model card. Different providers may use different prompts, tool setups, sampling budgets, or reporting conventions, so nearby headline scores are not automatically comparable. Use the figures to select evaluation candidates, then test accepted patches and domain answers under the same protocol.
MoE architecture implications
Moonshot documents a Mixture-of-Experts architecture with 1T total parameters and 32B active per token. That describes how capacity is routed; it does not by itself explain the listed API price or prove a quality, consistency, or efficiency advantage over dense models. Measure task acceptance, latency, memory use, and failure behavior on the deployment path you will use.
GPQA Diamond 90.5%: what it covers
GPQA Diamond tests difficult questions in biology, chemistry, and physics. Moonshot's reported 90.5% is one useful signal, but it does not translate directly into factual accuracy, hallucination rate, or safety in clinical, chemical, or other high-stakes work. Validate citations, uncertainty, refusal behavior, and expert-review requirements on your domain before relying on generated conclusions.
Cost scenarios
At 10M input + 3M output per month, Kimi K2.6's listed token rates total $9.50 + $12 = $21.50/month. At 50M input + 15M output, they total $107.50. These calculations exclude cache behavior, retries, hosting overhead, tool charges, and differences in accepted-task rates; compare total workload cost after a matched evaluation.
Use-case fit
Test for: Coding and science workloads where an open-weight option, the listed rates, or a 262,144-token context window are relevant shortlist criteria.
Test alternatives when: You need independent benchmark evidence, vision input, a provider-published maximum output, or controls and support not documented for your deployment path.
Decision checklist
Run a domain-specific evaluation with matched prompts, tools, sampling budgets, and review criteria. Score accepted patches or expert-reviewed answers rather than assuming provider headline benchmarks are directly comparable.
Check output volume and cache-hit rate. The listed $4/1M output rate can dominate total cost in verbose workflows; calculate cost per accepted result and include self-hosting operations if you evaluate the open weights.