Kimi K2.6 API pricing and reported benchmark results

Moonshot AI lists Kimi K2.6 at $0.95/1M input and $4/1M output with a 262,144-token context window. Its model card reports 80.2% on SWE-bench Verified and 90.5% on GPQA Diamond. Those are provider-reported benchmark results, not benchr measurements or a controlled cross-provider ranking.

By benchr Editorial Team · · Context, architecture, pricing, and benchmark provenance checked against Moonshot sources · View changelog

Input / 1MMoonshot AI
Output / 1MMoonshot AI
SWE-benchMoonshot-reported
GPQA DiamondMoonshot-reported

Pricing breakdown

kimi-k2-6 — official Moonshot AI pricing
TierRate / 1M tokens
Standard input$0.95
Standard output$4.00
Published context window262,144 tokens
benchr planning allowance200,000 tokens

How to read the two benchmark figures

SWE-bench Verified evaluates repository-level software repair, while GPQA Diamond contains difficult graduate-level science questions. Moonshot reports 80.2% and 90.5% respectively in the Kimi K2.6 model card. Different providers may use different prompts, tool setups, sampling budgets, or reporting conventions, so nearby headline scores are not automatically comparable. Use the figures to select evaluation candidates, then test accepted patches and domain answers under the same protocol.

MoE architecture implications

Moonshot documents a Mixture-of-Experts architecture with 1T total parameters and 32B active per token. That describes how capacity is routed; it does not by itself explain the listed API price or prove a quality, consistency, or efficiency advantage over dense models. Measure task acceptance, latency, memory use, and failure behavior on the deployment path you will use.

GPQA Diamond 90.5%: what it covers

GPQA Diamond tests difficult questions in biology, chemistry, and physics. Moonshot's reported 90.5% is one useful signal, but it does not translate directly into factual accuracy, hallucination rate, or safety in clinical, chemical, or other high-stakes work. Validate citations, uncertainty, refusal behavior, and expert-review requirements on your domain before relying on generated conclusions.

Cost scenarios

At 10M input + 3M output per month, Kimi K2.6's listed token rates total $9.50 + $12 = $21.50/month. At 50M input + 15M output, they total $107.50. These calculations exclude cache behavior, retries, hosting overhead, tool charges, and differences in accepted-task rates; compare total workload cost after a matched evaluation.

Use-case fit

Test for: Coding and science workloads where an open-weight option, the listed rates, or a 262,144-token context window are relevant shortlist criteria.

Test alternatives when: You need independent benchmark evidence, vision input, a provider-published maximum output, or controls and support not documented for your deployment path.

Decision checklist

Run a domain-specific evaluation with matched prompts, tools, sampling budgets, and review criteria. Score accepted patches or expert-reviewed answers rather than assuming provider headline benchmarks are directly comparable.

Check output volume and cache-hit rate. The listed $4/1M output rate can dominate total cost in verbose workflows; calculate cost per accepted result and include self-hosting operations if you evaluate the open weights.

Frequently asked

What do Kimi K2.6's benchmark figures establish?

Moonshot's model card reports 80.2% on SWE-bench Verified and 90.5% on GPQA Diamond. They are provider-reported results, not benchr measurements or a controlled cross-provider ranking. Compare candidates under matched prompts, tools, budgets, and acceptance criteria.

What is the MoE architecture in Kimi K2.6?

Moonshot documents a Mixture-of-Experts model with 1T total parameters and 32B active per token. That describes routing and active capacity; it does not by itself prove lower task cost, better quality, or greater edge-case variance than a dense model.

Why do this page's 262,144-token and 200,000-token context figures differ?

Moonshot publishes 262,144 tokens as the model context window. benchr uses 200,000 only as a conservative planning allowance, not as the provider's hard maximum. Validate usable length, quality, memory, and latency with representative long requests.

Changelog

  • — Corrected the published context from 200,000 to 262,144 tokens, retained 200,000 only as a labelled planning allowance, corrected the API ID, and labelled benchmark figures as Moonshot-reported.
  • — Expanded with dual-benchmark analysis, MoE architecture context, GPQA implications, and cost scenarios.
  • — Published. Pricing verified at platform.moonshot.cn.

Sources