By benchr Editorial Team · Published September 8, 2026Provider-published facts rechecked against the official sources on September 8, 2026

Kimi K2.7-Code: fewer thinking tokens, and a flat 2x for speed

Two decisions sit in this release, and neither of them is a leaderboard position.

A coding agent's reasoning trace shortens while a parallel track runs at double the rate.
Benchr model field plate Kimi K2.7-Code 1T / 32B · modified-MIT
Model researchThe visual for Kimi K2.7-Code: fewer thinking tokens, and a flat 2x for speed pairs evidence layers and comparison routes.
Input / 1M$0.95$0.19 on a cache hit
Output / 1M$4.004.2x the input rate
Context256K262,144 tokens
Weightsmodified-MIT1T total · 32B active

benchr rechecked Moonshot's pricing page and the official model card on September 8, 2026. Two things in this release are worth a decision, and neither of them is a leaderboard position.

The thinking-token claim is the cost claim

Moonshot says K2.7-Code reduces thinking-token usage by approximately 30% against K2.6 while improving on long-horizon coding work. Thinking tokens bill as output, and output here is $4.00 against $0.95 input — a ratio of 4.2 to 1.

So a 30% reduction in thinking tokens is not a footnote about efficiency. On a workload dominated by reasoning, it is close to a 30% cut in the expensive half of the bill, and it arrives without a rate change. That is the opposite of the usual pattern, where a newer model costs more per token and you argue about whether it earns the difference.

It is also the claim you can check yourself most easily. Run the same repository task on both, count output tokens, and the ratio is either there or it is not.

High-speed costs exactly double, everywhere

The hosted high-speed variant, Kimi K2.7-Code-Highspeed, called with kimi-k2.7-code-highspeed, is not a different model. Same 262,144-token context, same 1T/32B architecture, and every published rate is precisely twice the base.

Published rates per 1M tokens
Ratekimi-k2.7-codehigh-speedMultiple
Input, cache miss$0.95$1.902.0x
Input, cache hit$0.19$0.382.0x
Output$4.00$8.002.0x
Context window262,144262,144same

A flat 2x across every line makes the decision unusually clean: you are buying latency and nothing else, at a fixed and legible price. There is no capability argument to have, no context trade, and no benchmark difference to weigh. Either the wall-clock time is worth double, for this workload, or it is not.

Moonshot publishes no maximum-output figure and no license for the high-speed variant on the pricing page, so benchr records neither. The base model's modified-MIT weights are not evidence about a hosted endpoint's terms.

Cache hits are where the input bill actually lives

Cache-hit input is $0.19, exactly a fifth of the $0.95 miss rate. For coding work that repeatedly sends the same repository context, the difference between a well-structured prompt prefix and a poorly structured one is a 5x swing on the input side.

That is a bigger lever than the choice between the two variants for most teams, and it is entirely under your control. Measure the hit rate before comparing models; a low hit rate makes every model look expensive and hides which one is actually cheaper for the work.

The published scores, and what they are

Six figures come with the model card. All six are Moonshot's own; benchr has reproduced none of them.

Provider-reported results for Kimi K2.7-Code
BenchmarkScoreK2.6
MCP Mark Verified81.1not published
MCP Atlas76.0not published
Kimi Code Bench v262.050.9
Program Bench53.6not published
Kimi Claw 24/7 Bench46.942.9
MLS Bench Lite35.1not published

Two of the six carry a K2.6 comparison, and those are the informative ones: Kimi Code Bench v2 rises 50.9 to 62.0, Kimi Claw 24/7 rises 42.9 to 46.9. Note that the largest published gain is on Moonshot's own named benchmark, which is the normal shape of a vendor release and a reason to weight your own tasks over any of these.

Who this fits

What the official record supports
If you areThe record says
Running long-horizon coding agents and watching output costTest it. The thinking-token reduction is the claim, and output is the expensive half
Blocked on latency rather than costThe high-speed variant is a clean 2x for speed alone, with no other difference
Required to self-host or audit weightsThe base model publishes modified-MIT weights at 1T total, 32B active
Sending large repeated repository contextFix the cache hit rate first; $0.19 against $0.95 is a bigger lever than the model choice
Choosing on published benchmarks aloneDo not. Four of six have no prior-generation comparison and the largest gain is on the vendor's own suite

Frequently asked

What does Kimi K2.7-Code cost?

$0.95 per 1M input tokens on a cache miss, $0.19 on a cache hit, and $4.00 per 1M output tokens. The hosted high-speed variant is exactly double each: $1.90, $0.38 and $8.00.

What is the difference between K2.7-Code and the high-speed variant?

Price. Both publish a 262,144-token context and the same 1T total / 32B active architecture, and every high-speed rate is precisely twice the base. You are buying latency, which benchr has not measured.

Are the weights open?

The base model's card publishes modified-MIT weights at 1T total and 32B active parameters. Moonshot publishes no license for the hosted high-speed endpoint, so benchr records none for it.

Is the 30% thinking-token reduction a benchr measurement?

No. It is Moonshot's claim against Kimi K2.6, as are the two K2.6 benchmark comparisons. benchr has run no tests against either variant.

Changelog

  • September 8, 2026 — Published after rechecking the pricing page and model card, and after adding the six provider-published benchmark figures to the model record.

References

  1. Official pricing documentation: https://platform.kimi.ai/docs/pricing/chat-k27-code
  2. Official model card: https://huggingface.co/moonshotai/Kimi-K2.7-Code
  3. Official model documentation: https://platform.kimi.ai/docs/models