GPT-6 Astra API pricing

A dated GPT-6 Astra rate card for Standard, long-context, Batch, Flex, and Fast processing, with the whole-request 272K boundary made explicit.

By benchr Editorial Team ·

Standard / 1M$10 / $50input / output
Cached / write$1 / $12.50per 1M
Long-context>272Kwhole request
Context1.05M128K output

Official rate table

Rates derived directly from OpenAI's published base prices and processing multipliers; source checked September 5, 2026
LaneInput / 1MCached / 1MCache write / 1MOutput / 1M
Standard · ≤272K$10$1$12.50$50
Standard · >272K$20$2$25$75
Batch / Flex · ≤272K$5$0.50$6.25$25
Batch / Flex · >272K$10$1$12.50$37.50
Fast · ≤272K$20$2$25$100
Fast · >272K$40$4$50$150

Crossing 272K changes every token in the request

The premium does not apply only to the excess. OpenAI says prompts with more than 272K input tokens use 2x input and cache rates and 1.5x output for the full request. The table applies that rule, plus the documented 50% Batch/Flex and 2x Fast multipliers. Tool-call fees are separate.

Frequently asked questions

What is the Standard GPT-6 Astra price?

$10 input, $1 cached input, $12.50 cache writes, and $50 output per million tokens.

Are Batch and Flex the same price?

OpenAI documents both at 50% of Standard rates.

How much does Fast mode cost?

Fast is 2x the applicable rate, including the long-context rate when the prompt exceeds 272K input tokens.