
The workload
Any SaaS product that calls OpenAI's API has to price its own subscription around a moving input cost, because the per-token rate is the largest variable cost in an AI feature and it can change without advance notice to the founder building on it. Verified: OpenAI's pricing page, retrieved 16 September 2026, lists GPT-6 Astra, described as the company's most capable model, at $10.00 per million input tokens, $1.00 per million cached-input tokens, and $50.00 per million output tokens.
What the documents show
Verified: OpenAI's own pricing documentation, also retrieved 16 September 2026, states the same standard rate for GPT-6 Astra and adds a long-context tier, priced at $20.00 per million input tokens, $2.00 per million cached-input tokens and $75.00 per million output tokens once a request crosses the model's context threshold. The documentation also states that regional processing, or data-residency, endpoints carry a 10% uplift on eligible models released after 5 March 2026. The pricing page separately lists a second current model, GPT-5.6 Sol, at $5.00 per million input tokens and $30.00 per million output tokens, roughly half the flagship model's input rate, confirming that model choice, not only usage volume, is a direct lever on cost.
The operating cost
Verified: for GPT-6 Astra, a request using 1 million uncached input tokens and 1 million output tokens costs $10.00 plus $50.00, or $60.00 total, before any long-context or regional surcharge; using cached input at $1.00 per million instead of fresh input drops the input-side cost to a tenth of the standard rate. This is an estimated calculation applying the published rate to round numbers; actual per-customer cost depends on token volume per request, cache-hit rate, and whether requests cross into the long-context or regional-uplift tiers, none of which the pricing page itself can specify for a given product.
The stop condition
This is editorial, since the documents name no threshold of their own: a subscription price built around today's per-token rate should be revisited whenever OpenAI changes the listed price for the model in use, since the page carries no notice period and no prior rate is preserved on it once superseded.
- Which specific model is the product actually calling, and does its rate match the one in the founder's own cost model?
- What share of requests are realistically served from cache, and has that share been measured rather than assumed?
- Is the product exposed to the long-context or regional-uplift tiers, and has that surcharge been priced into the subscription?
A per-token rate is a snapshot, not a contract, and the only defensible practice is re-checking it against the live page at the same cadence a founder reviews their own subscription pricing.
Sources & reading trail
Lists per-token input, cached-input and output rates for the current flagship and secondary models.
Source published: Not established · Retrieved: 16 September 2026
Confirms the same standard rate, adds the long-context tier pricing, and states the regional-processing uplift.
Source published: Not established · Retrieved: 16 September 2026
Vendor documentation, regulator records and founder-published documents establish the entry; the workload reading and the stop condition are Solo Product Office editorial analysis. This retrospective draft does not imply the site published on the event date.