GPT-6 Luna
An efficient reasoning model for focused, high-volume tasks, with text and image input, a 1.05M-token context window, and adjustable reasoning effort.
Source: Official model specifications and supported features, API release date: September 22, 2026, 1024-token cache minimum and default 30-minute minimum lifetime, Reasoning token billing, Pricing, context tiers, and service tiers
- Text
- Image
- Text
Pricing
Every billing leg from the live catalog, in USD per 1M tokens.
| Billing leg | Standard | Long context | Batch |
|---|---|---|---|
| Input | $0.1 / M | $0.2 / M | $0.05 / M |
| Output | $0.5 / M | $0.75 / M | $0.25 / M |
| Cache read | $0.01 / M | $0.02 / M | $0.005 / M |
| Cache write | $0.125 / M | $0.25 / M | $0.0625 / M |
Long-context rates apply above 272,000 input tokens.
Cost calculator
Billing-grade math over the rate card above, including the thinking tokens most estimates miss.
Hidden reasoning tokens bill as output. 1.0x = thinking off. Medium-effort reasoning typically lands near 3x; heavy reasoning runs 6x to 8x. Set it to match your workload.
Provider
Where this model runs, and what it charges.
| Platform | Input / M | Output / M | Cache read / M | Batch | Access |
|---|---|---|---|---|---|
OpenAI | $0.1 | $0.5 | $0.01 | ✓ | First-party API |
Price history
Has this model gotten cheaper or more expensive?
More from OpenAI
Other models from the same provider.
Stop estimating. Track it.
SuperPenguin meters every GPT-6 Luna call your team makes, prices it with this exact rate card, and shows spend by feature, customer, and team.
Rates are generated from SuperPenguin's live pricing catalog.