OpenAI cuts GPT-5.6 Luna price 80% and Terra 20%, adds Sol Fast mode three weeks after launch
Three weeks after GPT-5.6 reached broad API access, OpenAI cut the cost of its two lower-tier models sharply on July 30 and added a new premium throughput configuration for its flagship, widening the price gap between tiers more than any adjustment at launch.
What changed
Luna, the smallest and fastest model in the GPT-5.6 family, dropped 80%: input tokens fell from $1.00 to $0.20 per million, and output tokens from $6.00 to $1.20 per million, per OpenAI's announcement. Combined input-plus-output cost goes from $7.00 to $1.40 per million tokens.
Terra fell 20%: input from $2.50 to $2.00 per million tokens, output from $15.00 to $12.00 per million tokens, for a combined $14.00. Sol Standard stayed at $5.00 input and $30.00 output, unchanged.
OpenAI also introduced Sol Fast mode, priced at $10.00 per million input tokens and $60.00 per million output tokens. The company said Fast mode delivers up to 2.5 times the throughput of Sol Standard at identical model intelligence, per OpenAI. OpenAI attributed both price cuts to inference efficiency improvements made since the family launched.
Why it matters
At $1.40 combined per million tokens, Luna now costs less than Google's Gemini 3.5 Flash-Lite ($2.80 combined) and a fraction of Gemini 3.6 Flash ($9.00 combined), per VentureBeat's frontier pricing comparison. Teams running high-volume, latency-sensitive workloads such as document routing, summarization, or lightweight real-time agents can now access an OpenAI frontier-series model at price points that were previously available only from smaller providers like DeepSeek and Xiaomi.
Terra's new combined rate of $14.00 matches Google's Gemini 3.1 Pro Preview pricing for context windows up to 200,000 tokens, per the same VentureBeat analysis. That puts Terra in direct price parity with a Google mid-tier offering rather than sitting above it.
Sol Fast points in the opposite direction. At $70.00 combined per million tokens it is the most expensive configuration in the frontier market, according to VentureBeat. It targets latency-sensitive production systems willing to pay for throughput rather than capability.
Context and reactions
GPT-5.6 initially reached a limited set of U.S. government partners in late June 2026 before broader API access opened on July 9. The July 30 price cuts arrived roughly ten days after Google released Gemini 3.6 Flash and Gemini 3.5 Flash-Lite, and close to Anthropic's release of Claude Opus 5 at the same sticker price as Opus 4.8. Sam Altman described the changes as "major price cuts today" on X, per VentureBeat.
Each of the three leading providers is taking a distinct route to lower the effective cost of production AI. OpenAI cut per-token rates directly. Google paired lower prices with token-efficiency claims on agentic workloads. Anthropic held its rates and replaced the underlying model with a more capable version at the same price.
What to watch next
Whether Anthropic and Google respond with direct per-token cuts on Claude Fable 5 or Gemini 3.1 Pro is the clearest next signal. OpenAI's inference efficiency attribution for the Luna and Terra reductions could also become the basis for a deeper technical guide on self-optimizing inference stacks.
Sources
- Advancing the price-performance frontier with GPT-5.6: OpenAI announcement, July 30, 2026
- AI price wars: OpenAI cuts GPT-5.6 Luna prices by 80% as model competition shifts toward cost: VentureBeat, July 30, 2026
- OpenAI cuts GPT-5.6 Luna and Terra prices by up to 80%: Quartz, July 30, 2026
