Just three days after Chinese AI startup Moonshot AI released Kimi K3 open weights, the largest free AI model ever at 2.8 trillion parameters, OpenAI is making a move of its own. The company has slashed prices on two GPT-5.6 models by as much as 80%, signaling that the AI race is no longer defined solely by who builds the biggest or smartest model. It’s becoming a contest over who can deliver the most intelligence for the lowest cost.
The timing is hard to ignore. Kimi K3 pushed free, open-weight AI into territory once dominated by a handful of well-funded frontier labs, raising fresh pressure on proprietary model providers to prove their value. OpenAI’s sweeping price cuts arrive just three weeks after GPT-5.6 launched, reflecting a broader shift across the industry as AI companies compete on price, efficiency, and real-world performance as much as raw capability.
OpenAI said on July 30 that it is reducing prices across much of the GPT-5.6 family, with the changes taking effect immediately. GPT-5.6 Luna, the company’s fastest and lowest-cost model, receives the largest adjustment. Input pricing falls from $1 to $0.20 per million tokens, an 80% reduction, and output pricing drops from $6 to $1.20 per million tokens.
In an announcement on X, OpenAI said:
“We are committed to pushing the model frontier across cost efficiency, capability, and speed. Starting today, we are reducing prices for GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20%, and offering a faster option for GPT-5.6 Sol in the API. Luna and Terra’s lower prices are reflected in how usage is counted in Codex and ChatGPT Work, so your usage goes further.”
We are committed to pushing the model frontier across cost efficiency, capability, and speed.
Starting today, we are reducing prices for GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20% , and offering a faster option for GPT-5.6 Sol in the API.
Luna and Terra’s lower prices are… pic.twitter.com/rFhK7XKedp
— OpenAI (@OpenAI) July 30, 2026
GPT-5.6 Terra, positioned as the everyday production model, now costs $2 per million input tokens and $12 per million output tokens, down from $2.50 and $15. OpenAI left pricing for GPT-5.6 Sol unchanged at $5 per million input tokens and $30 per million output tokens.
Artificial Analysis Intelligence Index Score
The AI price war is entering a new phase
The company introduced a new Fast mode for Sol through its API, replacing Priority Processing. OpenAI says the new option delivers up to 2.5 times the processing speed of standard requests at twice the standard price, with no loss in model intelligence. The discounted pricing for Luna and Terra carries over to ChatGPT Work and Codex usage, allowing organizations to consume fewer credits for the same workloads. Auto-review capabilities in the ChatGPT app and Codex CLI are moving from GPT-5.4 to GPT-5.6 Luna, a change OpenAI expects will reduce costs by roughly tenfold.
Rather than crediting cheaper hardware or lower demand, OpenAI says the savings came from improvements inside the system itself. According to the company, engineers used GPT-5.6 Sol to optimize its own serving infrastructure, lowering production serving costs by 20% through GPU kernel improvements and improving token generation efficiency by more than 15% with enhanced speculative decoding.
OpenAI says the gains extend beyond model optimization. Better hardware routing, improved inference software, smarter context management, and models that complete tasks with fewer unnecessary steps all contributed to lower operating costs. In one customer example shared by the company, GPT-5.6 Luna processed 2.2 times more context while generating 8.5 times fewer output tokens than GPT-5.4 mini, reducing overall costs by 87%.
Why OpenAI is cutting GPT-5.6 prices just weeks after launch
The pricing changes arrive shortly after the GPT-5.6 family became generally available on July 9 following an earlier preview period. OpenAI introduced three tiers to cover different workloads. Sol targets advanced coding and agentic tasks, Terra serves as the primary production model for most applications, and Luna focuses on high-volume deployments where cost efficiency matters most. All three models share the same large context window and knowledge cutoff around mid-February 2026.
The timing reflects a broader shift taking place across enterprise AI. Organizations are moving beyond experimentation and paying closer attention to operating costs. Agentic systems often consume far more tokens than traditional chatbot interactions, making inference expenses a growing concern for businesses deploying AI at scale. Lower token prices can significantly reduce the cost of running autonomous workflows over weeks or months.
Competitive pressure is building from multiple directions. Open-weight models continue improving, giving developers lower-cost alternatives for many workloads. Frontier AI companies are responding by squeezing more performance from existing infrastructure instead of relying solely on larger training runs or new hardware. That approach allows providers to pass efficiency gains directly to customers without waiting for the next generation of GPUs.
OpenAI highlighted benchmark data based on the Artificial Analysis Intelligence Index v4.1, showing GPT-5.6 Luna delivering the highest intelligence score among compared models at a relatively low cost per task. Customer feedback included in the announcement pointed to higher cache reuse rates, lower operating expenses, and the ability to deploy agentic workflows that had previously been too expensive to justify.
The broader message extends beyond one pricing update. AI companies spent much of the past two years competing to build the smartest models. The conversation is shifting toward who can deliver the best value per dollar. Faster inference, lower token consumption, and better infrastructure efficiency are becoming competitive advantages in their own right.
OpenAI’s latest price cuts suggest that intelligence alone is no longer enough. For developers and enterprises, the winning model may be the one that delivers strong performance without turning every AI-powered workflow into an expensive monthly bill.



