OpenAI Cuts GPT-5.6 Luna Price by 80%
OpenAI cut GPT-5.6 Luna by 80 percent and Terra by 20 percent on July 30, three weeks after launch. What the new numbers mean for a small monthly bill.
OpenAI cut the price of two GPT-5.6 models on July 30, 2026. The cheapest tier, Luna, dropped 80 percent to $0.20 per million input tokens and $1.20 per million output, down from $1 and $6. The mid tier, Terra, fell about 20 percent to $2 and $12, from $2.50 and $15. The flagship, Sol, was left alone at $5 and $30. The cuts were reported by CNBC and confirmed across subsequent coverage. Two weeks on, the numbers below have held, with no further change to Luna or Terra pricing as of mid-August 2026.
The timing was the unusual part. GPT-5.6 launched on July 9, 2026, which made this a repricing three weeks into a model family life rather than the usual slow drift downward.
The new numbers
Model | Input, per million tokens | Output, per million tokens | Change |
|---|---|---|---|
Luna | $1.00 to $0.20 | $6.00 to $1.20 | Down 80 percent |
Terra | $2.50 to $2.00 | $15.00 to $12.00 | Down about 20 percent |
Sol | $5.00 | $30.00 | Unchanged |
Why now
Reporting attributes the move to competitive pressure at the cheap end of the market rather than to a change in how the models work. Analysis published by Yahoo Finance put the new Luna price against DeepSeek V4 Pro at $0.435 and $0.87 per million with a promotional discount applied, and noted Chinese models had reached 46 percent of United States enterprise token usage on the OpenRouter platform. OpenAI framed the reduction as passing on efficiency gains in how it serves the models.
What this changes for someone building small things
Headline percentages are less useful than arithmetic on a real workload, so here is one.
Take a modest internal tool that sends a 4,000 token prompt and gets back 800 tokens, run 500 times a month. On Luna at the old price that is $2 of input and $2.40 of output, about $4.40 a month. At the new price it is $0.40 and $0.96, about $1.36. On Sol, unchanged, the same workload is $10 and $24, about $34.
Two things follow from those numbers. The first is that at the cheap end, model cost has stopped being the thing that decides whether a small project is viable. A tool costing under two dollars a month to run is effectively free next to the time spent building it. The second is that the gap between tiers is now large enough to be worth thinking about: Sol costs roughly 25 times what Luna does per token, so putting the strongest model on a task that does not need it is a real decision rather than a rounding error.
The multiplier that catches people is not the per token price but the volume, and volume is mostly driven by how much context you send on every single turn. A long running conversation re-sends its entire history each time, which is why the cost of a session grows faster than the number of questions in it. That mechanic is worth understanding before you leave anything running in a loop, and we covered it in how context windows work.
Why the cheap tier is where the fight is
Frontier pricing has held reasonably steady while the bottom of the market has fallen hard. That split makes sense once you look at what the tiers are for. The flagship is bought by people who need the best answer and will pay for it. The cheap tier is bought by people running high volume, repetitive work where the model is a component rather than the product, and in that market the buyer will switch providers over a price difference without much loyalty.
For anyone building on top of these models, the practical implication is to avoid assuming today prices in any plan longer than a quarter. They have moved repeatedly, and consistently in one direction.
What to do about it
If you built something on the cheap tier before July 30 and never revisited it, your bill has already fallen. Nothing to do.
If you defaulted to a flagship model out of caution, the price gap is now wide enough to justify testing whether a cheaper tier handles the task.
If you are choosing a model for a new project, pick on output quality for your specific task first and price second. The prices move. Your requirements do not. For a fuller breakdown by task type, see which AI model to use for which task.
If you have not checked what you are actually spending across providers recently, this is a reasonable time to do it. See how to audit your AI tool spend in an afternoon.
It is also a reminder that model lineups change underneath you. Providers retire models on published schedules, and an application pinned to a specific version eventually needs attention, as we noted when Claude Opus 4.1 reached its retirement date.
For most people reading this, the honest summary is that inference pricing is no longer the constraint on building something small and useful. The constraint is knowing what to build and specifying it well, which is where our guide to building an app with AI starts, and how that total compares to the alternative is broken down in AI app builder versus hiring a developer.
Price is only one input to picking a model day to day. See how to know when to upgrade to a newer AI model for the rest of the decision.
How did this land?
About the author

Senior Editor, AI & Product
Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.


