What actually changed in AI model pricing on 30 July 2026?
OpenAI cut GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20%, effective 30 July 2026. Luna now costs US$0.20 per million input tokens and US$1.20 per million output tokens. Terra costs US$2.00 and US$12.00. Sol pricing is unchanged.
Here is the honest framing before the numbers. Nobody switches their default model because a blog post told them to. You switch when the cheaper option stops being obviously worse at your specific work, and the July cut is the first time that threshold has moved for a large group of practitioners.
The catch is that "cheaper" and "cheap enough" are different questions, and most of the coverage of this announcement answers only the first one.
This page answers the second: what you actually pay per task, where the cheap tier genuinely fails, and which of these models deserves to be your default given the kind of work you do.
How much do GPT-5.6 and Claude models cost per million tokens?
As of 31 July 2026, GPT-5.6 Luna is US$0.20 input and US$1.20 output per million tokens, Terra is US$2.00 and US$12.00, and Sol is listed at US$5.00 and US$30.00. Anthropic's Claude Opus 5, launched 24 July 2026, is US$5.00 and US$25.00. Claude Fable 5 is US$10.00 and US$50.00.
Current list prices, per million tokens
--- GPT-5.6 Luna: US$0.20 in / US$1.20 out (was US$1.00 / US$6.00 before 30 July)
--- GPT-5.6 Terra: US$2.00 in / US$12.00 out (was US$2.50 / US$15.00)
--- GPT-5.6 Sol: US$5.00 in / US$30.00 out (unchanged)
--- Claude Opus 5: US$5.00 in / US$25.00 out (launched 24 July 2026)
--- Claude Fable 5: US$10.00 in / US$50.00 out
At an exchange rate of roughly HK$7.8 to the US dollar, Luna's output tokens cost about HK$9.40 per million and Fable 5's cost about HK$390 per million. That is a 41-fold spread across models you can call from the same script.
OpenAI's own announcement also replaced Priority Processing with Fast mode for Sol, which delivers up to 2.5 times faster responses at twice the price with no change in intelligence. That is a latency purchase, not an intelligence purchase, and it is easy to misread as an upgrade.
One important note on scope. These are API and platform token prices. If you pay a flat monthly subscription for ChatGPT or Claude, this cut changes your quota consumption rather than your bill. For the per-seat side of the decision, see our breakdown of enterprise AI seat pricing.
What does one real task actually cost after the cut?
A typical practitioner task, summarising a 20-page document and drafting a 900-word output, consumes roughly 15,000 input tokens and 1,500 output tokens. On Luna that is about US$0.0048, or HK$0.04. On Terra it is about US$0.048, or HK$0.37. On Opus 5 it is about US$0.11, or HK$0.87.
Scale that to a real workload and the picture sharpens.
Cost of 500 such tasks per month, at HK$7.8 to the US dollar
--- GPT-5.6 Luna: about US$2.40, or HK$19
--- GPT-5.6 Terra: about US$24, or HK$187
--- GPT-5.6 Sol: about US$83, or HK$645
--- Claude Opus 5: about US$56, or HK$437
--- Claude Fable 5: about US$150, or HK$1,170
These are arithmetic from published list prices, not benchmarks, and your real token counts will differ with prompt length and reasoning depth. Treat them as an order-of-magnitude guide rather than a quote.
The number that should reset your intuition is the gap between the first and last line. Running the same 500 tasks on Fable 5 instead of Luna costs roughly 60 times more per month. For most content, classification and summarisation work, that premium buys very little.
OpenAI's own framing supports this. The company states that Luna delivers performance comparable to models that were frontier-class a year ago at roughly six cents on the dollar per task, and at nearly nine times the speed.
Which model should be your default, by type of work?
Default to Luna for high-volume, well-specified work where errors are cheap to catch. Default to Terra for everyday professional output where a human reviews the result. Reserve Sol, Opus 5 or Fable 5 for work where a wrong answer costs more than the model does, which is a much smaller share of tasks than most people assume.
Verdict by user type
--- High-volume content or ops work (tagging, classification, summarising, first-pass drafting, data cleanup): Luna. The quality gap on well-specified tasks no longer justifies paying ten times more.
--- Client-facing writing and analysis reviewed by a human: Terra. Notion reported that Terra delivered comparable quality to GPT-5.5 at half the cost per task and in 60% less time in their evaluations.
--- Long-document reasoning and nuanced editing: Claude Opus 5. At US$5 and US$25 it undercuts Fable 5 by half, and Anthropic positions it as close to Fable 5's frontier intelligence.
--- Work where a single wrong answer is expensive (legal summaries, financial figures, regulatory text, anything published unreviewed): Sol or Fable 5, and add human review anyway.
--- Latency-critical interactive work: Sol with Fast mode, accepting double the token price for up to 2.5 times the speed.
The pattern worth internalising is that the correct answer is rarely one model. Multiple teams cited by OpenAI describe routing: a strong model to define the plan, a cheap model to execute the well-specified parts. Cognition built Luna into Devin Fusion specifically to handle routine work alongside larger models.
How do you test whether the cheap model is good enough?
Take twenty real tasks you have already completed, run them through the cheaper model, and count how many outputs you would have sent without editing. If the pass rate is above roughly 80% and the failures are obvious rather than subtle, the cheaper model is safe to default to.
The subtlety matters more than the rate. A model that fails loudly is safer than one that fails plausibly, because you will catch the first and ship the second.
Try this prompt to run the comparison:
--- You are being evaluated for a routing decision. I will give you a task I have already completed to a professional standard.
--- TASK: [paste your real task and source material here]
--- Produce your best output for this task.
--- Then, separately, answer three questions about your own output. First, list anything you were uncertain about and would want a human to verify. Second, state which claims in your output you cannot support from the source material provided. Third, rate your confidence from 1 to 5 and explain the rating in one sentence.
--- Do not soften the self-assessment. An honest low rating is more useful to me than a confident wrong answer.
The self-assessment section is what turns this into a routing test rather than a quality test. A cheap model that reliably flags its own uncertainty can be trusted with more work than an expensive model that never does.
Where does the cheap tier actually fall down?
Cheap models fail on tasks requiring sustained multi-step reasoning, ambiguous instructions that need interpretation, domain judgement where the right answer is contested, and any work where a plausible-sounding error survives review. Price cuts do not change these failure modes. They only change how much you pay to encounter them.
Four specific limitations are worth naming honestly.
Ambiguity is where cheap models lose most. Given a vague brief, a stronger model asks a clarifying question or hedges; a cheaper one commits confidently to one reading. If your prompts are loose, the savings evaporate into rework.
Output token counts vary more than you expect. A model that writes verbosely can cost more in practice than a nominally pricier model that writes tightly. Blitzy reported that Luna handled 2.2 times more context with 8.5 times fewer output tokens than their previous model, which is a reminder that per-token price is only half the equation.
Fast mode is a speed purchase, not a quality one. Paying twice the Sol price for up to 2.5 times the speed makes sense for interactive products and makes no sense for batch work that runs overnight.
Prices move. Luna's price dropped by 80% in a single announcement. Any routing decision you hard-code today should be revisited quarterly, and any cost model you build should read prices from a config file rather than from your memory of this article.
One more caveat specific to this page. UD does not publish a fixed price for AI workflow deployment because the work is scoped per client, and quoting a number here would be inventing one. The model prices above are published, verifiable and current as of 31 July 2026. Anything you read about UD's own pricing should come from a conversation, not from a table.
What is the correct next step?
Run the twenty-task comparison before changing anything. If Luna clears your bar, switch your high-volume work and keep a stronger model for the tasks that failed. If it does not clear the bar, you now know exactly which capability you are paying for, which is more useful than a general sense that expensive models are better.
The practitioners who benefit most from a price cut are not the ones who switch fastest. They are the ones who already know which of their tasks are well-specified enough to route cheaply, because that knowledge is what turns a price change into a saving.
Model prices will keep falling. What will not fall is the cost of not knowing which of your own workflows are reliable enough to automate. That is the part worth working on this quarter.
We know AI's cold edges. We know your real challenges. 28 years with UD, turning technology into a partnership with warmth.
Work Out What Your AI Workload Should Actually Cost
Choosing a model is one decision. Building a workflow that routes tasks to the right model, and keeps running when prices change, is the harder one. We'll walk you through every step, from task audit and model routing to deployment and cost monitoring.
Written and reviewed by the UD AI team, Hong Kong. Prices verified against OpenAI and Anthropic published rates on 31 July 2026.