NEWS · JULY 31, 2026 · ARTIFICIAL INTELLIGENCE

OpenAI cut the GPT-5.6 Luna API price by 80 percent

On July 30, 2026 OpenAI cut the GPT-5.6 Luna API price by 80 percent and the Terra price by 20 percent. The flagship Sol price did not change. The same day, the Priority Processing service tier was renamed Fast mode. The gap between the cheapest tier and the flagship widened from five times to twenty five times.

01 · WHAT HAPPENED?

The cut took effect the same day

OpenAI announced it in a single line on its own API changelog: "Starting July 30, GPT-5.6 Luna costs 80% less, while GPT-5.6 Terra costs 20% less." The entry is dated July 30, 2026 and defines no transition period, so the new prices applied on the day of the announcement.

The numbers run as follows. On GPT-5.6 Luna, a million input tokens went from $1.00 to $0.20 and a million output tokens from $6.00 to $1.20. On GPT-5.6 Terra, input went from $2.50 to $2.00 and output from $15.00 to $12.00. The flagship GPT-5.6 Sol was left alone at $5.00 per million input tokens and $30.00 output. The new figures are confirmed on OpenAI's official pricing documentation, and the pre-cut figures in independent reporting.

The timing is worth noting. Per OpenAI's own changelog the GPT-5.6 family was released on July 9, 2026, so the cut arrived exactly three weeks later.

02 · DETAILS

The headline number is not the whole bill

Those prices are the short context rates. OpenAI's pricing documentation lists a separate and higher long context tariff: on Luna $0.40 per million input tokens and $1.80 output, on Terra $4.00 and $18.00, on Sol $10.00 and $45.00. A long context Luna call therefore costs double the input rate in the headline. The coverage we reviewed reported only the short context number.

The same document defines variance in two more directions. Cached input drops to a tenth of the input rate, which is $0.02 per million tokens on Luna. Batch and Flex tiers are half of standard, so Luna comes to $0.10 input, $0.60 output and $0.01 cached input. Fast mode is twice standard across all three tiers. The upshot is that one model can bill anywhere between $0.01 and $0.40 per million input tokens depending on how the call is made, and the 10 percent uplift defined for regional processing endpoints can push that ceiling slightly higher.

Fast mode itself is not a new product but a rename. OpenAI's pricing page states that Priority Processing was renamed Fast mode on July 30, 2026, and that both "priority" and "fast" remain accepted in the service_tier field. Existing integrations need to change nothing. For Sol, the changelog describes up to 2.5 times faster speeds than standard processing at twice the price. Sol's price held steady, but the tier did not stand entirely still: a speed improvement delivered through Fast mode sits on that side.

The reason for the cut is efficiency rather than capacity. As The Decoder reports it, OpenAI attributes the gain to GPT-5.6 Sol optimising its own GPU software: a 20 percent cut in deployment costs and more than 15 percent better token generation through speculative decoding. The same source notes that the price pressure comes from cheaper Chinese models and Microsoft's lower cost MAI models, and cautions that aggressive pricing carries a risk if it slows revenue growth across the sector.

03 · WHY IT MATTERS

The spread between tiers widened from five to twenty five

Before July 30, Sol's input price was five times Luna's. After the cut it is twenty five times. On its own that produces an architectural consequence: the routing layer deciding which job goes to which model is worth considerably more than it was three weeks ago. A setup that sends bulk repetitive work to Luna and only genuinely hard tasks to Sol is now a serious lever on cost.

To make the number concrete: a job costing $1.00 per million input tokens on Luna on July 9 now costs $0.20, and $0.10 with the Batch discount. For high volume repetitive work such as generating catalogue copy, classifying inbound messages or summarising long records, that range changes the economics outright. The analyst commentary InfoWorld reported points the same way: moving pilots into production gets easier, and multi step agent flows that were hard to justify economically start to add up.

There is also a signal about where the competition is being fought. OpenAI cut the cheap tiers, left the flagship price untouched and made the existing paid speed tier faster on the flagship. It is defending price at the top and competing on price at the bottom. What it did not do is simplify: with short versus long context, caching, Batch, Fast mode and a 10 percent uplift defined for regional processing all in play, a budget built on the headline $0.20 will be wrong for long context or Fast mode traffic.

04 · TÜRKİYE

What it means for businesses in Türkiye

The assessment below is our reading, not something stated in the sources. None of the sources we reviewed contains a Türkiye specific breakdown, a lira equivalent or a local availability note. The only geography related entry in the pricing documentation is a 10 percent uplift on regional processing endpoints, and it names no country. So this is a change to the global tariff rather than a separate announcement for calls made from Türkiye.

Four practical notes. First, routing: because the spread between tiers has widened, a system built on a single model is leaving more money on the table than before, and separating bulk work onto the cheap tier is the fastest gain. Second, Batch: night jobs, catalogue updates and bulk classification that are not latency sensitive halve the bill once they move to the Batch tier. Third, measurement: build the budget on your own mix of short and long context traffic rather than the headline price, or forecast and invoice will drift apart. Fourth, an assumption to avoid: the sources state no end date or introductory period for the new prices, but it should not be read as a permanent commitment either, because a line item divided by five in three weeks can move the other way. That last sentence is our own comment.

The UNALSOFT take

In our agentic AI work, model selection is not a brand preference but a decision made inside the flow. Which step is solved by the cheap tier, which step genuinely earns the strong model, and the handover point between them are all designed up front. This cut raises the value of that design: it is a good week to recalculate the flows that were shelved on cost grounds. The question to ask is not which model to use, but which job belongs on which tier.

Cost came down, the decision sits in the architecture.

Let's work out together which job belongs on which tier.

Message on WhatsApp