HomeBlog › Three Price Cuts in 72 Hours
AI Pricing

Three Price Cuts in 72 Hours

OpenAI cut GPT-5.6 Luna, Anthropic set Opus 5 at roughly half of Fable 5, and DeepSeek shipped V4 under an MIT license, all inside three days. That is not a sale. That is a floor being found.

One signal a day. No noise. A 3-minute read when something genuinely shifts.
By Tyron Dizon · August 15, 2026 · 5 min read
OpenAI cut GPT-5.6 Luna, Anthropic set Opus 5 at roughly half of Fable 5, and DeepSeek shipped V4 under an MIT license, all inside three days. That is not a sale. That is a floor being found.
Source: Tech Startups, AI Weekly, LLM Stats (August 2026)

I have watched a lot of AI pricing announcements. Most of them are theater: a lab ships a shiny new model, prices it like a luxury good, then quietly trims six months later when the next one lands. Last week that pattern broke, and I do not think enough people noticed how strange it was.

The week the discount became the price

Inside roughly seventy-two hours, three separate things moved in the same direction. On August 13, DeepSeek released V4 with its weights under an MIT license, which is about as permissive as software licensing gets. On August 14, Anthropic cancelled a planned price increase for Claude Sonnet 5 and made $2 / $10 permanent. The same day, OpenAI cut pricing on GPT-5.6 Luna substantially, and Anthropic positioned Claude Opus 5 at roughly half the price of its higher-end Fable 5.

Three moves. All down. No coordination required, because they are all responses to the same thing.

Who is actually applying the pressure

The reporting is unusually blunt about the cause. Both cuts are attributed to the same force:

Lower-cost Chinese competitors including DeepSeek and Moonshot AI gain users among companies looking to control increasingly large inference bills.

Here is the analogy I keep coming back to. For most of the last century, a long distance phone call was a small, deliberate expense. You watched the clock. Then the underlying cost of moving a voice across a wire collapsed, and within a few years the idea of paying by the minute to talk to another country stopped making sense at all. Nobody announced the end of long distance. The price just fell until the category dissolved into the background.

That is roughly what is happening to inference. The frontier labs spent two years selling something that felt scarce and premium. Then competitors arrived selling something good enough for a large share of the work at a fraction of the price, and in DeepSeek's case published the recipe. Once a serious rival gives the thing away, the question stops being what is this worth and becomes what does it cost to serve. Those are very different numbers.

What gets cheap does not stay special

If you are building anything on top of these APIs, this cuts two ways at once.

The good news is that the bill for work you already do just went down without you touching a line of code. Anything you designed under 2025 token prices is quietly cheaper to run today than it was last month.

The uncomfortable news is that everyone else got the same discount. If the core of what you sell is access to a model with a markup on top, that spread is compressing toward zero and it is not coming back. The durable parts are the ones that do not get cheaper: the workflow around the model, the data only you have, the integration into systems that are genuinely painful to connect, and the judgment about what to build in the first place. The model was never the hard part. It just used to be the expensive part, and expensive is easy to mistake for defensible.

The thing cheap tokens quietly unlock

There is a second story from the same week that only makes sense alongside the first.

On August 10, Anthropic disclosed that an unreleased research version of Claude improved a longstanding lower bound in number theory: the fraction of Riemann zeta zeros known to satisfy the Riemann Hypothesis went from 41.6% to 67.2%, by synthesizing recent papers. The detail that matters to me is not the mathematics. It is the setup. The work ran inside Claude Code, across two sessions, burning 31 million output tokens.

Strip out the number theory and the shape is this: an agent, in an ordinary coding harness, given a very long leash and a serious token budget, produced something that had resisted people for a long time. Not a bespoke research supercomputer. The same kind of tool a lot of us already have open in a terminal.

Most of us have trained ourselves to keep agent runs short. That instinct came from a period when long runs mostly compounded errors and ran up a bill. Both halves of that are weakening at the same time: models hold a thread better than they used to, and the price of thinking for a long time just fell three times in one week. If you have a genuinely hard, well bounded problem that you have been rationing effort on, the arithmetic that told you to keep it cheap may no longer hold.

Ten models, six providers, two weeks

Underneath all of it is churn. Ten models shipped from six providers in the first half of August alone, most recently Gemini 3.7 Flash on August 13, with ByteDance's Seed 2.1 Turbo on August 10 and xAI's Grok Imagine Image 2.0 on August 8.

At that pace, picking a model stops being an architecture decision and becomes an operations one, closer to choosing a shipping carrier than choosing a database. Anything hard coded to one specific model name is technical debt the moment you write it. A thin layer that maps task types to models, changeable in one place, pays for itself in about a month at current release rates.

The part that does not add up yet

Falling prices are a wonderful thing to be a customer of and a difficult thing to be a seller of. OpenAI's last private valuation was $852B against roughly $2B in monthly revenue, and reported losses of about $1.22 for every $1 earned. The company confidentially filed for an IPO on June 8 and the filing has not surfaced publicly yet.

So prices are being cut into a market where at least one leader loses money on every dollar of revenue. That can persist for a long time. It cannot persist forever. The signal I am watching for is simple and it is the opposite of last week: the first time a major lab raises a production tier price, or moves to per seat licensing that is not metered by tokens. As long as everything moves down, this is a land grab and building aggressively is rational. The first move up is the moment the land grab ends and the lock in phase starts.

Until then, the tokens are cheap, the models are new every fortnight, and the ceiling on what an agent can do with a long enough run is higher than most of us have bothered to test. That is a good week to be building.

Three downward moves in 72 hoursFrontier model pricing, August 13 to 14, 2026AUG 13DeepSeek V4 weights released under MIT licenseAUG 14Sonnet 5 price rise cancelled, $2 / $10 permanentAUG 14GPT-5.6 Luna cut; Opus 5 at about half of Fable 510models shippedfrom 6 providersin the first halfof August 2026ALL PRICE MOVES: DOWNSource: Tech Startups, AI Weekly, LLM Stats, August 2026
Source: Tech Startups, AI Weekly, LLM Stats (August 2026)

One signal a day. No noise.

A 3-minute read when something genuinely shifts in AI, automation, or defense tech. Free, most weekdays.

Free, most weekdays. No spam, unsubscribe anytime.

Sources

  1. Tech Startups - Top tech news today, August 14, 2026 - https://techstartups.com/2026/08/14/top-tech-news-today-august-14-2026-apple-anthropic-deepseek-google-ibm-pony-ai-openai-spacex-uber-more/
  2. AI Weekly - AI news today - https://aiweekly.co/ai-news-today
  3. AI Weekly - Anthropic news - https://aiweekly.co/ai-news-today/anthropic-news
  4. LLM Stats - AI news - https://llm-stats.com/ai-news
  5. LLM Gateway - model release timeline - https://llmgateway.io/timeline
  6. Tech Journal - OpenAI IPO, what to expect from the public S-1 - https://techjournal.org/openai-ipo-public-s1-what-to-expect
  7. Yahoo Finance - OpenAI confidentially files for IPO with SEC - https://finance.yahoo.com/markets/stocks/articles/openai-confidentially-files-ipo-sec-223341186.html

Quick answers

Why did OpenAI and Anthropic cut AI prices in the same week?

Reporting attributes both moves to the same pressure: lower-cost competitors including DeepSeek and Moonshot AI winning users among companies trying to control large inference bills. OpenAI cut GPT-5.6 Luna pricing substantially and Anthropic positioned Claude Opus 5 at roughly half the price of Fable 5.

What did Anthropic decide about Claude Sonnet 5 pricing?

Anthropic cancelled the planned Sonnet 5 price increase and made $2 / $10 permanent, one day before the wider round of cuts.

What was the Riemann Hypothesis result Anthropic disclosed?

On August 10, 2026, Anthropic disclosed that an unreleased research version of Claude improved the longstanding lower bound on the fraction of Riemann zeta zeros satisfying the Riemann Hypothesis from 41.6% to 67.2% by synthesizing recent papers. The work ran inside Claude Code across two sessions and used 31 million output tokens.

How fast are new AI models being released in 2026?

Ten models shipped from six providers in the first half of August 2026 alone, including Gemini 3.7 Flash on August 13, ByteDance's Seed 2.1 Turbo on August 10, and xAI's Grok Imagine Image 2.0 on August 8. At that pace, model selection behaves more like an operations decision than an architecture one.

Tyron Dizon is a Chief Product Officer, AI product builder, and Techstars-backed SaaS founder based in Baguio City, Philippines. He previously co-founded and served as CPO of SanityDesk and now builds AI products, automation systems, SaaS platforms, and rapid prototypes. About · Work · Resume · LinkedIn