Your AI Bill Doubles on January 1
Google shipped a near-frontier model at $0.75 per million input tokens and published the date the price doubles. Meanwhile Anthropic started selling capability in exchange for data retention. The model you pick is now a finance and governance decision, not a benchmark decision.

On September 2, Google released Gemini 3.8 Flash at $0.75 per million input tokens and $3.75 per million output tokens. That is cheap. What makes it interesting is not the number, it is the footnote: on January 1, 2027, those prices become $1.50 and $7.50.
Google did not bury that. It published the expiry date alongside the launch. Which means for the first time the industry has handed us something close to a written schedule for how introductory AI pricing actually works.
A price with a fuse on it
Think of an introductory APR on a credit card. The rate on the paperwork is real, you genuinely pay it, and the number in your budget eighteen months from now is a completely different number. Everybody understands this about credit cards. Almost nobody has internalized it about tokens.
This was the second cost floor event in three days. Anthropic's Sonnet 5 introductory pricing ended September 1. Then Google undercut the field with a model it explicitly recommends for software engineering, autonomous agents, and multi-step reasoning, with a one million token context window.
And the benchmarks are not a consolation prize. Gemini 3.8 Flash posted 71% on DeepSWE v1.1 and 89.4% on Terminal-bench 2.1. That second one deserves attention, because Terminal-bench measures something with consequences. A model that is good at using a terminal is a model you can hand real work to, the kind that touches files, calls APIs, and changes state. Frontier-adjacent performance at a fraction of frontier price is how coverage described it, and for once the marketing framing is roughly honest.
So the arithmetic that matters: an agent workload costing $40 a month in tokens on a premium model can drop to single digits here. Then, on New Year's Day, it doubles.
Every introductory price in this market is a four month loan. The capability is permanent. The discount is not.
If you priced anything, internally or for a customer, off September 2026 token costs, your margin halves on January 1 unless something in your setup can swap models without a rewrite.
The other price tag: observability
The same week, Anthropic did something structurally stranger. It shipped Claude Fable 5.1 and Claude Mythos 5.1, and the important detail is that they are the same underlying model.
Fable 5.1 ships with safeguards that block or limit performance in risky areas like cyber and bio. Mythos 5.1 is the ungated version, and using it requires accepting a 30-day data retention policy for safety monitoring by default. Fable 5.1 also arrived cheaper and with fewer false positives on safety refusals, which is a genuine quality of life improvement for anyone whose content sits near a sensitive topic and keeps getting refused for no good reason.
Picture two identical cars on the lot. One has a speed governor fitted. You are allowed to buy the one without it, but only if you agree to a dashcam that uploads. Nothing about the engine changed. What changed is the terms.
That is a new commercial pattern and it deserves a name: capability sold in exchange for observability. And notice where it puts the decision. "Which model should we use" used to be a performance question you settled with a benchmark table. It is now a data governance question, and the person who has to answer it is whoever is building, not whoever owns the data.
Why the labs are building gates at all
There is context for the gating, and it is not subtle. OpenAI stated that its Astra model is the first of its models to reach the "Critical" cybersecurity capability level under its own Preparedness Framework. In OpenAI's own definition, Critical means the model can identify and develop functional zero-day exploits in hardened real-world systems without human intervention.
In expert assessments, Astra found previously unknown vulnerabilities in a hardened browser and operating system and chained them into a working sandbox escape and a local privilege-escalation chain to root. OpenAI says it slowed development over security concerns, built safeguards, and is now limiting access to Astra's most powerful cyber capabilities while clearing the model for release.
Read those three stories together and you get the shape of September 2026. Capability is getting cheap and it is getting genuinely dangerous, and the labs' answer to both is tiering. Cheap tier with an expiry date. Capable tier with a retention agreement. Dangerous tier with a gate and a waitlist.
What I would actually do about it
Three things, and none of them are exotic.
- Stop hardcoding model names. Route by task class instead. Classification, extraction, drafting, and reasoning are four different jobs with four different price sensitivities, and each should point at a named model you can change in one place.
- Treat model choice as a data question in writing. Which tier is allowed for which class of data, and what happens if a compliance review forces a downgrade. That is a one page document and it is going to save someone a very bad quarter.
- Put a repricing date on the calendar. Not January 1, because that is too late to renegotiate anything. October.
What to watch
Whether Anthropic or OpenAI answers with a cut in the same tier inside thirty days. If they do, sub-$1 per million input becomes the default assumption for production agents and everything priced above it needs re-baselining now rather than in January.
Whether Google announces an extension of the introductory window. If nothing lands by early December, treat January 1 as real and plan the doubling.
And whether other labs copy the retention-for-capability structure. If a second major lab ships a tier that trades data retention for an unlocked model before year end, that stops being an Anthropic quirk and becomes how the industry sells capability. At which point everybody building on top of these models needs an answer ready, because the customers will start asking.
One signal a day. No noise.
A 3-minute read when something genuinely shifts in AI, automation, or defense tech. Free, most weekdays.
Free, most weekdays. No spam, unsubscribe anytime.Sources
- LLM Stats - Gemini 3.8 Flash model listing - https://llm-stats.com/models/gemini-3.8-flash
- Tech Insider - Gemini 3.8 Flash launch pricing and January 1 increase - https://tech-insider.org/gemini-3-8-flash-launch-pricing-2026/
- DataCamp - Gemini 3.8 Flash benchmarks - https://www.datacamp.com/blog/gemini-3-8-flash-cyber
- Neowin - Google launches Gemini 3.8 Flash with frontier-level performance at a fraction of the price - https://www.neowin.net/news/google-launches-gemini-38-flash-with-frontier-level-performance-at-a-fraction-of-the-price/
- Anthropic - Claude Mythos - https://www.anthropic.com/claude/mythos
- MacRumors - Anthropic Claude Fable 5.1 - https://www.macrumors.com/2026/09/01/anthropic-claude-fable-5-1/
- CNBC - OpenAI Astra cyber model - https://www.cnbc.com/2026/09/01/open-ai-astra-cyber-model.html
- SecurityWeek - OpenAI's Astra becomes first model to cross critical cybersecurity threshold - https://www.securityweek.com/openais-astra-becomes-first-model-to-cross-critical-cybersecurity-threshold/
- Axios - OpenAI Astra's cyber Critical rating - https://www.axios.com/2026/09/01/openai-astras-cyber-critical
- TechCrunch - OpenAI says it slowed Astra model development over security concerns - https://techcrunch.com/2026/08/07/openai-says-it-slowed-astra-model-development-over-security-concerns/
Quick answers
How much does Gemini 3.8 Flash cost?
It launched on September 2, 2026 at $0.75 per million input tokens and $3.75 per million output tokens. On January 1, 2027 those prices are scheduled to become $1.50 and $7.50, a doubling in both directions.
Is Gemini 3.8 Flash actually good, or just cheap?
It posted 71% on DeepSWE v1.1 and 89.4% on Terminal-bench 2.1, with a one million token context window. Google recommends it for software engineering, autonomous agents, and multi-step reasoning, so the benchmarks that matter most are the agent-relevant ones.
What is the difference between Claude Fable 5.1 and Claude Mythos 5.1?
They are the same underlying model. Fable 5.1 ships with safeguards that block or limit performance in risky areas such as cyber and bio, and also arrived with lower costs and fewer false positives on safety refusals. Mythos 5.1 is the ungated version, and using it requires accepting a 30-day data retention policy for safety monitoring by default.
What does OpenAI's "Critical" cyber rating for Astra mean?
Under OpenAI's own Preparedness Framework, Critical means the model can identify and develop functional zero-day exploits in hardened real-world systems without human intervention. In expert assessments Astra found previously unknown vulnerabilities in a hardened browser and operating system and chained them into a working sandbox escape and a privilege-escalation chain to root. OpenAI says it slowed development, added safeguards, and is limiting access to the most powerful cyber capabilities while releasing the model.