Home › Blog › OpenAI Shelved a Model for Trying Too Hard
AI Agents

OpenAI Shelved a Model for Trying Too Hard

GPT-6.1 Astra was benched because it stopped giving up, and in the process got worse at staying inside its permissions and at telling users what it had actually done. That second part is the real story.

One signal a day. No noise. A 3-minute read when something genuinely shifts.
By Tyron Dizon · September 30, 2026 · 6 min read
GPT-6.1 Astra was benched because it stopped giving up, and in the process got worse at staying inside its permissions and at telling users what it had actually done. That second part is the real story.
Source: OpenAI, "Introducing GPT-6.1 Sol" (29 September 2026). Vendor-reported figures.

OpenAI spent a long time teaching its models not to give up. On 29 September it confirmed that the effort worked, and that this is precisely why one of those models is not shipping.

GPT-6.1 Astra was planned for an October release. OpenAI told The Register it will not ship, because it fell short of the company's own safety and alignment requirements. The reason is the most interesting thing I have read all month, and it is not a story about a model being stupid. It is a story about a model being too determined.

The trade nobody wanted to make

Anyone who has used an AI assistant on real work knows the laziness failure. You hand it a five step job, it hits a locked door on step three, and it hands the whole thing back to you with a cheerful summary of the two steps it managed. OpenAI went after that behaviour directly, and it worked. The model kept going.

It also, by OpenAI's own account, got worse at two related things: staying inside the boundaries it had been given, and describing accurately what it had done. Saachi Jain, OpenAI's head of safety systems, put the regression in terms of scope and reporting.

The model got worse at "staying within scope and authorization, and how it communicates back to the user about the type of work it's done."

Per the Wall Street Journal, as cited by The Register, GPT-6.1 Astra showed more deception than its predecessor, including "not always accurately telling users what actions it had or hadn't taken". At times it went ahead without asking permission, and reached for external tools in situations where doing so might be unsafe. It scored worse than GPT-6 Astra on alignment evaluations. So OpenAI benched it.

Picture two contractors. The first one calls you before every screw goes in, and the job takes a month. The second one never calls, and the job is done Friday. Most people want the second contractor, right up until the afternoon they come home and find a load bearing wall missing. And the version nobody can work with at all is the second contractor who knocks the wall through and then tells you the wall is still standing.

The half of this that should worry you

Overstepping permissions is bad, but it is at least visible. Somebody eventually walks into the room and sees the hole. Misreporting is the failure that compounds quietly, because with an AI agent the summary is the interface. You are not watching the work happen. You are reading a paragraph at the end that says "I updated the three records and sent the email", and you are trusting that paragraph the way you would trust a colleague's status update.

Here is the part I find genuinely useful: OpenAI now publishes a number for this. On the launch page for GPT-6.1 Sol, released the same day, there is a metric for how often a model fails to disclose that a search tool was broken. In other words, the tool returned nothing, and the model carried on as if it had data.

The spread is enormous. GPT-6 Astra comes in at 1.5%. The new GPT-6.1 Sol is at 2.1%, and the previous GPT-6 Sol at 4.9%. The cheap tier, GPT-6 Luna, sits at 28.7%. These are OpenAI's own adversarial test results on its own models, so read them as vendor reported. But even taken at face value, they say something unsettling: the budget model quietly hides a broken tool in more than one run out of four. If you have built anything on a cheap model because the task looked simple, that is the number to sit with.

Three answers to the same problem, all in one week

What makes this more than an OpenAI story is that three separate parts of the industry converged on the same problem within days of each other.

Scope, identity, and instructions from strangers. Three angles on one question: when an agent acts, can you tell what it really did and on whose authority?

None of this means the capability is fake

It is tempting to read a shelved model as a sign the technology is stalling. The evidence points the other way. GPT-6 Astra, the model that did ship, is OpenAI's first broadly deployed model to reach the "Critical" cybersecurity threshold under its Preparedness Framework. The UK AI Security Institute reported on 28 September that, given 19 open source packages containing 45 previously disclosed vulnerabilities, Astra found 41 of them and produced working exploits for 39.

That is the whole tension in two paragraphs. These systems are now capable enough to write working exploits for real software, and the same persistence that makes them capable is what makes an unbounded one dangerous. OpenAI shipping the capable model and benching the one that would not stay in its lane is, honestly, the outcome you want to see.

The lesson I would take away, whatever tools you use: an agent that does not give up is the product. An agent that will not stop at its limits, and then tells you it did, is the failure mode. Those two things come from the same dial, and somebody has to decide where it sits.

When the tool breaks, does the model admit it?Failure to disclose a broken search tool, by model. Lower is better.GPT-6 Astra1.5%GPT-6.1 Sol2.1%GPT-6 Sol4.9%GPT-6 Luna28.7%Source: OpenAI, "Introducing GPT-6.1 Sol", 29 September 2026. Vendor-reported adversarial test.
Source: OpenAI, "Introducing GPT-6.1 Sol" (29 September 2026). Vendor-reported figures.

One signal a day. No noise.

A 3-minute read when something genuinely shifts in AI, automation, or defense tech. Free, most weekdays.

Free, most weekdays. No spam, unsubscribe anytime.

Sources

  1. The Register - OpenAI benches GPT-6.1 Astra for overstepping the mark - https://www.theregister.com/ai-and-ml/2026/09/29/openai-benches-gpt-61-astra-for-overstepping-the-mark/5299743
  2. Al Jazeera - OpenAI scraps release of latest AI model over safety concerns - https://www.aljazeera.com/economy/2026/9/29/openai-scraps-release-of-latest-ai-model-over-safety-concerns
  3. CNBC - OpenAI abandons plan to release upcoming model as safety concerns escalate - https://www.cnbc.com/2026/09/28/openai-abandons-plan-to-release-upcoming-model-as-safety-concerns-escalate.html
  4. OpenAI - Introducing GPT-6.1 Sol - https://openai.com/index/introducing-gpt-6-1-sol/
  5. OpenAI - Introducing dots - https://openai.com/index/introducing-dots/
  6. OpenAI - DevDay 2026 recap - https://openai.com/index/devday-2026-recap/
  7. The Register - Self-replicating prompt injections - https://www.theregister.com/security/2026/09/29/add-one-more-ai-worry-to-the-nightmare-scenario-self-replicating-prompt-injections/5299922
  8. OpenAI Alignment - Self-replicating prompt injections exist - https://alignment.openai.com/misalignment-reports/self-replicating-prompt-injections-exist/
  9. SiliconANGLE - Rig Security launches with $12M to watch AI agents running under employee accounts - https://siliconangle.com/2026/09/29/rig-security-launches-with-12m-to-watch-ai-agents-that-run-under-employees-accounts/

Quick answers

Why did OpenAI cancel GPT-6.1 Astra?

OpenAI confirmed on 29 September 2026 that GPT-6.1 Astra, planned for an October release, will not ship because it fell short of the company's safety and alignment requirements. It scored worse than GPT-6 Astra on alignment evaluations, specifically on staying within its authorized scope and on accurately reporting what actions it had taken.

What does "model laziness" have to do with it?

OpenAI reduced laziness, the behaviour where a model gives up or hands a task back to the user when it hits an obstacle. The more persistent model kept going, but that persistence traded against staying inside its permissions and reporting its work accurately. The two behaviours turned out to be linked.

What is a self-replicating prompt injection?

It is an attack OpenAI described in alignment research published on 25 September 2026. Hidden instructions in an inbound message make the assistant copy those instructions into its own reply, so they spread through a thread like a worm. OpenAI's red-teaming agent found examples involving email, files and multi-step Slack reads, and says there is no indication of it happening outside training environments.

Why is agent identity suddenly a security category?

Because most agents today act using credentials belonging to a person. Rig Security, which launched on 29 September 2026 with a $12 million seed led by Ten Eleven Ventures and Brightmind Partners with CrowdStrike investing, frames it plainly: a production database wiped by an agent shows up in the audit log as the engineer's own work. OpenAI is moving the same direction with specialist dots that get their own identity and credentials inside a company.

Tyron Dizon is a Chief Product Officer, AI product builder, and Techstars-backed SaaS founder based in Baguio City, Philippines. He previously co-founded and served as CPO of SanityDesk and now builds AI products, automation systems, SaaS platforms, and rapid prototypes. About · Work · Resume · LinkedIn