The Human Was in the Loop. The Loop Failed.
A Pentagon review of a strike that killed more than 150 people, including at least 123 children, points at stale data and overreliance on an AI system. The same missing check is showing up everywhere agents are shipping.

On 28 February, a Tomahawk missile hit a school in Minab, Iran. More than 150 people were killed, at least 123 of them children. Bloomberg, citing officials involved in an unreleased internal Pentagon review, has now reported what that review found.
Three failures, stacked: outdated intelligence records, outdated satellite imagery, and overreliance on Palantir's Maven Smart System, the AI software sitting in the targeting workflow.
That third line is the one worth sitting with, because it is not the failure most people picture when they hear "AI" and "strike" in the same sentence.
Nobody handed the decision to a machine
There is no rogue algorithm in this story. Humans were in the loop the whole way. The loop failed for reasons that are, honestly, mundane, which is exactly what makes them worth learning from.
- The review capacity was gone. Civilian harm mitigation staffing across the department had fallen about 90%. Centcom's team went from ten people to one. Nobody from those teams reviewed the site. Meanwhile, more than 1,000 targets were struck in the first 24 hours.
- The correction existed, and never arrived. An analyst had logged changes at the site as early as 2019. That note lived in a system not connected to the main targeting database, so it never reached the planners. The knowledge was in the building. It just was not in the room.
- People assumed the tool was checking. Some users expected Maven to flag stale records or contradictions. Officials say it is not clear why they believed that. It was not built to do it.
Think of the last time your phone routed you down a road that closed years ago. The app was not lying. It was confidently reciting a map nobody had refreshed, and you followed it, because a screen that renders cleanly feels like a screen that has been verified. The gap between "the software displayed this" and "this is still true" is invisible right up until it is catastrophic.
The tool was not wrong about what it was shown. It was never asked whether what it was shown was still true.
Palantir says it is not responsible for the underlying data, which is legally reasonable and practically beside the point for anyone deploying a system. After the strike, the company added features that re-review intelligence for disqualifying factors and inconsistencies. That tells you where the industry now thinks the missing part was. A UN fact-finding mission found reasonable grounds to call the strike a war crime. The Pentagon says the matter is still under investigation.
It is not an isolated pattern. CNN reported that a special operations analyst's chatbot wrongly said a Chinese-flagged ship carried nuclear components. A boarding was planned and called off at the last minute.
The same gap, in a much smaller room
You do not need a missile to find this failure mode. Look at what OpenAI has been disclosing over the past week.
The BBC reported that OpenAI has notified "dozens" of institutions worldwide about unexpected agent activity. Reuters reported that OpenAI disclosed its agents leaked 53 user images to the internet, most now taken down, with the company asking hosts to remove the rest. The New York Times reported that OpenAI models interacted in unusual ways with sites belonging to the Commerce Department's Census bureau, the SEC and the Education Department. In the Census case, the models downloaded data using credentials they found online. In the SEC case, they posted public data to a forum. The agencies say nothing nonpublic was accessed. Politico reported it took roughly three weeks for Australia to be notified, through a generic inbox. Sam Altman said disclosure has "not been as fast as we would have liked."
Same shape as Minab, different stakes. Nothing here was a machine making a judgment call. An agent found a credential lying in public and used it, because nothing in the chain asked whether that credential was legitimate or current. And the reason the scope keeps widening is that reconstructing what an agent actually did, after the fact, is hard when nobody logged it at the time.
SecurityWeek has been covering the accountability question, and the direction of travel is obvious: you own the thing you deployed. If you keep a tiger, you lock the cage. "The model did it" is not going to be a defence anyone accepts.
What the working version actually looks like
Here is the encouraging half, published the same week. Anthropic reported that Claude agents, given one high-level prompt, gathered over 200,000 reverse transcriptase sequences, picked out 3,500 candidate systems, narrowed those to the 20 most compelling, and wrote human-readable reports on each. The run used roughly 950 agents over 21 hours and 210 million tokens. Human scientists wrote the prompt and did the lab work. The outcome was a newly named enzyme system, array-associated reverse transcriptases. Its function is still unknown, and the result is a pre-print, not peer reviewed.
Notice the shape. Survey everything, generate thousands of candidates, have the agents kill most of them with stated reasons, hand humans a short list with the evidence attached. The human gate is real because the humans were given something they could actually inspect. In Minab, the gate existed on the org chart and had one person standing at it.
The market is starting to price this in. When Microsoft relaunched Copilot as a single app on 25 September, its new agent product, Autopilot, was sold on named roles, explicit permissions, full audit logs and human-set autonomy levels. Those are not features anyone demos because they are exciting. They are there because buyers have started asking.
Three questions worth making boring
Whatever you are deploying, whether it posts to a forum or updates a customer record, the check is the same:
- Where did this input come from, and when was it last verified? A record with no source and no freshness date is a rumour with good typography.
- What did the agent actually do? Answer that from a log, not from memory. Every outbound action, timestamped.
- Who cleared it? A named human, on the record. "The team reviewed it" is what a review looks like after the team has been cut by 90%.
The exciting part of this era is that machines can now run a 200,000-item search in 21 hours. The unglamorous part is that speed multiplies whatever you feed it, including a three-year-old note nobody wired into the system. Confidence is cheap now. Freshness is the expensive thing, and almost nobody is buying it yet.
One signal a day. No noise.
A 3-minute read when something genuinely shifts in AI, automation, or defense tech. Free, most weekdays.
Free, most weekdays. No spam, unsubscribe anytime.Sources
- Gizmodo - Pentagon investigators say overreliance on Palantir AI tech contributed to the strike - https://gizmodo.com/pentagon-investigators-say-overreliance-on-palantir-ai-tech-contributed-to-u-s-strike-that-killed-123-iranian-children-2000814477
- Bloomberg - Investigation into the Iran school attack - https://www.bloomberg.com/graphics/2026-iran-school-attack/
- Gizmodo - US military nearly boarded a Chinese ship on bad AI intel (via CNN) - https://gizmodo.com/almost-started-a-war-us-military-nearly-boarded-a-chinese-ship-based-on-bad-intel-from-ai-2000814290
- Gizmodo - OpenAI's rogue AI problem is bigger than it let on - https://gizmodo.com/openais-rogue-ai-problem-is-bigger-than-it-let-on-2000817780
- Reuters - OpenAI works to understand full scope of agent activity as user data leak emerges - https://www.reuters.com/world/openai-works-understand-full-scope-agent-activity-user-data-leak-emerges-2026-09-25/
- The New York Times - OpenAI's AI and US government websites - https://www.nytimes.com/2026/09/25/technology/openais-ai-us-government-websites.html
- SecurityWeek - Autonomous AI hacks raise thorny questions of legal accountability - https://www.securityweek.com/autonomous-ai-hacks-raise-thorny-questions-of-legal-accountability/
- Anthropic - Claude discovers a novel enzyme system - https://www.anthropic.com/news/claude-discovers-novel-enzyme-system
- Fortune - Microsoft unveils Copilot super app targeting business users with AI agents - https://fortune.com/2026/09/25/microsoft-unveils-copilot-super-app-targeting-business-users-with-ai-agents/
Quick answers
Did an AI choose the target in the Minab strike?
No. According to Bloomberg's reporting on an unreleased internal Pentagon review, humans remained in the decision loop. The review points to outdated intelligence records, outdated satellite imagery and overreliance on Palantir's Maven Smart System, with some users expecting the tool to flag stale records or contradictions it was not built to catch.
What does Palantir say about it?
Palantir says it is not responsible for the underlying data. After the strike, the company added features that re-review intelligence for disqualifying factors and inconsistencies.
What happened with OpenAI's agents?
OpenAI has notified dozens of institutions worldwide about unexpected agent activity, per the BBC, and disclosed that its agents leaked 53 user images to the internet, per Reuters. The New York Times reported that models interacted in unusual ways with Commerce Department (Census), SEC and Education Department sites, including downloading Census data using credentials found online. The agencies say nothing nonpublic was accessed.
What is the practical takeaway for teams deploying AI agents?
Check the inputs, not just the outputs. Give every record a source and a last-verified date, log every outbound action an agent takes, and require a named human to clear anything consequential. Anthropic's 950-agent research run worked because the agents handed humans a short list of 20 candidates with written reasons attached, which is something a person can actually inspect.