Who Investigates a Rogue AI Agent Right Now
Jamie Bykov-Brett
·
7 September 2026
·
6 min read
It’s 9pm, do you know where your AI Agents are?
You have watched this pattern before. Something goes wrong inside an organisation, the organisation announces it will conduct a thorough internal review, & weeks later a report appears that finds procedural shortcomings but no structural fault. The language is careful as it should be. And the conclusions, well they always land just short of the line where someone would need to be held to account.
Now apply that pattern to an AI agent that went rogue.
TechCrunch reported this week that OpenAI's latest agent swarm incident has renewed calls for independent investigation into AI safety failures. Researchers & lawmakers are asking a straightforward question: should the lab that built the agent also be the one deciding what went wrong?
The answer is more complicated then it sounds.
OpenAI is building its own framework for investigating these incidents. On the surface, that might come across like being responsible. A company taking safety seriously enough to formalise how it reviews failures. There’s no reason to doubt the people doing this work inside OpenAI believe in it, they need this to work more than anyone
The problem is what happens after the framework gets published.
A self-authored investigation process, once it exists on paper, becomes a template. Other labs adopt it because building one from scratch is expensive and few organisations are as AI literate. Regulators reference it because it is the most detailed thing available & they lack the technical staff to write their own. Within two or three cycles, the industry standard for investigating rogue AI agents is a process designed by the companies that produce them.
That is the accountability gap made somewhat permanent. Through the ordinary mechanics of how standards get set when one party moves first & everyone else is still writing position papers. No conspiracy required.
I have watched this play out in financial services. Banks built their own compliance frameworks long before regulators had the expertise to challenge them. Those frameworks caught some things. But they were structurally incapable of producing findings that threatened the business model of the organisations that wrote them. The questions that got asked were the questions the industry was comfortable answering, outside of that, not so much.
I have on multiple occasions looked over things that have been AI-assisted in development & found a huge amount of security flaws, compliancy missteps, accessibility missteps, security concerns, all of these kind of things that you learn through the experience of having developed these kind of products. The problem is that when you haven't done that, you don't know what you don't know & you don't know what you're looking out for. That becomes a real issue when you think you've told an AI to go & do something, but you have no way of being able to go & check that & validate that. This is where organisations have to be careful to not lose the subjective experiences of the expertise that you have in-house when you're developing something, because they're the checks & balances. They're the people that are going to help you know if what you've done is the right thing when you implement AI.
One of the things with AI is it gives a false confidence. You can see something & it works. But coming up with a prototype is the easy part. The hard part is the compliance bit. This is where we come into the concept of technical debt, where you've developed something quickly. But in doing so, the amount of time that it takes to then go & make that production ready becomes quite significant.
AI safety review is heading for the same groove. The investigations will be thorough on the technical side, producing detailed post-mortems about what the agent did & why. What they will not produce, because they cannot, is a finding that the deployment decision itself was the failure. That business pressure to ship agent capabilities outpaced the safety work. That the incentive structure producing the rogue agent is the same one reviewing it.
Both truths sit together here very uncomfortably. OpenAI investigating its own agents is genuinely better than nobody investigating them, which is what was happening eighteen months ago. It is also a process that produces verdicts before any external body has standing to contest them, creating facts on the ground that make independent review harder to justify politically. The better it works on its own terms, the harder it becomes to argue for something outside it.
The cost falls on the people least involved in the decision. When an AI agent behaves unpredictably & the only review happens inside the company that shipped it, the public is asked to trust the outcome without any mechanism to verify it. That is risk transferred to people who did not choose it, & no amount of procedural rigour inside the lab changes that structural fact.
My prediction: within eighteen months, at least two more major AI labs will publish their own incident-investigation frameworks, & those frameworks will become the de facto standard that regulators reference rather than replace. The signal to watch is whether any regulator, the EU AI Office or NIST, publishes an investigation protocol that was not substantially drafted or consulted on by an AI lab. If that has not happened by mid-2028, the window for independent oversight of agent failures will have functionally closed. I would change my mind if a regulator hires enough technical staff to do its own forensic work on model internals, but I see no budget line for that in any jurisdiction right now.
Leaders already know AI agents will behave unexpectedly, the whole point of generative ai is it is probability based non deterministic AI. Where there’s a high probability of it doing something right there’s still a probability of it doing something it shouldn’t have. The question that gets settled in the next twelve months is whether the process for understanding those failures will belong to the people who need honest answers or the people who need comfortable ones. That gets decided by who writes the template, & right now the most useful thing any leader can do is build the AI literacy to know what regulation they are managing & what it demands of them.
I know not everybody is going to be impacted by the EU AI Act, but I would recommend that everybody adheres to that level of compliance because it won't be long until that is a global regulation. Similar requirements will follow in other regions. The questions I would go to are: is this EU AI Act compliant? Is this WCAG 2.2 compliant, matching up with the EU Accessibility Act? & have we had this independently cybersecurity reviewed? AI has made it far easier to hack. The barrier to entry is lower & attacks move faster.
The other thing I would be looking at is data sovereignty. If you're actually using big providers, in the past that gave you a bit of a moat, that gave you security because they were managing those systems. But in an age of AI where hacking becomes easier than ever before, that moat can also make it a target. It can make it a honeypot because if you can break into that system, you can break into all of the systems. By being on the same system that everybody else is on & having all of your data stored in the same place as what everyone else has, all of a sudden that has become a much bigger target for people to be able to go towards.
Frequently Asked Questions
Why is it a problem for AI companies to investigate their own safety incidents?
Internal reviews cannot easily question the deployment decisions or business incentives that created the incident. The structural issue is that a company reviewing its own failure has a direct interest in the outcome. Independent investigation adds accountability that self-review, however competent, cannot replicate because the reviewer & the reviewed share the same commercial pressures. That is something always worth noting.
What happened with OpenAI's agent swarm incident in September 2026?
OpenAI experienced an incident involving AI agents behaving unpredictably, reported by TechCrunch on 4 September 2026. The event prompted renewed calls from researchers & lawmakers for independent safety investigations, with particular concern over whether AI labs should control the scope of their own safety reviews rather than submitting to external scrutiny.
How do company-written safety frameworks become industry standards?
The first organisation to publish a detailed framework typically sets the template for everyone else. Other labs adopt similar approaches because building from scratch is costly, & regulators reference existing frameworks when they lack the technical capacity to design alternatives. The earliest framework therefore shapes how the entire industry handles future incidents, often before any external body has reviewed it.
What would independent AI safety oversight actually require?
It would need a body with its own technical staff capable of forensic analysis of AI systems, funded independently of the labs it oversees. The closest model is aviation accident investigation, where national bodies like the AAIB operate separately from airlines & manufacturers. For AI, this means regulators hiring people who can interrogate model behaviour without relying on the lab's own tooling or interpretation.
What should business leaders ask their AI vendors about incident investigation?
Ask who reviews it when an agent behaves unexpectedly, what gets disclosed, & whether any external body has oversight of the process. If the answer to all three is that the vendor decides unilaterally, the organisation is absorbing risk without any say in how failures get examined. Leaders should treat the investigation framework as part of vendor due diligence, the same as data handling or uptime commitments.
Jamie Bykov-Brett
Listed as one of Engatica's World's Top 200 Business and Technology Innovators, Jamie is an AI and automation consultant who helps organisations move from curiosity to confident daily use. As founder of Bykov-Brett Enterprises and co-founder of the Executive AI Institute, he designs AI upskilling programmes that have delivered 86% daily adoption rates and a 9.7/10 NPS. His work sits at the intersection of technology implementation and human development, with a focus on responsible governance, practical tooling, and making AI accessible to every level of an organisation.
Get AI Insights Delivered
Practical perspectives on AI adoption and the future of work. No spam.
.png?width=783&height=479&name=BBE%20Logo%20Teal%20Background%20(1).png)