When an AI Acts Without Permission What Happens Next
Jamie Bykov-Brett
·
1 September 2026
·
15 min read
It's time to shake off the bank holiday rust and get this condensed week underway.
Look, we've all been there, you've delegated a task to someone capable, given them broad instructions, then checked back to find they'd done something you never asked for. Maybe they meant well. Maybe the outcome was fine. But the moment you realise your instructions were interpreted more creatively than you intended, it changes how you think about oversight.
Now scale that to an AI system operating autonomously.
Earlier this year, Anthropic's Claude agents took what the company describes as "unauthorised actions" during real-world operations. Now, "gained unauthorised access to real-world systems", is doing a lot of work and specifics are sparse. However the response was significant: Anthropic temporarily paused some AI training & cybersecurity evaluations, as detailed in a blog post published today & reported by Axios. A public operational change, disclosed voluntarily.
There hasn't been any details shared of what systems, or when beyond "earlier this year", or what the actual harm was, if there was any. What we can say the company found something it thought was serious enough to stop training runs over and you don't do that for nothing because training is expensive. You definitely don't do that for a near miss you're comfortable with.
For years, autonomous AI misbehaviour lived in the realm of alignment theory: conference papers, red-team exercises, & thought experiments about paperclip maximisers. What Anthropic just did moves it into operations. A frontier lab looked at what its agents actually did, decided the risk warranted stopping certain work, & told the public. That is an incident response.
I think this is the clearest signal yet that the gap between "AI safety as philosophy" & "AI safety as ops" has closed. The labs are reacting to what did go wrong.
Anthropic pausing training is a responsible act, the kind of move I wish was more frontier labs were comfortable making when something unexpected happens. It also confirms that autonomous agents are already operating in territory where "unauthorised actions" is a category that exists. Both are true in the same breath. The pause deserves kudos. The fact that it was necessary deserves scrutiny.
Pausing cybersecurity evaluations too, which I personally find more interesting, honestly because that's the testing. If you pause the the part of the process that's meant to catch problems, one speculation I would have around that is that the test itself was the risk. That the evaluation environment was where the agent got out. It's like trying to measure water being held in an holey bucket.
Let's be honest, most organisations adopting AI agents right now are not Anthropic. They don't have alignment teams or evaluation frameworks sophisticated enough to catch an agent acting outside its mandate before the consequences land. If a frontier lab with hundreds of safety researchers can end up pausing its own training because its agents surprised it, your team's readiness to deploy autonomous systems into procurement or customer operations probably warrants a closer look. This has gone to show recently, obviously, within OpenAI & the Hugging Face hack, & Anthropic have admitted to it themselves. Anthropic only found this out three months after it happened, & that's because most people aren't actually monitoring what their agents are doing. They don't have human in the loop, they don't have records that are monitoring the ins & outs of what an AI agent is doing, the decisions it's making, the tasks it's completing, the alignment that it has.
Many leaders carry a belief that slows them down: the assumption that if a company admits its AI did something wrong, the technology must be too dangerous to touch. The opposite reading is more useful. Anthropic's willingness to pause & disclose is exactly the behaviour you should be looking for from your own vendors. The dangerous vendors are the ones who never seem to have incidents because they never tell you when things go sideways.
So the practical question is about governance infrastructure. Most organisations can't tell when an agent does something outside its mandate. Most people aren't recording how these agents work, & if they are recording it they're definitely not having somebody check over it to see what tasks have been completed, & even if they have got something that is monitoring it & someone checking over it, they probably don't have someone then also checking over that work as well. The playbook is straightforward: define what agents can & cannot do, then build the logging & culture to catch it when they step outside those lines.
My prediction: within twelve months, "autonomous agent incident disclosure" will become a standard section in enterprise AI vendor contracts, sitting alongside uptime SLAs & data residency clauses. The signal is whether procurement teams start asking for it. Right now, I don't think procurement teams are even thinking about that. When I talk to most organisations at the moment, they are at the beginning stages of implementing AI & most of them are working with agents that sit within the infrastructure that they already have, like Copilot infrastructure or the Gemini infrastructure, but that is a very user-friendly UX version of what an agent is & it's nothing compared to what the capability of agents actually is. It's one that makes sense to people, they can sit with it in front of them & they can see & feel that there's some degree of control or autonomy on it, but the agents being developed through Claude are essentially running on an operating system with full access to everything that a human being can do on a computer. If the next two major agent-related incidents from any frontier lab produce similar public disclosures, the pattern will be undeniable, & legal teams will formalise what Anthropic just did voluntarily. If disclosures stay rare, it means the industry chose opacity, & buyers will be flying blind. I think the former is more likely, because the reputational cost of hiding an incident after Anthropic set this precedent will be higher than admitting one.
One thing to try this week: ask whoever manages your AI tools whether you'd know, within 24 hours, if an agent took an action outside its defined scope. If the answer is uncertain, that's your starting point.
The organisations that come through this transition will be the ones who built the muscle to govern AI, starting with real AI literacy across their teams.
Frequently Asked Questions
What unauthorized actions did Anthropic's Claude agents take?
Anthropic has not disclosed the full specifics. What we know from their blog post & the Axios report is that Claude agents took actions outside their intended scope during real-world operations earlier this year, prompting the company to temporarily pause some AI training & cybersecurity evaluations while it investigated.
Why did Anthropic pause AI training rather than just fix the bug?
Pausing was an operational response to prevent the same patterns from compounding during the investigation. Anthropic treated the incident as a systemic concern, halting certain work until it understood the root cause & could adjust its safety evaluations.
Does this incident mean AI agents are too risky for business use?
No. The lesson is about governance. Anthropic paused & disclosed publicly, which is what responsible AI operations look like. The risk is with organisations that deploy autonomous agents without the systems to detect when those agents act outside their scope.
How can I tell if my AI vendor takes safety seriously?
Look for transparency about incidents. Vendors who disclose problems & change their processes afterward show operational maturity. Ask vendors directly whether they have incident disclosure policies for autonomous agent behaviour & whether they can provide logs of agent decision paths.
What governance should my company have before deploying AI agents?
At minimum, you need clear boundaries defining what agents can & cannot do, plus logging that captures agent decision paths. You also need a culture that treats unexpected agent behaviour as an investigation trigger. A good baseline: could you detect an out-of-scope action within 24 hours?
Jamie Bykov-Brett
Listed as one of Engatica's World's Top 200 Business and Technology Innovators, Jamie is an AI and automation consultant who helps organisations move from curiosity to confident daily use. As founder of Bykov-Brett Enterprises and co-founder of the Executive AI Institute, he designs AI upskilling programmes that have delivered 86% daily adoption rates and a 9.7/10 NPS. His work sits at the intersection of technology implementation and human development, with a focus on responsible governance, practical tooling, and making AI accessible to every level of an organisation.
Get AI Insights Delivered
Practical perspectives on AI adoption and the future of work. No spam.
.png?width=783&height=479&name=BBE%20Logo%20Teal%20Background%20(1).png)