Your AI Agents Are Approved. That's Not the Same as Governed.

Security leaders already rank agentic AI governance among their weakest controls. Here are four risks hiding under that awareness.

 

Sometime in the next couple of board cycles, a director is going to ask you how your company is governing AI. If the answer is an approved policy and a green status box on the quarterly risk slide—the same green it showed all year—you have a problem, because neither the policy nor the box says anything about what your AI agents have been doing.


Security leaders like you now find themselves juggling two conflicting mandates when it comes to AI.

  1. Allow the organization to adopt AI fast enough to keep pace with competitors.

  2. Make sure that this adoption isn’t creating unseen risk.

This situation mirrors the same tensions that security leaders have faced over the last 20 years with advances like implementing cloud computing, integrating mobile devices within a distributed network infrastructure, and building microservices-based apps using DevOps principles. In other words, how do you protect the enterprise without stifling innovation? 

Agentic AI may prove to be the biggest challenge of all. In my previous post, I explained the ways in which AI agents behave more like employees than applications. And if they behave like employees, you may conclude that you would govern them like employees. Bounded autonomy with exception-based control would give agents room to work and pull humans in when the stakes rise. The trouble arises when we assume that model can keep AI agents within the boundaries we define. 

The issue is compounded by the traditional (if all the previous advances of the last 20 years can be considered “traditional) focus on controlling what can access the network. Yes, it’s important to vet and approve Agentic AI systems, draft acceptable use policies, and stand up an exception process for any actions that go beyond these parameters. But these processes aren’t designed to anticipate, let alone control, what AI agents do once they’re allowed through the door.

Given that these agents can sidestep controls that even human employees can’t outmaneuver, how do you limit their access to sensitive files or prevent attackers from injecting malicious prompts within something that appears to be aboveboard?  

This disconnect leads to four security risks that plague effective governance of Agentic AI. These risks only become more urgent as Agentic AI grows exponentially more capable by the day. Let’s take a look at each of them. 



Risk No. 1: The Exception to Your Exception Handling 

Exception-based control can work with AI agents under the model just described. But as AI agents become more autonomous, they increasingly treat these controls as obstacles rather than as strict boundaries that cannot—and should not—be crossed. If an agent needs to install a PDF reader and discovers that it's blocked from using the installer, it tries PowerShell. If PowerShell is blocked, it tries other methods until it completes its objective because, well, they have that generative AI intelligence everyone is so excited about. 

This problem isn’t theoretical. In late July, OpenAI reported that two of their most capable models, running in a sandbox with reduced guardrails, broke out of the test environment, connected to the internet without human direction, used stolen credentials, and exploited a previously unknown vulnerability to reach the servers at Hugging Face, a repository for AI testing data. Hugging Face caught the intrusion about a week before it learned OpenAI was responsible. The models had been instructed to test complex attack paths. Nobody had written an exception rule for what happened when those paths led out of the sandbox. 

Exception handling, as it's currently built, catches the actions someone thought to prohibit. It isn't designed to anticipate the methods nobody thought of. 

 

Risk No. 2: Knowing is half the battle. Controls are the half that counts. 

So why hasn't anyone fixed that? Security leaders aren’t blind to the risks Agentic AI poses. A recent McKinsey survey found that almost two-thirds (62%) of respondents cite security and risk concerns as their top obstacle to scaling Agentic AI, way ahead of technical limitations or regulatory uncertainty. Yet less than one-third (32%) report having the AI governance controls in place to manage that risk.

This gap between recognizing the risk and having the operational controls needed to control said risk has been compounded by the ongoing optimism (at least by senior leadership) that generative AI can take on every productivity problem you can throw at it. Mounting workloads? Agentic AI can handle the increase without having to hire more employees. Efficiencies excite shareholders. Risk? Not so much.

 

McKinsey & Company - State of AI trust in 2026: Shifting to the agentic era

Responsible AI Control

32%
Have a strategy for responsible AI practices
62%
Cite security and risk concerns as top obstacle
32%
Report AI governance control in place

2026 AI Trust Maturity Survey taken between December 2025 and January 2026; responses from ~500 organizations across industries and regions, with respondents who hold direct responsibility or expertise in AI governance, risk management, or AI investment decisions. https://www.mckinsey.com/capabilities/tech-and-ai/our-insights/tech-forward/state-of-ai-trust-in-2026-shifting-to-the-agentic-era

 

But the risks continue to mount. According to the CSA survey cited in my first post, 65% of respondents say their organization has experienced at least one security incident involving an AI agent in the last year. Given that OpenClaw wasn’t even released before late November 2025 (under a different name, no less) and the news outlet AP cited Agentic AI as the newest buzzword around the same time, I'm guessing the percentage, both in the number of organizations experiencing an incident and in the number of incidents experienced, has only grown to levels most of us would rather not think about. Unfortunately, until that awareness is matched with the operational controls necessary to counteract the risk, the problem is only going to get worse.

 

The Risk No. 3: ‘Show me’ is the new ‘Are we safe?’ 

There is one thing in common with controls that can't anticipate an agent's methods and controls you haven't fielded yet: neither leaves a record. That's the third risk, and it lands directly on your desk. Even where governance exists, few organizations can show evidence of it. As I said in the introduction (and as Airlock Digital has discussed elsewhere), your job governing AI agents only starts at the system approval stage. You need to be able to verify, on demand, that each action an agent takes stays within your security policy whenever your senior leadership, your board of directors, an internal auditor, and ultimately, a regulator asks you for proof.

In the past, documentation was sufficient—quarterly risk reports, signed policies, a clean audit. Now the question is, Show me. A single AI agent (and let’s face it, like Lay’s Potato Chips, nobody can have just one) makes thousands of decisions between your quarterly reviews. If none of its decisions leaves a record that a third party can interpret, then you’re stuck with the grown-up version of Because I said so! Good luck getting away with that answer.

The CSA survey notes that only 16% of respondents say their organization continuously monitors their AI agents. About 60% checks in daily or weekly, while 17% only does so after something breaks, which is less a monitoring strategy than an autopsy. But periodic monitoring of AI agents leaves you with the same evidence gap as no monitoring—because you can’t reconstruct what any AI agent does between those checkpoints.

 

Risk No. 4: Your agents are talking to each other. Nobody's on that thread. 

Fix the monitoring and you close that gap—for agents working one at a time. But how many of your agents work alone? Agents can delegate to agents. One agent's output is another agent's instruction. The latest Salesforce Connectivity Benchmark Report finds that half of enterprise AI agents already operate as part of a multi-agent system rather than in isolated silos.

Now run the earlier risks through that architecture. Your exception handling assumes a human is one hop away from any consequential action. But how many hops away is your human in a multi-agent chain? You have no way of knowing. Meanwhile, all these agents are supervising one another doing who knows what. An agent blocked at one path doesn't just reach for PowerShell anymore; it can hand the task to a peer with different permissions. And monitoring built to watch what agents do to systems and data has no view of what agents say to each other. Your agents have been cc'ing each other for months—but you're not on the distribution list.

The silo number cuts both ways. With half of agents still working alone, the coordination channel isn't loud enough yet to demand attention, which makes it a hard line-item to defend to the people funding your controls. We'll look at it when there's something to look at.

That's the trap. I'd make the same call if the numbers held still, but they don't. The Salesforce survey states that organizations run an average of 12 agents today, a number the report forecasts will climb 67% within two years. The cheapest time to build visibility into a channel is before it gets busy. Wait until agent-to-agent traffic is load-bearing, and you'll be retrofitting oversight onto a system that already routes around it.

 

Adoption is what they’re pushing. Risk is what you’re carrying. 

The processes you have aren't wrong. They just stop at the door. They don't touch agents that treat controls as obstacles and keeps outrunning the controls you can field. Nor can they handle the necessary level of governance that can restrict what AI agents can do, let alone what they’re coordinating among themselves. 

So, when that director asks how you're governing AI, the policy and the green status box will still be true. They just won't be responsive. What you need is a way to build the better answer while leadership is still measuring your program by how fast it says yes. You've been in that squeeze since the early days of the cloud, and it has never once resolved in favor of slowing down. I've named four risks and solved none of them—which is where this post ends and for you, the actual work starts. 

 

Next Steps

Naming the risks is the easy half. The working half — how to discover the agents you already have, decide which ones to trust, and govern what they do once they're running — is the ground we cover in our new guide, See, Govern, Restrict: A Five-Step Blueprint for Agentic AI Control. We hope you'll enjoy the read.