The Model Companies Won't Save You

AI Security
Replace or connect this featured image
Replace or connect this featured image
Summary:

Why agentic AI security is a problem the frontier labs have already told us they can't solve, and what enterprises should do instead.

When a sophisticated orchestration of AI agents possibly directed by a state actor was used to attack a large number of websites earlier this year, the response from the model company involved was instructive. The underlying tooling was open source, they said. Their agents were merely the orchestration layer. This was a problem for the security industry to solve.

The reality is that they were right. In the fast moving and at times uncomfortable world of agentic AI, this was an accurate description of where responsibility for agentic security actually sits. One the industry has been strangely reluctant to accept, including, at times, the model companies themselves.

The reality

Having stepped back from securing agentic attacks, the same frontier AI labs haven't stepped back from attempting to secure agents. Reports from across the industry describe an emerging arrangement that amounts to ‘send us the prompt, the chain of thought, the output, the logs and we'll deal with it.’

This is not a problem which is going to be solved by the LLM companies. In fact, the biggest risk isn't that the model companies fail. It's that enterprises believe they're capable of succeeding and build their agentic deployments on that assumption.

Current pathways to security remain flawed

This isn't a criticism of the research. Some of it is genuinely impressive, for example, Anthropic's interpretability work on attribution graphs and its efforts to identify and suppress specific concepts inside models. But even pushed to its theoretical extreme, interpretability fails as a security system, for a simple reason: in the context of something as mundane as booking a train ticket or making a payment, there is a whole universe of dangerous concepts, and it is practically impossible to enumerate them all.

Chain-of-thought monitoring fares no better. The model companies' own research shows that the reasoning a model displays is not necessarily reflective of what it is actually doing. Some describe this as models "lying" but that anthropomorphises a stochastic system beyond what's reasonable. Lying requires intent. When a probabilistic text generator outputs a wrong number, did it deceive you, or were you unlucky? You cannot build enterprise security on a concept that is fuzzy.

Every one of these approaches shares the same architectural limitation: they operate upstream of execution. Guardrails influence what an agent is likely to do. They cannot guarantee what it will do the moment it invokes a tool. Probabilistic controls applied to a non-deterministic system will reduce the rate of harmful outcomes. They will never guarantee safe actions.

By any of the metrics we've used to consider systems secure before, any of the behaviours we've required of humans or of networks, none of this would pass muster.

Authenticated, authorized, and still dangerous

If you want a picture of where this leaves enterprises, consider the "SearchLeak" vulnerability disclosed in Microsoft 365 Copilot Enterprise this year. A single link could cause a Copilot agent to retrieve sensitive enterprise information and transmit it externally using its own legitimate capabilities.

The significance wasn't that the agent lacked access or authorization. It had both. It was authenticated, operating through approved workflows, doing things it was fully entitled to do. Traditional safety layers were present and functioning as designed.

That is the enterprise reality the "leave it to the labs" assumption collides with: an agent can pass every check and still execute actions that create serious risk for you. The gap isn't in the model's behaviour. It's in the absence of any independent mechanism that decides, at the moment of execution, whether a specific action, by a specific agent, using a specific tool, under a specific policy, should be allowed to happen at all.

The right approach

None of this means the model companies should stop. Their research into chain of thought, harmful content and bias genuinely matters, and enterprises benefit from every improvement. But, as an industry we need to accept that these approaches will not deliver the deterministic control and security we desire.

Security has been here before. Operating system vendors improved, and endpoint security still became an industry. Cloud providers hardened their platforms, and cloud security still became a category because the platform provider's responsibilities and the customer's risks are not the same thing, and never will be. Every generational shift in computing has demanded a security layer built independently of the platform being secured. Agentic AI is no different.

What that layer looks like follows directly from the problem. If risk is defined by the customer, not the model provider, then policy must be authored by the customer. If agent reasoning cannot be read or trusted, then enforcement must operate independently of it. And if guardrails are probabilistic, the control that sits at the point of execution must be deterministic: the action either meets the policy, or it cannot happen.

That is the layer Outerlimit has built. In today’s agentic era, Outerlimit provides the world’s first decentralized security and authorization layer designed specifically for securing Agentic AI.

Recognizing a fundamental shift in enterprise risk, Outerlimit is extending Zero Trust to the agent action layer, binding identity, authorization, and action into a single operation at the moment of tool execution. Meeting Enterprise customers where they are today, Outerlimit guides organizations from discovery and observability to deterministic enforcement.

The model companies won't save you. They've told us as much.The good news is they don't need to, securing what agents do was never their job. It's ours, working with you.