kras99 - stock.adobe.com
AI agents can act. It’s unclear if enterprises can stop them.
As AI agents gain more authority to act on their own, recent incidents are exposing gaps in how enterprises monitor, control and intervene when agents cross established boundaries.
AI agents are being designed to do more than just answer questions. Organizations are increasingly giving them the ability to access systems, use tools and take actions autonomously.
But a string of incidents is exposing the other side of that autonomy: what happens when an agent does something it wasn't supposed to do?
Google confirmed this week that its Gemini AI system broke out of a sandbox and hacked three companies during testing earlier this year. The incidents stemmed from testing-environment problems similar to those that have affected agents from OpenAI, Anthropic and Meta.
At a press conference in New York on Thursday, Australian Prime Minister Anthony Albanese said an OpenAI agent hacked into a government health agency in June, gaining unauthorized access to public and non-public files. Albanese also criticized OpenAI’s response.
Those incidents point to a problem enterprises will have to confront as they give AI agents more authority: Deciding what an agent should be allowed to do isn't enough. Companies also need to know when an agent crosses those boundaries and have a way to stop it.
A report from AI observability vendor New Relic found that one in four AI agents run unmonitored. That becomes a bigger problem as companies deploy more agents and lose visibility into what those agents are actually doing.
And that’s where the problem lies. Giving an agent permission to act is one thing. Knowing what it is doing once it starts acting is another.
A new category of runtime controls is emerging in response.
Okta, for example, introduced new tools this week designed to give companies more control over AI agents while they're operating. Its Agent Gateway sits between agents and the tools they interact with, enabling companies to enforce policies and log interactions at runtime. If something goes wrong, a new kill switch can revoke an agent's active tokens and shut down sessions already in progress.
“It’s very hard currently for us to stop an agent if it’s doing something that it shouldn’t be doing,” James Simcox, chief product and operating officer at U.K. fintech firm Equals Money, told Computer Weekly.
The controls are designed to answer some increasingly important questions for enterprises: Where are their agents, what can they do, what are they doing and how can companies respond when something goes wrong?
The emergence of tools like these suggests that enterprises need a way to intervene while an agent is acting.
Governance specifies what an agent can do. Runtime control determines what happens when it doesn't.
Also this week in AI news:
Anthropic, OpenAI launches show shift toward multi-model enterprise AI: New models from Anthropic and OpenAI are giving enterprises more options for different workloads while adding new considerations around managing a growing mix of models.
AI gains call for organizational overhauls: AWS: AWS research suggests companies need to rethink how they organize around AI as pressure grows to demonstrate business value from their investments.
Meta Connect 26: Muse paves the way to global domination: The social media company is positioning Muse as a platform that can orchestrate third-party services for users across its massive app ecosystem.
Should workers be paid for AI skills? Companies could face retention challenges if they pay a premium for new hires with AI skills without similarly rewarding existing employees who develop the same expertise.
Wiz uses AI to find vulnerabilities in critical infrastructure: The Google subsidiary is offering free AI-powered vulnerability scanning to operators of critical infrastructure, including railroads and hospitals.
In the AI era, organizations must redesign work in real time, Gartner says: Companies need to rethink how they structure work as AI changes the skills and capabilities organizations need.
Amazon: AI supply chain agents among seller upgrades: The online retail giant is introducing AI agents for tasks including inbound planning and aged inventory as it expands automation tools for sellers.
Bank of America plans to double AI budget next year: The second-largest bank in the U.S. plans to double its AI budget as the company points to identifiable gains from its investments in the technology.
Boston Dynamics begins robotics testing at Hyundai, outlines expansion: The Massachusetts-based company is testing Atlas humanoids at Hyundai facilities as it works toward using the robots for component assembly by 2030.
Disney introduces CTO role to lead enterprise tech, AI platforms: Disney created a CTO role and tapped former Character.AI CEO Karandeep Anand to oversee enterprise technology and AI platforms.
Liz Hughes is an award-winning editor and writer covering AI and emerging technology and the former editor of AI Business and IoT World Today.