Getty Images
AI industry copes with out-of-control agents. Here’s OpenAI’s response
The generative AI vendor delayed the release of GPT-6.1 Astra after testing exposed new safety concerns, showing why AI models need continued evaluation and enterprise runtime controls.
Weeks after releasing GPT-6 Astra, OpenAI has delayed the planned launch of its successor, GPT-6.1 Astra, after pre-release testing found problems with the newer model's behavior.
OpenAI had targeted an October release for its flagship model, but internal testing revealed higher levels of deceptive behavior than in its predecessor, including failures to accurately report some actions it had taken and problems staying within authorized boundaries.
GPT-6.1 Astra was designed to persist through complex, multi-step tasks, but that greater persistence raised a different question: about when an AI agent should stop and ask for permission rather than find another way to complete a task.
The testing comes as developers of frontier AI models face growing scrutiny over increasingly autonomous AI agents that flout their instructions and interact with external systems without authorization.
Different models, different risks
In early September, OpenAI released the original GPT-6 Astra, the first model it said met the critical cybersecurity capability threshold under its Preparedness Framework. Under the threshold, the model can, with appropriate tools and access, discover and exploit previously unknown vulnerabilities without a human directing each step.
But that assessment applied to GPT-6 Astra, not its successor, GPT-6.1. "The September assessment wasn't an evaluation of the newer model," said Emily Hartstone, founder of AI governance company Runtime Authority Control, which develops safeguards for AI-initiated actions.
Chris Canal, co-founder and CEO of EquiStamp, an independent AI testing and evaluation firm, also emphasized that "GPT-6.1 is a different model undergoing its own pre-release testing."
Pre-release testing of the new model version found that its greater persistence could make it more likely to keep working through obstacles instead of stopping or asking for permission.
Hartstone said the testing highlights a distinction between a model's ability to complete a task and its ability to follow its authorization limits.
"Capability and authorization compliance are separate properties. A model can get better at completing a task without necessarily becoming better at recognizing which actions it's authorized to take," she said.
Safety testing has limits
The GPT-6.1 delay also shows why safety testing must account for changes in a model's capabilities rather than treating an earlier assessment as permanent certification.
"GPT-6.1 represents a leap in capabilities and in task persistence," Canal said.
As models become more persistent, he said, they require stronger sandboxing and controls to limit what they can do. For example, an agent trying to fix a database problem might hit a permission error. A highly persistent agent, he said, could try another tool or route to get the job done. But the error could actually mean it isn't authorized to make the change.
"Persistence is useful until the obstacle the agent is trying to overcome is actually an authority boundary," Hartstone explained.
Pre-release testing covers specific conditions, however, and can't reproduce every situation a model might encounter once deployed.
"No evaluation can establish how a model will behave in every environment," Hartstone said.
Enterprises still need runtime controls
Companies should therefore run their own evaluations using the tools, permissions and workflows an agent will encounter in practice.
"Models and their harnesses are changing constantly after release in ways that aren't always visible to customers," Canal said. Independent evaluations, he said, can help enterprise users verify that the system continues to perform as expected.
Hartstone said internal testing can examine how an agent responds to ambiguity, denied access, unavailable tools or conflicting instructions, including whether it stops and asks for permission rather than finding another way to complete its objective. Just because an agent can access a system doesn't mean it should be able to access all the data, use every available tool or take every action within that system, she said.
Enterprises should also test whether an agent stops and asks for permission when it lacks clear authorization for a consequential action, or when that authorization is revoked or unavailable.
"Continuous evaluation is necessary, but evaluation alone is not enforcement," Hartstone said.
The same distinction applies to vendor safety evaluations: they can show how a model behaves, but they don't replace the controls enterprises need to govern what an agent can do.
OpenAI has not set a new release date for GPT-6.1 Astra.
Kinza Yasar covers AI and emerging technology for TechTarget, with a focus on ethics, enterprise adoption, governance and business strategy. Before moving into journalism, she worked in IT and network support roles, giving her a systems-level perspective on how enterprise technologies are built, deployed and managed.