Getty Images/iStockphoto
AI Safety is Falling Behind the Pace of Development
Businesses need continuous monitoring, stronger guardrails and human oversight to keep increasingly capable AI systems under control.
Earlier this month, Anthropic researcher Jacob Coxon left the company over concerns about the risks posed by increasingly capable AI, warning that the technology could “kill us all by the end of the decade.”
His departure came as concerns about AI systems behaving in unexpected or harmful ways are becoming harder for the industry to ignore. Perhaps most notably seen in OpenAI’s models launching an unprecedented attack on Hugging Face earlier this summer, an incident the company said was a “warning shot” for the world about the potential power of unchecked AI.
For AI evaluation expert Brooke Hopkins, the answer lies in treating safety as an ongoing process rather than a final check before deployment. The founder of AI agent observability company Coval AI, Hopkins previously led the team at robotaxi vendor Waymo responsible for ensuring autonomous vehicles can operate safely and reliably. She is now focused on applying similar principles to voice AI, an area where she says companies are moving faster than they can identify failures, evaluate risk and understand when their systems can no longer be trusted.
In this Q&A, Hopkins discusses what the latest concerns around AI safety reveal about the industry's approach to evaluation, what voice AI can learn from autonomous vehicles, and how regulation could help establish stronger safeguards as the technology advances.
How serious do you think the risk Coxon identified is, and what do you think companies are currently getting wrong about AI safety?
Brooke Hopkins: The risk is real, but I don’t think we should treat AI as some uniquely unknown technology that’s impossible to control. What’s different is the speed and scale. AI can make existing problems like fraud, misinformation and misuse of sensitive information happen much faster and across many more interactions.
However, the speed at which we can respond and detect issues is also accelerating. Sifting through thousands of messages to detect malicious behavior is an ever-evolving task. So, the answer to these risks is also elevating safety capabilities.
The systems are changing, the models are changing and the ways people use them are changing. Safety needs to be part of the development loop, with continuous evaluation, strong telemetry and clear human accountability from the beginning.
Do you think current AI evaluation methods are good enough to identify genuinely dangerous behavior in increasingly capable models? If not, what is missing?
Hopkins: They’re getting better, but there’s still work to do. A benchmark can tell companies how a model performs on a specific test at a specific point in time. It doesn’t necessarily tell how that system will behave across millions of real interactions.
What’s missing is a much stronger focus on real-world behavior. That means continuous telemetry, adversarial testing and evaluations that look at how systems behave over longer interactions and in unexpected situations.
Human-in-the-loop for constant human alignment is also key. No eval suite catches everything, and it shouldn't have to. The teams building these systems need people actively reviewing outputs, red-teaming behavior and stepping in when something looks off, not just running a model through a test suite at launch and calling it done. Alignment should be an ongoing practice, and that means keeping humans close to how these systems actually behave in production, not just in a lab.
With voice AI, what can the industry learn from how autonomous vehicles were evaluated before being deployed in the real world?
Hopkins: Companies can’t rely on one type of test to confirm whether a system is ready for the real world. They need simulation, controlled environments, edge-case testing and real-world monitoring, followed by a gradual rollout.
Voice AI needs that same level of discipline. Companies need to understand how a system performs under normal conditions, but also what happens when conversations get messy, users behave unexpectedly or the model encounters something it wasn’t designed for. And once a system is live, evaluation can’t stop. Enterprises need to continuously monitor what’s actually happening in production and have humans in the loop when the stakes are high.
If regulation is part of the answer, what would useful AI regulation actually look like in practice?
Hopkins: The most useful regulation would focus on accountability and outcomes rather than trying to prescribe exactly how every company should build its technology. For higher-risk systems, I think it’s reasonable to expect companies to understand and document the risks, report significant incidents and be transparent about how they evaluate their systems and where those systems have limitations.
For the most capable or high-impact systems, there should also be a higher bar for independent evaluation and ongoing monitoring. The goal shouldn’t be to slow down innovation, but rather to make sure that if a company puts a powerful system into the world, there’s a clear understanding of what it can do, where it can fail and who is responsible when something goes wrong.
That’s also why industry-specific standards are important, since the risks and appropriate safeguards can look very different.
How can companies manage the tension between moving quickly to develop more capable AI and taking enough time to establish that those systems are safe?
Hopkins: I don’t think speed and safety have to be in conflict. In fact, the faster AI systems evolve, the more important it is that evaluation keeps pace.
The answer is to build evaluation into the development process rather than treating it as a gate at the end. Set clear standards before launch, test against those standards, roll out systems progressively and continuously measure what happens in production. If a system starts failing in ways that matter, companies should be able to intervene, fix it or roll it back.
Lots of tiny checks, built into every stage of development, can add up to safer and more reliable systems.
Scarlett Evans is a freelance writer with a focus on robotics and emerging technologies. Previously, she was assistant editor at IoT World Today, where she specialized in robotics and smart city technologies. Scarlett also has a background in the mining and resources sector, with experience at Mine Australia, Mine Technology and Power Technology.
Editor’s note: This interview was edited for clarity and conciseness.