Alex - stock.adobe.com

AI model incidents highlight security concerns

In this Q&A, Seth Johnson, CTO at Cyara, discusses how enterprises need to sharpen their security policies and infrastructure in the wake of recent AI frontier model incidents.

In recent months, the tech world has seen security incidents involving the industry's most advanced frontier models. Major AI companies -- including OpenAI, Anthropic and Meta -- disclosed cases where their AI systems exhibited unexpected behavior by accessing external systems and exploiting vulnerabilities beyond their intended operational boundaries.

One question that persists is whether the frontier models were "going rogue" or were simply following their instructions and doing what was intended.

Regardless, these incidents have sparked important conversations about AI safety, security protocols and the responsibilities of both AI developers and enterprise users in maintaining robust guardrails around increasingly powerful language models.

Informa TechTarget recently spoke with Seth Johnson, CTO at Cyara about how CIOs need to think about the evolving nature of the AI frontier models and the implications for enterprise cybersecurity. Based in Austin, Cyara provides software and services for testing and validating CX systems.

Editor's note: The following transcript was edited for length and clarity. 

What are your impressions of the recent stories about ChatGPT, Anthropic and Meta AI models going rogue and hacking into other systems?

Seth Johnson: The first thing is I've been impressed -- for lack of a better word -- that these organizations came forward and basically owned it. The disclosure is a big step, and recognizing that 'we did something wrong and we can do better' is a key to improving things. Anthropic did it after the fact and saw that there were a couple of incidents that had been disclosed that they didn't even know about and then informed a couple of their customers. It was a bit of egg-on-face, certainly, but that's the kind of information you want those large organizations to lead with.

Seth Johnson, CTO at CyaraSeth Johnson

Was there anything specific about the way these models operated that led to the vulnerabilities?

Johnson: At the end of the day, they weren't necessarily model problems -- they were what we call harness failures. The ecosystems in which these operate were not protected or validated the way they needed to be. Specifically, the Anthropic situation was done in an environment where they told the model and the test to solve a particular problem and that it was not connected to the internet. But it was in fact connected to the internet, which it never should have been. That should have been something that was addressed and validated and assured ahead of the tests that were being run. [The model] found out that it did have access to the internet, so as it was trying to do its job and solve its problem. It went out and wreaked havoc outside the environment in which it was operating. It's interesting to see that both situations were slightly different, but they could have been avoided by scoping the environment using the instructions you provided.

How do you do that?

Johnson: You can indicate with prompts that this is how we want you to behave, but you also need to ensure that the environment variables align with the instructions that you provided the model. Sometimes there could be competing instructions where you give the model specific instructions, and it says 'this is more important than that' and goes outside the bounds of some other instruction. So, you need to be sure that those boundaries are in place, and saying something like 'this is the simulation' isn't sufficient. You need to ensure that those items are present in the environment.

Were these models going rogue or were they just doing what they were supposed to be doing and found vulnerabilities?

Johnson: I'd say the latter. [Based on what we've seen], it appears they were doing exactly what they were instructed to do. It just had too much leash, I suppose, in terms of where it could operate and what it could do.

Companies need to understand that they need to build technology and capabilities around providing those guardrails to reduce that blast radius.
Seth JohnsonCyara

How do companies need to think about these new systems that are continually adapting and evolving?

Johnson: If you're a customer of one of these frontier models -- like OpenAI's, Anthropic's or Google's -- and you're using their LLMs to do things, you can't expect that they are going to be your guardrails. It's an LLM -- it's supposed to be able to tell you all kinds of things. You need to be accountable and responsible for building those guardrails and ensuring that the information your applications, bots or otherwise are requesting is within the area you want them to be. It's building guardrails to ensure that out-of-bounds questions are not answered. Companies need to understand that they need to build technology and capabilities around providing those guardrails to reduce that blast radius. And certainly, they need to have testing and monitoring after the fact to ensure that it's actually doing it.

Do enterprises need to think about security differently than they have in the past?

Johnson: If we look specifically at the Hugging Face situation, it found and then exposed a vulnerability. So, it goes back to basic security best practices, such as patching your systems and making sure you are not vulnerable to something like this. Obviously, an ounce of prevention is worth a pound of cure. You need to be able to detect that a vulnerability was exposed, but also what's happening on the inside. That's where observability and monitoring technologies come into play to help detect anomalies, pattern changes and usage changes that can help you see that something's up here. You don't know what it is, but you should take a look and go in to investigate further. You can't prevent someone from trying, but you can limit the opportunities they have to be successful and then have systems in place to tell you when they were successful and that you need to get involved.

These were very high-profile cases. Is that a wake-up call for the industry to say this is the time to start getting your house in order?

Johnson: Yes, certainly. Everybody should already be aware that these types of things can happen. This just happened to be an AI agent playing the nefarious actor instead of a group of hackers. So, organizations should already be aware of those types of activities. This is just now coming from a different threat surface. But how you can ensure that your AI is behaving and doing the things you want it to -- and not the things you don't -- is something people absolutely need to be aware of.

Will we see more of these frontier model security incidents?

Johnson: I'd be shocked if this was the last. Hopefully, companies will take it seriously, learn from it and apply the learnings to prevent it from happening again. But I would not be surprised to see it find another way at some point. You close one hole, and ultimately, some things find another. We'll continue to see, on occasion, unique circumstances where a weakness gets exposed and something happens.

Jim O'Donnell is a news director for TechTarget, where he covers IT strategy and enterprise ESG.

Dig Deeper on CIO Strategy