What CISOs need to know about AIOps security
AI offers compelling benefits for IT, but organizations have to weigh operational advantages against risks that can spread unchecked through the network.
AI is a literal game-changer. It has changed everything from how we write software to how we test it. It's influenced marketing, support, sales, HR and the executive suite, and even reshaped how we write documents and send emails.
AIOps -- the use of AI to support and bolster IT ops -- has also gained significant traction over recent years. In particular, we've seen the emergence of AIOps security use cases to enhance, contextualize and refine threat intelligence and event management. It's no surprise. According to a study published by the European Journal of Computer Science and Information Technology, companies using AIOps reduced their resolution times by 62% and their alert volumes by 91%. Equally important, study respondents said they were able to predict 87% of potential service disruptions before they could affect users.
These results, coupled with advancements in LLMs and agentic AI, promise even more compelling use cases for AIOps in the future.
Yet, as more enterprises evaluate AIOps, they must also consider its security implications. AIOps can streamline operations, increase resiliency and underpin security use cases such as monitoring and threat hunting. But AIOps can also trigger risk without oversight and planning. It can fail in unexpected ways and spread risks across the most critical production operations areas and support processes.
Fortunately, by proactively planning, incorporating AIOps into systemic planning and control, and accounting for it in their risk strategy, organizations can account for potential AIOps security failure modes and stay alert to where and how risks can arise.
Let's examine what can go wrong with AIOps and how companies can adapt their programs as a result.
What can go wrong?
LLMs today increasingly support IT and technology operations. In the past year, research has spotlighted LLM-based AIOps anomaly detection and system monitoring, and there have been attempts to use the model for root cause analysis, data analysis, reporting and case creation. Cisco, meanwhile, is embracing something it dubbed AgenticOps, describing it as "a new paradigm for modern IT operations, powered by AI agents."
Yet, the combination of LLMs and AIOps can generate significant risks. In one case, an agentic integrated development environment initiated a mass file-deletion event in which it purged an entire filesystem. In another, an AI-powered tool accidentally deleted an entire production database -- an event made worse by the fact that it occurred during a code freeze.
As dire as these situations were, they occurred without adversarial input. The stakes become even higher when considering how susceptible LLMs are to prompt injection and other attacks. Even without direct control over the context window, researchers have demonstrated unwanted outcomes. For example, research from RSAC Lab and George Mason University demonstrated that attackers can manipulate telemetry information within the environment and crate unwanted outcomes.
Lastly, while not intended for AIOps use cases directly, the recent OpenClaw security challenges illustrate how adversarial attack techniques can subvert LLM guardrails to bring about unintended operational behaviors.
How to prepare
As AIOps use cases move beyond predictive AI and into generative modes that embrace LLMs and agentic AI, what can we do to stay ahead and ensure risks are accounted for?
First, understand how, where and for what purpose ops teams are using these technologies. That's easy to say, but much harder to do in practice -- particularly given organizational complexities and how rapidly existing vendors incorporate AI functionality into new releases of products already deployed.
A few strategies can help here. One is to just ask. Where a reliable and open line of communication exists, discuss usage candidly and openly. Enlist other teams' help to keep you informed. This is a two-way street, so offer your services in return -- for example, by performing and documenting a risk analysis to assist in audit response.
Another option is to use routine interactions or periodic non-optional tasks as a data-gathering strategy. If you have existing check-in or coordination meetings, use them to stay informed. If you engage for other reasons -- e.g., business impact analysis, disaster recovery testing, annual policy review, audit responses, etc. -- use those as a vehicle to capture usage information. Alternatively, partner with your technology audit team and enlist its help to identify operational changes when and how they happen.
As you do this, understand and think through the data sources and context to which an agentic component or LLM might be exposed. Any data that gets supplied to the AI -- among them logs, telemetry and infrastructure-as-code artifacts -- are all important to evaluate and determine how they could be potentially subverted.
It's true that being forewarned is being forearmed. Learning where AIOps processes are and where/how autonomy is being extended is as valuable as understanding what data elements drive those use cases.
Once you identify where AIOps is deployed, ensure risks are controlled and aligned with organizational tolerances. Numerous factors can influence this process. For example, a customer-facing production environment will have a much different risk tolerance compared with an internal, QA or development-oriented one.
Think through and assess AIOps security risks thoroughly. Keep a written record to track data sources, specific product interactions, operational processes and other factors. A baseline reference is valuable.
Ed Moyle is a technical writer with more than 25 years of experience in information security. He is a partner at SecurityCurve, a consulting, research and education company.