Cybersecurity vendors turn to multi-model, open AI to bypass guardrails

Commercial AI guardrails are blocking legitimate cyberdefense activities. Security vendors are fighting back with multi-model frameworks and open-weight LLMs.

Recent high-profile security incidents highlight a growing problem across the cybersecurity industry: Commercial AI models frequently mistake legitimate security analysis for malicious activity. To prevent upstream AI safety guardrails from stalling critical security workflows, security vendors are now actively decoupling their products from single-model dependencies.

When Hugging Face found itself under attack by rogue AI agents, incident responders -- confronted with 17,600 intrusion logs -- ran the investigation through their own AI-assisted pipeline. But when the security team asked Claude Opus and Fable to analyze attack commands, exploit payloads and command-and-control artifacts, the models largely refused.

"Their safety guardrails treated reverse-engineering an exploit the same as launching one," Hugging Face engineers wrote in a blog post. Instead, they rerouted the pipeline through an open-weight model, GLM-5.2, on their own infrastructure.

Hugging Face's experience is far from an isolated case. As commercial LLMs enforce increasingly rigid safety alignment, security teams and software vendors find themselves trapped when models flag benign threat research as malicious behavior. To maintain reliable incident response capabilities, vendors are insisting on architectural agility -- establishing multi-model frameworks and self-hosted, open-weight LLMs to bypass upstream refusal gates.

"Adversaries will use whichever AI model gives them an edge, and they will switch the moment a better one appears," Greg Clark, CEO of threat detection firm ExtraHop, told TechTarget. "Defenders need the same freedom. They shouldn't be tied to one provider's changing policies or capabilities."

Navigating AI refusal gates with multi-model flexibility

To navigate the unpredictable tripwires laid out by commercial AI labs, security vendors are turning to multi-model architectures that allow SecOps workflows to automatically swap or re-route prompts when a model declines a query.

The Agentic SOC Alliance, for example, is a coalition of 15 cybersecurity vendors founded in July. The group has proposed a reference architecture that separates a SOC's security telemetry and governance from underlying AI models, making execution layers easily interchangeable when a model breaks a request.

"A workflow that can move between frontier models, smaller specialist models and deterministic automation gives the vendor more options when access changes or an intervention prevents completion," said Andrew Braunberg, analyst at Omdia, a division of Informa TechTarget.

A recent lab exercise at ExtraHop illustrates the real-world need for multi-model flexibility, according to Paul Giorgi, vice president of global engineering. When his team asked AI agents to investigate exploitation of a Windows DNS vulnerability, many models refused. Naming the CVE and asking about signs of active exploitation triggered guardrails.

"What finally worked was describing the same investigation as observable network behavior -- a DNS server making connections it has no business making -- with a structured, step-by-step task," Giorgi said. "That version completed on most models. One major model refused no matter what."

The takeaways, he added, are that models have tripwires and refusal gates in different places, and no single model is best at every SOC role.

Research from Palo Alto Networks underscores the necessity of a multi-model approach across security use cases, including vulnerability detection. In testing across complex environments, the company found that no single AI model detected more than 40% of vulnerabilities, with major models showing surprisingly little overlap in findings. The company recently put that research into practice with its Unit 42 Continuous Frontier AI Defense service, which distributes queries across Anthropic's Claude Mythos, OpenAI's GPT-5.6-Cyber and open-weight models.

Bypassing third-party guardrails with open-weight models

While multi-model routing offers operational flexibility, other vendors are taking model control a step further by building directly on open-weight foundations. These models are publicly available for users to download, modify and run on their own infrastructure.

CrowdStrike -- a member of the Agentic SOC Alliance, along with ExtraHop -- recently debuted SafeMind, an agentic harness-model system with complementary red-team and blue-team capabilities. CrowdStrike purpose-built the system's underlying adversarial and defensive models, Red Tempest and Blue Solano, on Nvidia's open-weight Nemotron models and trained them on its own data.

"Moving to open-weight models gives CrowdStrike full control of its solution, with no external vendor able to throttle, restrict, or revoke access to the capability its defensive product depends on," Omdia's Braunberg said.

For enterprise security leaders assessing AI-driven SecOps tools, a vendor's underlying model strategy is quickly becoming a core evaluation metric, according to Omdia. Products tied to a single commercial LLM risk operational disruption whenever upstream providers adjust safety policies or experience outages. As vendors balance multi-model frameworks with specialized open-weight deployments, resilience is likely to hinge on keeping AI execution layers fluid and adaptable.

Alissa Irei is an Informa TechTarget news reporter covering cybersecurity.