When Anthropic introduced Project Glasswing in April, the tech industry braced for impact. Five months in, DevSecOps teams are wading through the glut of discoveries to find out what's worth patching – and so far, it isn't much, according to one cybersecurity researcher.
Anthropic updated its Vulnerability Disclosure ledger for the first time in August. A small percentage of initial Mythos vulnerability findings made it into the ledger, and fewer than 1% of the vulnerabilities Mythos found were marked as fixed, according to Patrick Garrity, a security researcher at vulnerability prioritization tool maker Vulncheck.
Garrity and Vulncheck have also been tracking whether vulnerabilities discovered with AI are actually exploited in the wild. This research combined vulnerabilities attributed to Anthropic's model with data from the Berkeley Vulnerability Research Initiative and correlated them with Vulncheck's known-exploited vulnerability data.
Of the 1,061 vulnerabilities attributed to AI-assisted discovery, only 14 had been listed as exploited in the wild as of June 2026. That's roughly 1%.
"You discover a vulnerability -- that doesn't mean it's necessarily useful for an attacker. It doesn't mean it's necessarily going to be exploitable with certain conditions," Garrity said.
The discrepancy between the volume of findings and the rate at which they've been patched points to a significant bottleneck around the human triage step, where DevSecOps teams determine what findings matter, how bad they are, whether they should fix them at all, and how they should fix the vulnerability if necessary, software security experts said.
"There was so much agita around [the] 'Vulnpocalypse', around everybody being able to find new zero-day vulnerabilities and exploit them, the impact that was going to have," said Neil Carpenter, an independent security evangelist. "And then the secondary piece of, 'How do we fix so many vulnerabilities?' It becomes a burden on defenders, both developers and blue teams downstream."
People think that AI is magic, that a frontier model is going to magically do this stuff for you. It isn't.
Michele ChubirkaSr. principal security architect, Red Hat
Teams can use AI to help them parse the vulnerabilities that have piled up after Glasswing, but that's far from removing the human from the loop, according to Michele Chubirka, a senior principal security architect who participated in Project Glasswing at Red Hat. Chubirka and platform engineers at the company built a workflow engine which she described as "part harness, part integration of deterministic security tools" to create a scanning framework for all the repositories under the vendor's OpenShift product organization.
"People think that AI is magic, that a frontier model is going to magically do this stuff for you. It isn't," Chubirka said, emphasizing that she was speaking in her personal capacity, not on behalf of Red Hat.
AI in production workflows also introduces entirely new security risks and bugs of its own, according to Garrity.
"No one is talking about... the risks [that arise] when you deploy these technologies," Garrity said. "The bigger shift I'm seeing is actually the attack surface expanding substantially. AI products themselves are being targeted and hit very fast."
Glasswing results: humans 'rate-limiting'?
Garrity's analysis showed that Anthropic reported 26,153 total findings; 2,736, or 10.5%, had reached the ledger. 2,096, or 8%, had been reported to maintainers. 202, or 0.8%, were marked fixed. 245 had been withdrawn. Anthropic's ledger described independent human triage as a "rate-limiting step."
Garrity's analysis of Anthropic's ledger also pointed to a discrepancy in the assessment of vulnerability severity between Claude and actual maintainers. Claude concluded that 91.5% of findings were of high or critical severity, but maintainers' assessment of the same group found only 51.3%.
"The Glasswing approach came out and was like, everything is critical and high. What we're finding is the maintainers are scoring things not all critical and high," said Garrity.
For example, in May, Daniel Stenberg, the founder and lead developer of the Linux curl command-line utility project, wrote in a blog post that Mythos had identified five confirmed vulnerabilities, but that his team boiled it down to just one.
"The other four were three false positives (they highlighted shortcomings that are documented in API documentation) and the fourth we deemed 'just a bug,'" Stenberg wrote. "The single confirmed vulnerability is going to end up a severity low CVE planned to get published in sync with our pending next curl release 8.21.0 in late June. The flaw is not going to make anyone grasp for breath."
The Claude team might not have had the specific expertise necessary to guide Mythos in accurately assessing vulnerabilities, said Garrity.
"This wasn't their subject domain. And so naturally they took their tool and threw it at a ton of open source projects and it came back with a ton of findings," he said. "If you don't have the person with the subject matter expertise building the [agent] skills to do the task, ultimately you're going to end up with a lot of noise."
Anthropic reps did not respond to a request for comment as of press time.
AI vulnerability discovery creates triage bottleneck
AI creates a lot of noise, but that doesn't mean it's not useful at all, Garrity said. When people disclose vulnerabilities to Vulncheck, they can say whether AI assisted them in finding it.
"[About] 45% of people self-select that," said Garrity. "That's just the people willing to say they are [using it]."
An Omdia survey report from August stated that "nearly nine in 10 organizations report that agentic AI is having a moderately (40%) or significantly (49%) positive effect on improving visibility into the security hygiene and posture of IT assets, leading to better security decisions."
Still, it takes a lot of effort to cut through that noise and apply those findings in a meaningful way. Organizations spend a lot of time on triage, working with researchers to determine if a vulnerability is actually important, Carpenter said.
"The real benefit of agentic AI is in the ability to take a bunch of findings from some deterministic tools, from the code itself, to take that all together, to analyze it in a big bucket, make some judgment calls and some evaluations that might take a long time to get to [independently], [and] to chain findings in a novel way," he said. "That kind of stuff with a security professional can take a long time."
The value rises when it's paired with expert guidance for analyzing large swaths of data, Chubirka said.
"You have a base report, you have machine triage, you have humans eyeballing to help triage, to confirm or deny certain findings that maybe are questionable," she said.
Chubirka's team used AI to create patches after triage, but a significant portion of them needed an additional human touch.
"Sixty percent of the patches were generally good out of the box; 20% were like, 'Eh, maybe I wouldn't have done it that way -- it doesn't really conform to my design principles or a design guide,'" Chubirka said. "Then the other 20% were like, 'You're crazy town, banana pants.'"
Chubirka said she mined AI vulnerability findings based on the team's triage to verify and remediate them. That's the basic concept of Project Glasswing, she said: creating a repeatable VulnOps workflow. VulnOps is a continuous approach to vulnerability management, rather than handling vulnerabilities in batches.
But applying that principle is no mean feat and requires subject matter experts, she said, recalling 45 straight hours of testing for false positives and negatives during her team's remediation project.
"Even with AI, it was five people doing this work, ultimately," Chubirka said. "The only way you're going to do this at scale is to treat it a little bit like a factory."
"No one is talking about... the risks [that arise] when you deploy these technologies," Garrity said. "AI products themselves are being targeted and hit very fast."
Part of the danger is giving AI access to sensitive information and incorporating AI tools into infrastructure faster than security and governance teams can keep up.
"They're deploying the technologies and [giving it the] keys to the kingdom. Everyone's like, 'Oh, the AI needs to know all my information in Salesforce. It needs to have control of AWS. It needs to have control of all these different processes,'" Garrity said. "The whole concept of least-privilege access was basically just thrown out the window in a lot of these organizations."
Ben Lutkevich is an award winning writer for TechTarget.