A new analysis of public data from Anthropic’s Project Glasswing has highlighted a significant gap between the number of vulnerability findings generated by its Claude frontier model and those that ultimately prove to be real, serious, and worth fixing.

The distinction matters because it suggests that the bottleneck in vulnerability research may increasingly lie in validating new flaws and coordinating their remediation rather than in discovering them.

Barely 10% Have Made It to Disclosure Stage

Patrick Garrity, a security researcher at VulnCheck, recently analyzed Anthropic’s Vulnerability Disclosure Ledger, which is a public record tracking Project Glasswing-related findings as they move through the vulnerability disclosure and remediation process. The analysis showed that Anthropic’s Claude Mythos generated a total of 26,153 vulnerability findings across numerous software projects since Project Glasswing’s launch in April 2026.

Related:Microsoft Issues Emergency Fixes After Massive Patch Tuesday

However, only 2,736 of those findings, or slightly more than 10%, had made it into the disclosure ledger, meaning they have either been disclosed to the appropriate software maintainer or are in the process of being disclosed. Less than 0.8% of flaws, a mere 202, are currently patched, and 245 were withdrawn. Another 191 vulnerabilities were in the pre-disclosure stage and had not been reported to their maintainers yet.

The remaining nearly 90% of Claude Mythos-generated findings had not made it to the ledger yet, suggesting human validation and coordination have become a bottleneck in determining which AI-generated findings warrant disclosure and remediation, Garrity says.

The results are “not a surprise for those of us closer to understanding how coordinated vulnerability disclosure works,” Garrity tells Dark Reading. But it “is much different than the narrative frontier model providers have positioned,” which has largely focused on AI’s ability to dramatically accelerate vulnerability discovery.

“It seems like they are learning this through trial and error,” he says.

True Positives and Severity Assessments

Garrity’s analysis also raised questions about Anthropic’s claims regarding the accuracy of Mythos’ vulnerability findings and the model’s ability to assess their severity. He noted that the 202 findings marked as fixed in the vulnerability ledger are notably fewer than the 245 vulnerabilities marked as withdrawn. The numbers warrant closer scrutiny of how Anthropic defines and measures its claimed 91.4% true-positive rate, he wrote.

Similarly, Garrity found Anthropic’s AI to be substantially more aggressive in assessing severity of vulnerabilities compared with the actual maintainers of the affected software. Claude, for instance, assessed 91.5% of the findings that made it to the ledger as being critical or high severity. However, project maintainers themselves determined only 61.3% as being in this severity category.

Related:US Government Accuses Chinese AI Firms of Distilling Frontier Models

“My gut tells me the team didn’t prompt Claude with detailed instructions on how to determine severity, or if they did, it wasn’t well thought out, resulting in higher severity determinations,” Garrity says. “I’d like to see the CVSS metrics used to generate severity and the CWEs used to help better understand the actual weaknesses, both of which are industry standards expected when disclosing vulnerabilities.”

Signs of a Larger Issue?

The questions around AI’s ability to accurately assess vulnerability findings are not unique to Glasswing. Research by Contrast Security, for instance, showed significant variability in the results produced by AI-powered security scanners. Different runs of the same tool against the same codebase produced substantially different findings, and multiple scanners agreed on only a small percentage of vulnerabilities.

Jeff Williams, founder of Contrast Security and creator of the OWASP Top 10, says when his company ran three different AI scanners three times against the same 50,000-line codebase, the scanners generated different results each time. A Sonnet-based simple scan, for example, reproduced only 17% of its own findings across the three runs, while Opus reproduced 25% but had a nearly 30% swing in finding count between best and worst runs.

Related:AI’s Vulnerability Surge May Be More Manageable Than First Feared

The disparity highlights a potential shift in the economics of vulnerability research, Williams says. AI can make finding potential flaws relatively inexpensive, but determining which findings are real, actionable, and worth fixing can require substantially more time and resources. “Three AI scanners against the same codebase agreed on 5% of findings,” he says. “Scanning a 2-million-line codebase cost around $315 in API charges and triaging the results cost around $128,000.”

VulnCheck’s analysis comes as Project Glasswing and similar efforts to apply AI to vulnerability discovery across the industry have begun producing a massive volume of newly identified flaws. The trend is creating new challenges for security and vulnerability remediation teams that have to validate, prioritize, and patch those flaws. Microsoft’s record-setting September Patch Tuesday this week, which addressed 974 vulnerabilities, offers an example of the scale of the challenge that organizations face as AI accelerates vulnerability discovery.

“The world is still doing security at human speed,” Williams says.

But finding more issues faster does not help, he says, pointing to the low number of Project Glasswing findings that have actually reached a maintainer and the even lower number that are confirmed fixed. “And that is Anthropic with unlimited compute, and Apple, Google, and Microsoft as partners. That is what happens when you point a firehose at a funnel.”





Source link

#

Comments are closed