Project Glasswing finds security flaws but lacks transparency on impact and remediation success.
Summary
- Anthropic’s Project Glasswing claims superior vulnerability detection capabilities, but these claims lack independent verification
- The project’s status report shows many vulnerabilities discovered but almost none have been patched
- Anthropic refuses to release detailed data, asking stakeholders to simply trust their assertions
- Security expert Bruce Schneier questions the validity of claims that Mythos outperforms other models at vulnerability detection
Anthropic launched Project Glasswing in April with considerable fanfare, positioning its new model as a breakthrough tool for companies to identify and remediate vulnerabilities in their own software. The initiative generated significant media coverage, with many outlets repeating the company’s assertions without critical examination.
Claims versus reality
The widespread media coverage has established what Schneier describes as “common wisdom” that Mythos surpasses other models in vulnerability detection capabilities. However, this perception lacks supporting evidence from independent verification or comparative analysis.
Mixed results in practice
Anthropic’s recently published status report presents a puzzling picture. While the project claims success in discovering numerous vulnerabilities across various software systems, including some classified as dangerous, the remediation outcomes tell a different story. The report shows that almost none of the identified vulnerabilities have actually been patched.
Transparency concerns
The lack of detailed information from Anthropic raises questions about the project’s actual effectiveness. Security researcher Bruce Schneier notes irregularities in the available data and criticises the company’s reluctance to provide specifics, describing their “trust us” approach as problematic for proper evaluation of the technology’s capabilities.
Why it matters
CISOs evaluating AI-powered vulnerability detection tools need reliable data to make informed procurement decisions, but vendors making unsubstantiated claims without transparency create risk assessment challenges and potential false confidence in security capabilities.
What to do now
- Demand detailed performance data and independent validation when evaluating AI vulnerability detection tools
- Question vendors who refuse to provide specifics about their security tools’ effectiveness
- Focus on measurable outcomes like successful patch rates rather than discovery claims alone
