Anthropic found 29,000 possible bugs and ran out of people to check them
11 Oct 2026 · News
AI got good at finding software bugs faster than anyone could confirm them. The fix Anthropic landed on says a lot about what makes machine output worth trusting.
The bottleneck in AI bug hunting is no longer finding bugs. It is someone saying "yes, that one is real."
On Thursday Anthropic launched OSS Scanner, a free, opt-in service that scans important open-source projects with its strongest models and sends the maintainers whatever it finds. The catch is stated up front: the reports are "fully model-generated, without human review or triage," so some will be wrong.
The backstory is the interesting part. Over six months, Anthropic's models turned up more than 29,000 candidate vulnerabilities. Its people managed to review and triage about 6,000. Meanwhile, maintainers who got the first verified reports started asking for the rest, unchecked, patches and all. Anthropic says it has sent nearly 5,000 of those on request.
Think about how odd that is. Not long ago, open-source maintainers were complaining about floods of junk AI bug reports. Now some of them are asking for the raw feed.
What changed is less about the model being smarter and more about what comes attached. Each report includes a self-contained reproducer (a small program that triggers the bug), an explanation, and a candidate patch where possible. A maintainer at OpenSSL put it bluntly in Anthropic's post: when a report comes with a working exploit, "that's basically job done for an engineer as you can verify it right away."
That is the real shift. You do not have to trust a report that proves itself. You run it and see. A claim that "this code is unsafe" costs a maintainer an afternoon of doubt. A claim that ships with a crash you can reproduce in a minute costs almost nothing to check.
The numbers Anthropic shares look strong. Its own penetration testers checked 97 critical and high-severity findings across 48 projects: 85 met its bar, 11 were real but duplicates, and one was invalid. The team behind wolfSSL, an encryption library, said all but two of 74 reports were valid. Worth remembering: this is Anthropic grading its own early version, with quotes it chose to publish. The honest test is what maintainers say a year from now, when the scanner reaches projects with fewer volunteers and less patience.
There is also a quiet cost shift. The human review step did not disappear. It moved from Anthropic's security team to the maintainer's inbox, often someone unpaid. Making it opt-in, and saying plainly that nobody checked the output, is what keeps that fair. Projects choose the trade with their eyes open.
The broader lesson reaches past security. AI output gets useful at scale when it carries its own evidence, not when it sounds confident. A bug report with a reproducer, a math proof that a checker can verify, a code change with passing tests: those are things a busy human can accept quickly. Everything else is just more email.