AI Generated Bug Reports Overwhelm Bug Bounty Programs Forcing Suspensions and Triage Changes
HackerOne sees submissions up 76 percent year over year while legitimate findings hold at 25 percent amid rise of automated scanning

Read Time 6 minutes Tags Bug Bounty AI Security Cybersecurity AI Slop HackerOne Bugcrowd Vulnerability Disclosure AI Agents Security Economics Companies that pay hackers to find flaws in their software are being inundated with low quality reports generated by AI forcing some to suspend the programmes altogether Businesses that run bug bounty schemes have long relied on independent security researchers to spot vulnerabilities But the rise of AI tools is now overwhelming them with spurious submissions Bugcrowd whose customers include OpenAI T Mobile and Motorola said the number of reports it received more than quadrupled over a three week period in March with most proving to be false Curl a widely used tool to transfer data across the internet suspended its paid bug bounty programme in January citing an explosion in AI slop reports and lower quality submissions Cyber security experts say advances in generative AI are reshaping the economics of bug bounty programmes While the tools allow experienced researchers to find flaws more quickly they are also lowering the barrier to entry triggering a flood of automated or erroneous submissions that companies must sift through Software group Nextcloud suspended its bug bounty programme in April because of the massive increase of low quality reports It said it hoped to resume the programme once it had found a way to filter submissions effectively The surge in AI generated reports comes as Anthropic last month launched Mythos its new cyber AI model which it says can find software flaws faster than humans Companies running bounty bug programmes have started to introduce more stringent background checks to combat the problem as well as building AI agents to triage submissions HackerOne whose bug reporting platform serves Goldman Sachs Google and the US Department of Defense said it had introduced new agentic validation capabilities this year to help organisations manage high volumes of findings such as those generated by models like Mythos The company said submissions had jumped 76 percent in the year to March But it said the share of reports flagging legitimate vulnerabilities had remained steady over the past year at 25 percent HackerOne chief executive Kara Sprague said it had in recent weeks seen a rise in higher quality reports that had used AI She added that the rise in AI generated submissions was not a strong reason to say we don t want them altogether given that hackers were using the technology to spot more flaws Bugcrowd chief Dave Gerry said developments such as Anthropic s Mythos would assist human bug bounty hunters not replace them AI is going to help with a lot of things but we re never going to replace that human creativity he said How AI changes the bug bounty economics One Lower barrier to entry Open source scanners and agents can now chain together crawling fuzzing and static analysis with minimal setup Amateurs can submit reports that look plausible without understanding the underlying vulnerability This increases volume but reduces signal to noise ratio Two Automated end to end systems A third cohort described by Sophos CISO Ross McKerchar consists of experienced AI builders who run automated scanning and submission systems These systems generate thousands of reports with little human review creating absolute carnage for triage teams Three Human plus AI workflows Experienced researchers use agents to accelerate recon exploit development and report writing This raises the ceiling for legitimate findings but also contributes to the flood if results are not filtered before submission Four Model specific impact Anthropic s Mythos and OpenAI s GPT 5 5 Cyber can find flaws across operating systems and browsers When made available to defenders they improve patching speed When accessed by attackers or indiscriminate scanners they increase the volume of low quality reports Operational impact on programs One Program suspensions Curl suspended its paid bounty in January Nextcloud suspended in April Both cited time spent debunking slop and mental toll on staff as reasons The cost of triage now exceeds the value of payouts for some projects Two Triage overload Triage teams must distinguish real vulnerabilities from hallucinations misconfigurations and known issues that are out of scope The process is slow because each report requires reproduction and risk assessment Three Quality remains stable HackerOne reports that despite a 76 percent rise in submissions the share of reports flagging legitimate vulnerabilities stayed at 25 percent over the past year This suggests AI is expanding the pool but not yet changing the base rate of real bugs Four Platform response HackerOne has deployed agentic validation to pre screen findings and route likely false positives to automated checks Bugcrowd is building similar filtering before human review What changes to restore signal One Reputation and gating Require historical accuracy or identity verification before allowing high volume submissions Projects can rate limit accounts with low signal to noise ratios Two Mandatory evidence standards Require proof of exploitability with minimal false positive rate Reports without a working reproduction or clear impact should be auto rejected Three AI assisted triage Use agents to reproduce steps check against known issues and score severity before a human looks at the report This reduces burden but requires careful validation to avoid missing novel bugs Four Scope tightening Projects can narrow scope to reduce noise from generic scans Target specific components or endpoints where AI assisted research adds value Five Differential payouts Reward high quality reports with clear impact and novel technique more heavily than low effort findings This aligns incentives with signal quality Broader implications One Security posture improvement If triage improves AI assisted researchers can find more real bugs faster The net effect can be stronger software despite the noise Two Cost shift Small projects cannot absorb the triage cost They may move to private bounties or community auditing rather than public programs Three Adversarial adaptation Attackers also use agents to find zero days The arms race shifts from manual discovery to agent versus agent defense and validation Four Policy considerations If federal vetting of models is implemented as discussed in the White House AI security EO the pipeline for high capability cyber agents may face access controls This could reduce indiscriminate scanning but also slow defensive use For companies running bounties the immediate priority is to automate triage without discarding edge cases For researchers the priority is to use AI to increase depth not just volume For platforms the priority is to preserve trust by keeping the signal rate stable even as volume grows Do you think bug bounty programs should require human verification of exploitability before accepting AI assisted reports or would that exclude valuable automated findings Share your view in the comments
About the Creator
Enjoyed the story? Support the Creator.
Subscribe for free to receive all their stories in your feed.
Comments
There are no comments for this story
Be the first to respond and start the conversation.