Tech
Ai-generated Patches for Software Vulnerabilities Are Largely Ineffective, With Only 26% of Attempts Proving Usable

The research, published Thursday, tested six recently disclosed open-source vulnerabilities—including flaws in Linux, ActiveMQ, Chrome, EXIM, SpringAI, and Gemini CLI—across 6,080 patch attempts using Claude and an OpenAI Codex-based model. Results showed 21% of patches fixed the bug but altered application behavior, while 53.9% failed entirely, introduced new bugs, or both. The team dubbed these flawed outputs "FLAWED" (Fix-Like Artifacts With Embedded Defects), noting they superficially appear functional but don't fully resolve vulnerabilities and may add new issues.
Keith Hoodlet, head of Off-By-1-Labs, told ZDNET that AI is better suited for vulnerability discovery and triage, helping defenders prioritize impactful bugs, while human oversight remains critical for patch decisions. The team released its FLAWED tooling on GitHub for further research.
Source: ZDNet