101% more reported CVEs per day in 2026 than last year.

AI

AI vulnerability remediation: what AI fix pull requests can do

The slow part of vulnerability management was never finding the vulnerabilities. It was writing the fix, adapting the code to a breaking upgrade, and getting the change reviewed. That is the part AI is now good at. Whether its fixes help depends on what it is told to fix.

Updated

What AI remediation is

AI vulnerability remediation uses a language model to produce the fix for a known vulnerability. It comes in three shapes:

  • Code fix suggestions. A model proposes a patch next to a static analysis alert. GitHub Copilot Autofix, for example, "automatically generates code change suggestions for CodeQL alerts found on pull requests and on the default branch."
  • Dependency upgrades that adapt the code. A version bump is easy. The hard part is the renamed function or changed default the new version brings. A model can make those follow-on changes in the same pull request.
  • Agentic remediation. An agent carries the fix through several steps on its own: change the code, run the build and the tests, read the failures, try again, and open a pull request only when the checks pass.

Where AI fixes earn their keep

Rule-based update bots already open pull requests that bump a vulnerable package to a fixed version. They stall when the upgrade breaks something, because adapting the calling code is not a rule. That is where the backlog piles up: the security update that sits open for months because nobody has time to work through the major version change.

Reading a changelog, finding the call sites and rewriting them is the kind of work language models do well. An AI fix that arrives with the upgrade and the code changes, built and tested, turns a week of someone's attention into a review.

Where AI fixes go wrong

The clearest statements of the limits come from the vendors themselves. On Copilot Autofix, GitHub writes:

  • It "uses a generative model that is non-deterministic. Even with the same alert and code, it might fail to produce a viable suggestion, or the suggestion might vary across attempts."
  • "The system does not know which dependency versions are supported or secure, and may suggest fabricated dependencies published under statistically probable names."

The second point is not hypothetical. Research on code-generating models has measured how often they recommend packages that do not exist, names an attacker can register and fill with malware. A fix that installs a hallucinated package replaces one vulnerability with a worse one.

Three guardrails cover most of the risk:

  1. Build and test every fix before a person reviews it, so review time goes to fixes that work.
  2. Check every added or changed package against the registry and the version you expect.
  3. Keep a person on the merge. An AI fix is a contributor's pull request, not a decision.

Garbage in, pull requests out

An AI remediation tool fixes the vulnerabilities it is given. If that list is reconstructed from lockfiles or a scanned image, it inherits every gap in the reconstruction: tools the build downloads that no manifest declares, versions that differ from what the build resolved, packages restored from caches. The tool then opens confident pull requests for some vulnerabilities and never hears about others.

Some teams go further and ask a model to work out the dependencies themselves. That makes the inventory non-deterministic too, which is the one part of the process that should come out the same every time. AI-generated SBOMs covers why a guessed inventory is not evidence.

The pattern that works is deterministic data underneath and AI judgment on top: a recorded list of what each build actually contains, matched against advisories by lookup, and AI writing the fix.

Where CRACI fits

CRACI is the deterministic layer. It runs your GitHub Actions jobs on its own runners and records the packages each job actually fetched, including transitive dependencies and packages restored from caches, with a completeness state for every job. It re-evaluates monitored SBOMs continuously, so a new advisory shows up against the builds it affects, and security teams triage each finding and send it to the team that owns the fix.

That is the list an AI fix should start from, and after the fix merges, the next build's SBOM shows whether the fixed version is what the build really used. See vulnerability tracking for the full capability, and AI vs deterministic security for where each approach belongs.

AI vulnerability remediation: frequently asked questions

What is AI vulnerability remediation?

AI vulnerability remediation uses a language model to produce the fix for a vulnerability: a code change for a static analysis finding, or a dependency upgrade plus the code changes the upgrade needs. The output is usually a suggestion or a pull request for a person to review.

What is agentic remediation?

Agentic remediation gives an AI agent the tools to carry a fix through several steps on its own: read the finding, change the code, run the build and tests, and iterate until they pass, then open a pull request. The difference from a one-shot suggestion is that the agent checks its own work before a person sees it.

Are AI-generated security fixes safe to merge?

Treat them like any other contributor's pull request. GitHub describes Copilot Autofix as non-deterministic and warns that it may suggest fabricated dependencies, so review the change, run the tests, and check that every package it adds exists and is the one you expect.

Can AI find vulnerabilities in dependencies?

It is the wrong tool for that part. Whether a package version is affected by a known advisory is a lookup, and it should come out the same every time. Let deterministic tooling produce the list of vulnerable packages from a recorded SBOM, and use AI where judgment helps: writing and adapting the fix.

What does an AI remediation tool need to work well?

An accurate list of what is actually vulnerable, a build and test suite it can run the fix against, and a person on the merge. The first is the one most often missing: a fix pull request for a package the product does not ship is noise, and a vulnerable package the inventory missed never gets one.

Give your fixes the right starting point

Book a demo and run one of your workflows on CRACI. See every vulnerable package the build actually pulled in, recorded rather than guessed.

Book a demo