CI/CD
GitHub Actions cache: how it works and what it hides
The GitHub Actions cache stores files between workflow runs, so jobs skip downloads they already made. It makes builds faster. It also means packages arrive in your build without being fetched, which most records of a build never see.
Updated
How the cache works
The actions/cache action saves a set of paths under a key at the end of a job, and restores them at the
start of a later one. A typical npm setup caches the package manager's download cache and keys it on the lockfile:
- uses: actions/cache@v4
with:
path: ~/.npm
key: npm-${{ runner.os }}-${{ hashFiles('package-lock.json') }}
restore-keys: |
npm-${{ runner.os }}-
- run: npm ci
On restore, the action looks for an exact match to key first, then a partial match, then each
restore-keys prefix in turn. When several caches match a prefix, it restores the most recently created
one. That fallback is what keeps builds fast after the lockfile changes, and it also means a job can start from a
cache that was built for an older lockfile.
You often do not need the action at all. setup-node, setup-python, setup-java,
setup-go, setup-ruby and setup-dotnet have caching built in. For Node.js,
setup-node caches global package data for npm, Yarn and pnpm, not node_modules, and turns
npm caching on automatically when package.json declares npm as its packageManager.
Scope and limits
| Rule | What GitHub does |
|---|---|
| Which caches a run can restore | Its own branch and the default branch; pull requests also their base branch |
| What a pull request can write | Only caches scoped to its merge ref, not to the default branch |
| Storage | 10 GB per repository by default; organization and enterprise owners can raise it |
| Eviction | Entries not accessed for over 7 days, and the least recently used when storage is full |
| Secrets | Do not cache them: anyone with pull request access can read cache contents |
What the cache hides from your SBOM
A cache is a shortcut past the package registry. That is the point of it, and it is also why it confuses every record of what a build used:
- No download, no trace. Packages restored from a cache are installed without a request to the registry. A record built from the build's network traffic misses them unless it follows the cache.
- Old contents under a new key. A partial
restore-keysmatch restores a cache built for an earlier lockfile. The install then adds or replaces what changed, but what the job started from is not what the current lockfile describes. - More than your dependencies. Caches also hold build tools, compilers and generated files, which no application manifest lists.
- A scan cannot tell. A lockfile-based SBOM describes what the lockfile says, whether or not the job installed exactly that from its cache.
For how this fits the wider problem, see why a lockfile is not a record of what shipped and transitive dependencies.
Cache poisoning
Caches are shared across workflows in a repository, which makes them a path between them. In May 2024, security
researcher Adnan Khan showed that anyone who can execute code in the context of the default branch, through a
vulnerable pull_request_target workflow or a compromised dependency, can overwrite files in workflows
that restore caches. With the cache token from a running job, an attacker can fill the repository's cache so
GitHub evicts legitimate entries, then write poisoned ones under the same keys. A release workflow that restores
them runs the attacker's files with its own secrets.
To reduce the risk:
- Build releases from a clean state, or from caches that only the release workflow writes.
- Never run untrusted pull request code in a default-branch context.
- Keep secrets and tokens out of cached paths.
- Treat what a cache restored as an input to the build that needs the same scrutiny as a download.
How CRACI handles caches
CRACI is a GitHub Actions runner, so actions/cache and the setup-* caches work as they do
today, alongside container layer caching. Caching is also part of why CRACI runs builds about twice as fast as
GitHub-hosted runners, with faster hardware and shorter queue times.
What changes is the record. CRACI's package-aware proxy records the packages a job fetches, and evidence travels with caches: when a job restores a cache, CRACI carries the dependencies it knows about forward into that job's SBOM.
- Completeness per cache. Every cache has a completeness state, like every job. A job that restores an incomplete cache gets an incomplete SBOM, and a cache it produces carries that status forward, so a gap never turns silently into a complete record.
- No guessing from the archive. CRACI does not read package manifests inside the cache archive. It knows what is in a cache because it recorded the job that produced it.
- One cache per producing job keeps that evidence reliable, together with an egress policy set to
default: denyin the producing job. The cache and SBOM completeness guide has the details.
See build-time SBOMs for how the recording works, or book a demo to run one of your cached workflows on CRACI.
GitHub Actions cache: frequently asked questions
How long does GitHub keep an Actions cache?
GitHub removes cache entries that have not been accessed in over 7 days. When a repository goes over its storage limit, 10 GB by default, it also evicts caches in order of last access, oldest first.
Does setup-node cache node_modules?
No. The cache input of actions/setup-node caches the package manager's global package data for npm, Yarn or pnpm, not node_modules. The install step still runs, but it reads packages from the restored cache instead of downloading them.
Can a pull request read caches from the main branch?
Yes. A workflow run can restore caches from its own branch and the default branch, and a pull request can also restore caches from its base branch. Caches created in a pull request are scoped to its merge ref, so they cannot be written into the default branch's scope.
Is it safe to use the Actions cache in release workflows?
It is a trade-off. Research published in 2024 showed that anyone who can run code in a default-branch workflow can poison caches that other workflows restore, including release workflows. Workflows that publish releases or hold deployment secrets are safer building from a clean state or from caches only they write.
Does a restored cache change my SBOM?
It can. Packages restored from a cache are installed without being downloaded, so a record built from network traffic misses them unless it tracks the cache. And a scan of the lockfile does not tell you whether the cache held exactly those versions.
Keep your caches and a complete SBOM
Run your workflow on CRACI: caches still speed up the build, and the dependencies they carry stay in each job's recorded SBOM.
Book a demo