Files
shelfmark/.github/workflows
splitsec2 ca25448529 perf(docker): keep the heavy build layers cacheable across builds - save 11minutes per build (#1379)
With the amount of changes and testing I've been doing lately I noticed
how long the builds were taking and how much each one pulled, so I went
and looked at the Dockerfile. I think this balances cache and
efficiency.

Three changes needed to be stacked to make it happen.

**The version stamp sits above everything expensive.** `ARG
BUILD_VERSION` and `ENV BUILD_VERSION` are at the top of the `base`
stage and the value carries the commit sha, so it changes on every
commit. This invalidates the layer and everything below it, which means
the apt install, the dependency sync and the Chromium install. Nothing
in the build reads either variable. They're only used at runtime by
entrypoint.sh, tor.sh, wireguard.sh and genDebug.sh, so they can move to
the end of the final stages.

**`COPY . .` sits above the Chromium install.** It's in `base`, and the
`shelfmark` stage installs Chromium and the seleniumbase drivers after
it, so any source change rebuilds those too. Moving the source copy and
the runtime-paths block to the end of each final stage solves that.

**There's no cross-run cache.** Runners are ephemeral, so without
`cache-from` and `cache-to` every layer is rebuilt on every run whatever
the ordering, and a rebuilt layer gets a new digest even when the
content is identical. That's why reordering on its own doesn't change
the load.

Measured with a source-only change between two builds:

|  | Build | Pull |
|---|---|---|
| before | 13m 20s | 526.5 MB |
| after | 2m 25s | 3.7 MB |

15 of 18 layers get reused where it was 5. The cache sits at 0.86 GB,
which leaves room under the 10 GB repo budget for the uv caches in
ci.yml and e2e-platform.yml. I tried `mode=max` first and it built a bit
quicker at 1m 43s, but it used 5.15 GB of cache and the potential to
impact other workflows so it didn't seem worth 40 seconds, but that is a
single line fix if you want to.

This means the runtime-paths block is now duplicated, once per final
stage, and that's most of the diff. It has to sit below each stage's
heavy layers to do its job and I couldn't find a way around that short
of another shared stage, which looked like more complexity for
complexity's sake. Happy to take another run at it if you'd rather have
it 'DRY'.

I checked the built image against the current one. Same size, it boots,
/api/health returns 200, and BUILD_VERSION and RELEASE_VERSION are still
stamped correctly.

Also, I'll slow down on the PRs. Promise.
2026-09-25 18:07:05 -04:00
..