The “current” Haskell test runner Dockerfile uses haskell:9.2.7-slim-buster. Debian Buster is past end-of-life and the image no longer builds. (The apt-get update fails as the repo is no longer there.) The Docker image I pulled from Dockerhub clocks in at an impressive 2.65GB.
I attempted to rebuild the image using haskell:9.14.1-slim-bookworm as the base. Unfortunately, that clocked in at 10.2GB. I was able to reduce that to 8.83GB by using a builder layer and then copying over /opt. The /opt can be paired down a bit, eg dropping ~878MB by clearing out docs, but it’s still rather large. /opt/ghc/9.14.1 is 3.7GB and /opt/test-runner/.stack is 4.8GB, split between /opt/test-runner/.stack/pantry and /opt/test-runner/.stack/programs.
I’m not sure why it got 4-5x larger from 9.2.7 to 9.14.1 nor what all those files are and if something can be slimmed down.
Stack can only use a system GHC installation if its version is compatible with the configuration of the current project
So I think you need to update the resolver to avoid Stack installing two versions. The most recent LTS seems to be LTS Haskell 24.39 (ghc-9.10.3) :: Stackage Server which appears to be compatible with GHC 9.10.3. That GHC version was released just last September so that’s not too old given we’re using 9.2.7 which is a bit over three years old now.
Did you update the stack.yaml files for each test suite as well?
We’ve mentioned removing the docs in/opt/. What about cached files still inside that folder?
Stack uses Pantry for managing package snapshots so we’d need to keep the snapshots downloaded from the Hackage package repository. However, Pantry, part 1: The Package Index indicates an index of package metadata exists in a downloaded form which is also imported into a SQLite database, presumably in the .stack/pantry folder. Hackage has almost 18k packages, and each published revision gets a separate package metadata file so that surely takes up a chunk of space. We wouldn’t need these files because all the necessary packages are already installed and we don’t need to grab additional ones to compile student code.
I deleted /opt/ghc/9.10.3/lib/ghc-9.10.3/lib/x86_64-linux-ghc-9.10.3/ and was still able to get some exercises to run and pass the tests. However, that directory is present in the base image. If we use FROM haskell:9.10.3-slim-bookworm we would get that. I’d need to rebuild things FROM debian:slim-bookworm, replicate some of the layers/commands used to build haskell:9.10.3-slim-bookworm then merge that with parts of the haskell-test-runner/Dockerfile. This shouldn’t be too difficult but maybe a task for another day given the hour here.
A count of how many times (ignoring single hits) a base image is used across Dockerfiles:
My understanding is those all fetch the latest version of the tag. If a new image is pushed, they do not use the same image so we may have, say, 9 different ubuntu:24.04 images. However, as of right now those all point to the same hash. We could update them all with one hash and use one base image … at the cost of all the rebuilding.
Would it be reasonable to make PRs for all the test runners that use those images to pin them all to the same SHA? This would impage 38 test runners.
Hooray for shaving some more off. It’s curious why we’re still so bloated compared to before. I would have thought using a compatible resolver and updating the stack.yaml files would be enough. Something still seems off because even if we carve out these bits, the maintainers didn’t need to do that previously for a much smaller size. However, I don’t know much if anything about the Haskell ecosystem besides my Google hits so far.
As for the base images, I think that makes sense. If we pin them to a SHA, would Dependabot update the images for us? How do we go from alpine:3.23.4 to alpine:3.23.5 for example? When dealing with commit SHAs, one SHA is after another so Dependabot can easily bump to the next one. However the SHAs are per tag, right?
I don’t think Dependabot handles Dockerfile hashes. At any one point of time, a tag points to one hash. The tag can be updated to a different hash. The tag isn’t actually needed when using a hash; the hash uniquely identifies the image.
If we want to reduce the number of images we store, we would need to manually keep them in sync and update cross-track. If a track wants to update from alpine 3.23 to 3.24, they would ideally/hopefully use the same hash that we’re already using elsewhere for the 3.24 image.
That’s nifty. Though I’m not sure how that would work. Assuming all repos run Dependabot on the same schedule and the tag is stable across the “run”, that should keep all images on the same hash. If Dependabot takes long enough to process all the repos, we could end up with different hashes … though that shouldn’t be too bad if there’s only a few. I’m not sure how critical it actually is to regularly bump the image hash, though. My understanding is that new images are built as part of the PR process so runner repos that go months/years without a PR would be running an image that is months/years old.
What if we provided base images like exercism/debian-slim, exercism/alpine etc.? Then every individual tooling repo doesn’t have to get spammed with dependabot PRs all the time. Just update the base image hash in one place and trigger a rebuild of every tooling image. Should be possible to automate that with the workflow_dispatch trigger of github actions.
That’s an interesting idea. It should simplify the process a fair bit. The downside is that we wouldn’t actually drop older versions of base images until we trigger a rebuild.
We could use the latest tag for other stuff, too, but prefer explicit hashes and noisy updates via dependabot. Using explicit hashes for the images seems in line with that.
we wouldn’t actually drop older versions of base images until we trigger a rebuild.
We’ll yeah, but that’s the same problem as “we wouldn’t actually drop older versions of base images until we merge PRs to update hashes in every test-runner repo.” There needs to be automation either way.
We could use the latest tag for other stuff, too, but prefer explicit hashes and noisy updates via dependabot.
The difference being that we control the latest tag in this case.
You can start a workflow with an API request instead of a PR when the workflow_dispatch trigger is enabled. Updating the base images would trigger its own workflow that sends API requests for each tooling image, triggering all of those rebuilds at once.
I do not see how providing an own base image or using SHAs for tags can help with not replicating existing toolchain images for all of the 82 languages? PHP for example uses php:8.4.10-cli-alpine3.22 as the base image - simply because replicating and maintaining thousands of build script lines in php-test-runner for the tooling makes no sense. I think the same goes for node:22-bookworm-slim, python:3.13.5-alpine3.22 etc.
We can’t use a single base image for all images but we can reduce the number of base images used. We have an image for the PHP test runner and PHP representer. If they used the same base tag (which they don’t, by the way) but get built at different times, the tag can point at different hashes and user different base images. By pinning the hash, we use the same base image for both. Go was using three base images for the test runner, analyzer and representer. By pinning the hash, we reduce the number of base images from 3 to 1.
Many tracks use Alpine or Ubuntu as the base image. If Ubuntu has a monthly update and we assume the various test runners, representer and analyzers all had their last PR across 24 different months, we would be able to go from 24 different copies of Alpine/Ubuntu to 1 each. We’d still need multiple base images (Alpine, Ubuntu, PHP, Go, Python, Node, etc) but hopefully only one of each. (Or, rather, one Ubuntu 22, one Ubuntu 24 and one Ubuntu 26.)
IIRC base images like Node are shared across multiple tracks (Javascript, Typescript) so we’d actually have up to 6 different copies of node that can be reduced to one.
I like that idea. We could have a repo that builds the base Exercism images. Some Exercism images share a fair bit of common setup across multiple repos (eg the runner/analyzer/representer); that setup could potentially be baked into the Exercism base image and shared that way.
Switching to a setup like this might need to go through @iHiD, though. Erik would probably be able to set up a “base images” repo for us.