Module 06: Registries
Distributing Images
Meet Sam. Sam just built their first containerized app and it runs perfectly on their laptop, which is exactly the problem. You can build images (Module 02) and run full multi-container apps (Module 05). The missing piece is how images get from your machine to CI, to production, to a teammate's laptop. That answer is a registry, and this module covers everything from Sam's first push to pinning by digest for reproducible, secure deploys.
Where we left off
Go back to Module 00 for a moment. The very first diagram showed three actors: the Docker client
(your CLI), the Docker daemon (the engine running on your machine), and a
registry (the remote store). Every time you ran docker pull nginx back in
Module 01, the daemon silently reached out to Docker Hub (that's a registry) and downloaded layers. Sam, like you, has been
using registries from day one without thinking about them.
By Module 02 you learned to build images: layers, Dockerfiles, the cache. By Module 05 you could orchestrate full multi-container applications with Compose. But every image stayed local on the machine that built it. Sam hit the wall the moment a teammate asked to run the app too. The second you need to:
- deploy that image to a cloud server or Kubernetes cluster,
- share it with a teammate or a CI pipeline,
- roll back production to a known-good version,
you need to push it to a registry. That's this module.
Build (Module 02)
You can produce an image from a Dockerfile. Layers are local, cached, composable.
Run (Modules 01–05)
You can run containers, volumes, networks, full Compose stacks, all on your machine.
Share (this module)
Push to a registry so CI, production, and teammates can pull the exact same image.
What a registry is
A registry is a server that stores and distributes image layers. When you push an image, the
daemon uploads each layer as a blob identified by its sha256 content hash. When you pull,
the daemon downloads only the layers it doesn't already have locally.
That last sentence is important: layers are content-addressed. If two images share a base
layer, say both start from python:3.12-slim, that layer is stored once in the registry
and once on every machine that has pulled either image. A push doesn't re-upload layers that are already there. Docker
checks each layer's digest against the registry before uploading; only genuinely new or changed layers travel over
the wire.
Remember the layer diagram from Module 02? Every RUN and COPY
instruction produced a layer with a sha256 digest. The registry uses those exact same
digests to decide what to store and what to skip on push. Write a Dockerfile with cache-friendly layer order
and your CI pushes get faster too. Only the changed bottom layers upload each time.
Docker Hub is the default public registry. When you type docker pull nginx without
specifying a host, the daemon talks to registry-1.docker.io. The registry protocol itself
is an open standard (OCI Distribution Spec), which is why every major cloud and self-hosted solution speaks the same
push/pull API.
A registry exposes an HTTP API. Clients check which layers exist (HEAD /blobs/sha256:…),
upload missing ones (PUT /blobs/uploads/…), then push a manifest JSON that references
those layer digests. A pull reverses the process. This is why any OCI-compliant client works with any
OCI-compliant registry.
The major registries
You will encounter these on the job. Know the trade-offs.
| Registry | Best for | Auth model | Notes |
|---|---|---|---|
Docker Hubdocker.io |
Public open-source images, learning, small personal projects | Docker account / PAT | The default. Free tier has pull-rate limits (unauthenticated: 100/6h per IP, authenticated: 200/6h). Unlimited public pushes. 1 free private repo. |
GHCRghcr.io |
Projects hosted on GitHub; seamless GitHub Actions integration | GitHub PAT or GITHUB_TOKEN |
Free for public repos. Private images tied to org/repo permissions. Best choice if your CI is GitHub Actions, with no separate secrets for auth. |
AWS ECR<account>.dkr.ecr.<region>.amazonaws.com |
Workloads running on ECS, EKS, Lambda | IAM roles / aws ecr get-login-password |
No pull-rate limits within AWS. Pay per GB stored + transferred out. Scan on push via Inspector. Short-lived tokens (12h). |
GCP Artifact Registry<region>-docker.pkg.dev |
Workloads on GKE / Cloud Run | Workload Identity / service account JSON | Replaced GCR. Supports multi-format repos (Docker, npm, Maven). IAM-integrated. Free first 0.5 GB/month. |
Azure ACR<registry>.azurecr.io |
Workloads on AKS / Azure Container Apps | Service principal / managed identity | Geo-replication on Premium tier. Built-in vulnerability scanning (Defender for Containers). Task runner for building inside Azure. |
Self-hostedregistry:2 / Harbor |
Air-gapped environments, on-prem, full control over retention | Basic auth or LDAP (Harbor) | registry:2 is the official minimal CNCF distribution: one container, no UI. Harbor adds RBAC, scanning, replication, a UI. Both speak OCI. |
Start on Docker Hub or GHCR while learning, as Sam does here. When you join a company using AWS, use ECR. It integrates with IAM and eliminates the need to manage separate credentials. On GCP, use Artifact Registry. Self-hosting with Harbor makes sense when your organisation has compliance requirements that forbid data leaving its own infrastructure, or when you need fine-grained RBAC and image lifecycle policies in one place.
Image naming & tags
Before you can push, you need to understand the full image reference format. Every image reference has the same anatomy:
| Part | What it is | When omitted |
|---|---|---|
| Registry host | The hostname (and optional port) of the registry server. ghcr.io, docker.io, 123456789.dkr.ecr.us-east-1.amazonaws.com | Defaults to docker.io (Docker Hub) |
| Namespace | Your Docker Hub username, GitHub username, or organisation name. On ECR it's the AWS account + region prefix. | On Docker Hub defaults to library (official images like nginx) |
| Repository | The name of the image, typically the app name: app, api, worker | Required; there is no default |
| Tag | A human-readable pointer to a specific image manifest. Mutable, so the registry can move it to a different manifest at any time. | Defaults to :latest |
The :latest trap: a callback to Module 01
In Module 01 you were told never to rely on :latest for anything important.
Here is the full reason: :latest is just a mutable tag. It points to
whatever manifest was pushed most recently. Push a new version and :latest now means
something different. Two machines that pull :latest at different times can get completely
different images with no warning.
A Kubernetes deployment using image: myapp:latest will silently run a different version
than the one you tested the moment anyone pushes a new build. In the best case this causes confusing behaviour; in
the worst case it silently ships a regression to production. Pin a specific version tag, or better, a digest
(Section 5).
docker pull nginx expands to docker.io/library/nginx:latest.
The library namespace is Docker's official image program. When you push your own image,
you always need your own namespace: docker.io/yourname/app:1.0.
Login, tag, push, pull
Here is the moment Sam has been building toward. Four commands cover the entire publish flow. They are all together, in order, for a real GHCR example. Read the whole block first, then we break down each step.
# 1. Authenticate: use a Personal Access Token, not your password
echo "<YOUR_GITHUB_PAT>" | docker login ghcr.io -u <github-username> --password-stdin
# 2. Build your image locally (you already know this from Module 02)
docker build -t myapp:1.0 .
# 3. Tag it with the full registry reference
docker tag myapp:1.0 ghcr.io/<github-username>/myapp:1.0
# 4. Push: only layers the registry doesn't already have are uploaded
docker push ghcr.io/<github-username>/myapp:1.0
# --- on another machine / in CI / in production ---
# 5. Pull the image
docker pull ghcr.io/<github-username>/myapp:1.0
# 6. Run it
docker run -p 8000:8000 ghcr.io/<github-username>/myapp:1.0
Step-by-step explanation
docker login writes an auth token to ~/.docker/config.json. The daemon
attaches this token to every push/pull request to that registry host. Always pipe the token via
--password-stdin rather than passing it as a CLI argument. Otherwise it appears in your
shell history and in ps output on shared machines. Use a Personal Access Token
(PAT) or a CI-managed secret, never your account password.
docker tag does not copy or move anything. It creates a second name for the exact same image manifest that already exists locally. The image has one set of layers; it now has two names. Think of it as an alias.
docker push contacts the registry, walks each layer's digest, and asks "do you already have this?" Only missing layers are uploaded. Sam watches the output scroll by and sees something like:
The push refers to repository [ghcr.io/user/myapp]
a3f42e8c1b20: Pushed ← new layer, uploaded
7d3c1f9a2044: Layer already exists ← base layer already in registry
e9c3d8b71a11: Layer already exists ← shared with another image
1.0: digest: sha256:c8f2... size: 1234
docker pull is the reverse: the daemon downloads the manifest, checks which layers it already has locally, and fetches only the missing ones. This is why pulling a second image that shares a base is fast: the base layers are already there.
If your image is private, every machine that needs to pull it must be authenticated. On a developer laptop that
means docker login. In Kubernetes that means an imagePullSecret
(a Secret containing registry credentials that you attach to a Pod spec). This is a forward reference to
Kubernetes, so keep it in mind.
Tags vs digests: the reproducibility difference
This is a concept worth mastering. Tags and digests both refer to an image, but they have fundamentally different guarantees.
Tags: mutable pointers
A tag is a string label that a registry associates with a manifest. Anyone with push access can move it to a
different manifest at any time. :1.0, :latest, :stable, all mutable.
Digests: immutable content addresses
A digest is the sha256 hash of the manifest JSON. It can never change, because the content is
the address. Pull the same digest on any machine, anywhere, at any time: you get the identical image.
To pull or reference an image by digest:
# pull by digest: immune to tag mutation
docker pull ghcr.io/user/myapp@sha256:c8f2a91b3d4e7f6082c1a9b5d8e3f2a7b4c6d9e1f3a5b7c9d2e4f6a8b0c1d3e5
# inspect an image's digest locally
docker inspect --format '{{index .RepoDigests 0}}' ghcr.io/user/myapp:1.0
# see the digest returned after a push (last line of push output)
# 1.0: digest: sha256:c8f2... size: 1234
Tags are mutable; they can be moved to a different image at any time. Digests are immutable, since they are
the sha256 hash of the manifest itself, so they uniquely and permanently identify one
exact image. For reproducible, secure production deployments, pin the digest. Your CI pipeline should
record the digest from the push output and use it in the deployment manifest, not the tag.
In a Kubernetes deployment manifest, a digest pin looks like this:
# tag alone: mutable, not recommended for prod
image: ghcr.io/user/myapp:1.2.0
# tag + digest: best of both, human-readable AND immutable
image: ghcr.io/user/myapp:1.2.0@sha256:c8f2a91b3d4e7f...
Tagging strategy
A consistent tagging strategy is something engineering teams set up once and rely on forever. This is what professional CI pipelines look like.
| Tag pattern | Example | When to use it |
|---|---|---|
| Semantic version | 1.2.0 | Every release. Stable, human-readable, allows rollback by version name. |
| Git commit SHA | a3f42e8 | Every CI build. Links the image directly to the exact commit that produced it. Most useful for debugging. |
| Branch name | main, feature-login | Mutable pointer to the latest build on that branch. Good for staging environments. |
latest | latest | Acceptable as a convenience alias for the most recent stable build. Never as the only tag, never for prod deploys. |
| Date / build number | 2026-06-01, build-1234 | Some teams use these for auditability. Less common than git SHA. |
A solid CI pipeline applies multiple tags to the same image on every build. That means a single
push produces: a git SHA tag (traceable), a semver tag (releasable), and optionally moves latest.
The layers are stored once; multiple tag labels just point to the same manifest. There is no storage penalty.
# CI environment variables (GitHub Actions example)
SHA=$(git rev-parse --short HEAD) # e.g. a3f42e8
VERSION=1.2.0 # from package.json / pyproject.toml / tag
REGISTRY=ghcr.io/myorg/myapp
# build once
docker build -t $REGISTRY:$SHA .
# apply additional tags (no re-build, just aliases)
docker tag $REGISTRY:$SHA $REGISTRY:$VERSION
docker tag $REGISTRY:$SHA $REGISTRY:latest
# push all three: same layers, three manifest pointers
docker push $REGISTRY:$SHA
docker push $REGISTRY:$VERSION
docker push $REGISTRY:latest
The git SHA creates a direct, unambiguous link between the running container and the source code that produced it. When a bug appears in production, you can read the SHA from the container's image reference, check it out locally, and reproduce the exact build. Without a SHA tag you have to guess which commit is in the running image. SHA tagging is the single most useful debugging practice in containerised CI/CD.
Scanning & signing: intro to supply-chain safety
Pushing an image to a registry makes it available to anyone (or any system) that pulls it. Before you do that in a production context, you want to know two things: does this image contain known vulnerabilities? and can the recipient trust that this image came from you and hasn't been tampered with? That's vulnerability scanning and image signing, respectively.
Vulnerability scanning
Every layer in your image contains OS packages, language runtimes, and libraries. Any of those can have published CVEs. A scanner compares the image's bill of materials (SBOM) against a vulnerability database and reports what it finds.
# quick overview: shows CVE count by severity
docker scout quickview ghcr.io/user/myapp:1.0
# full CVE list
docker scout cves ghcr.io/user/myapp:1.0
# compare two versions to see what changed
docker scout compare ghcr.io/user/myapp:1.0 --to ghcr.io/user/myapp:1.1
Trivy (by Aqua Security) is the most widely deployed open-source scanner. It runs as a single binary or container, scans local images or remote references, and integrates easily into CI pipelines:
# scan a local image
trivy image myapp:1.0
# fail CI if any CRITICAL CVEs are found
trivy image --exit-code 1 --severity CRITICAL myapp:1.0
Scan your image as part of the CI build, before it reaches the registry. Catching a critical CVE
after the image is in production is a fire drill. Catching it in CI is a five-second fix. Most teams gate the push
step: trivy image --exit-code 1 --severity CRITICAL myapp:$SHA && docker push ....
Image signing and provenance (intro)
Signing answers: "is this image actually what it claims to be, and did it come from the pipeline I trust?" The CNCF tool cosign (part of the Sigstore project) attaches a cryptographic signature to an image's digest in the registry. A deploy pipeline can verify the signature before running a container. If the signature is missing or invalid, the deploy is rejected.
Docker also ships an older, built-in mechanism called Docker Content Trust (DCT), backed by The Update
Framework / Notary. Setting DOCKER_CONTENT_TRUST=1 makes the client sign tags on push and refuse to
pull any unsigned tag. It predates Sigstore and is largely superseded by cosign in modern pipelines
(cosign is keyless, registry-agnostic, and signs the immutable digest rather than the mutable tag), but you will still see
DCT referenced in legacy setups.
Attestations go further: they attach machine-readable provenance metadata (which git commit, which
CI run, which Dockerfile) to the image in the registry. GitHub Actions can generate these automatically with
actions/attest-build-provenance.
Supply-chain security (cosign, Sigstore, SBOM generation, policy enforcement with OPA/Kyverno) is covered in depth in Module 07, Security & Hardening. For now, the key takeaways are: always scan before pushing, and know that signing/attestations exist as the next level of trust beyond scanning.
Private registries & CI
Pushing from GitHub Actions
GitHub Actions has a built-in GITHUB_TOKEN that is scoped to the repository and has
write access to GHCR for that repo. You don't need to create a separate secret; just log in with it. Sam's next step is wiring this into CI so the push happens on every commit:
- name: Log in to GHCR
uses: docker/login-action@v3
with:
registry: ghcr.io
username: ${{ github.actor }}
password: ${{ secrets.GITHUB_TOKEN }}
- name: Build and push
uses: docker/build-push-action@v5
with:
context: .
push: true
tags: |
ghcr.io/${{ github.repository_owner }}/myapp:${{ github.sha }}
ghcr.io/${{ github.repository_owner }}/myapp:latest
For other registries (ECR, Docker Hub) you store the credentials as encrypted repository secrets
(secrets.DOCKERHUB_TOKEN, etc.) and reference them the same way. Never hard-code tokens in
the workflow file itself.
Pulling in production with imagePullSecrets
When Kubernetes pulls a private image, it needs credentials. You create a Kubernetes Secret of type
kubernetes.io/dockerconfigjson containing the registry auth, then reference it in your Pod
spec under imagePullSecrets. This is a forward reference to the Kubernetes phase, so just know
that the registry credential problem you solved locally with docker login has a direct
equivalent in every orchestrator.
Self-hosting with registry:2
Running your own registry is genuinely one command. The official registry:2 image is the
CNCF Distribution reference implementation: no UI, no RBAC, just the OCI-compliant HTTP API. Useful for
air-gapped labs or local testing:
# start a local registry on port 5000, persist data in a volume
docker run -d -p 5000:5000 --name registry -v registry-data:/var/lib/registry registry:2
# tag and push to it
docker tag myapp:1.0 localhost:5000/myapp:1.0
docker push localhost:5000/myapp:1.0
# pull from it on the same machine
docker pull localhost:5000/myapp:1.0
registry:2 has no UI, no user management, no vulnerability scanning, and no replication.
For a production on-prem registry that needs role-based access, a web UI, integrated Trivy scanning, and
cross-datacenter replication, use Harbor (CNCF graduated project). It wraps
registry:2 internally but adds all of the above.
Hands-on: do this now
Walk Sam's path yourself. The only way to understand push/pull is to do it and watch only the changed layers travel. This checklist should take 30 to 45 minutes.
- Create a free Docker Hub account at hub.docker.com, or enable GHCR on any GitHub repo (it's free for public repos).
- Run
docker login docker.io(ordocker login ghcr.io) and verify you get "Login Succeeded". - Take any local image you built in Modules 02–05. Tag it:
docker tag myapp:1.0 <yournamespace>/myapp:1.0. - Push it:
docker push <yournamespace>/myapp:1.0. Watch the output and note which layers say "Pushed" vs "Layer already exists". - Delete the local image:
docker rmi <yournamespace>/myapp:1.0. Then pull it back:docker pull <yournamespace>/myapp:1.0. Confirm it runs. - Make a trivial one-line change to the app (e.g., edit a string), rebuild, retag as
:1.1, and push again. Observe that only the changed layers upload; the base layers say "Layer already exists". - Find the digest: run
docker inspect --format '{{index .RepoDigests 0}}' <image>. Copy thesha256:...and pull by that digest directly. - Run
docker scout quickview <yourimage>(ortrivy image <yourimage>if Trivy is installed). Read the CVE summary. Note that even a minimal image may have findings. - Optional: start a local
registry:2container, push and pull tolocalhost:5000to see that the same commands work against a self-hosted registry.
The most important moment in this hands-on is watching the second push, where only one or two layers upload instead of all of them. This is the moment Sam's app stops being trapped on one laptop. That is the content-addressed deduplication that makes registries fast and storage efficient. The registry didn't need you to tell it what changed; it figured it out from the layer digests. If you understand why that works, you understand how registries work.