← Learn
Docker Course  /  Phase 1  /  Module 07
Day 7 ~2h · hands-on
Phase 1 · Docker · Day 7

Module 07 · Security & Hardening

Sam has built images, run containers, managed volumes, networks, Compose stacks, and pushed their first app to a registry. The demo works. Now comes the part that keeps Sam up at night: making all of that production-worthy. A container that works is not the same as a container that is safe to run in the world, and shipping the first one is where most developers learn that the hard way.

~2h · hands-on heavy builds on all prior modules non-root · capabilities · secrets · scanning Docker 29.5 · CachyOS
0

Where we left off

Module 00 stated a truth that everything else in this course traces back to: containers share the host kernel. They are not separate operating systems. The isolation comes from Linux namespaces (PID, NET, MNT, UTS, IPC, USER) and cgroups, both of which are features of the same kernel your host is already running. That single fact defines the entire security picture for Docker.

Module 02 touched the surface: drop root with USER appuser, keep secrets out of the context with .dockerignore. Those two steps were not just best-practice flavour. They were the beginning of a full hardening story, and this module finishes it. Sam already did the easy parts without quite understanding why; here we make the why explicit and add everything else.

Shared kernel = shared risk

A vulnerability in the Linux kernel is a vulnerability in every container on that host, simultaneously.

Root in container ≈ root on host

Module 02 dropped root. This module explains why that matters at a kernel level and adds the rest of the hardening stack.

Defence in depth

No single control is enough. Hardening is layering: non-root + capabilities + seccomp + read-only fs + scanning + secrets hygiene.

Scope of this module

This module covers hardening for containers you build and run yourself. For multi-tenant untrusted workloads (running other people's arbitrary code) a VM boundary or a kernel-level sandboxed runtime like gVisor or Kata Containers is the correct tool. Docker's namespace isolation alone is not considered a hard security boundary for that threat model.

1

The threat model (what you are actually defending against)

Before layering controls, Sam needs to understand what failure looks like. Otherwise hardening is just cargo-culting flags off a blog post. There are four categories of concern with Docker containers:

ThreatWhat it meansPrimary control
Container escapeA compromised process inside the container breaks out to the host filesystem, host network, or host process table. The shared kernel is the path.Non-root, drop caps, no --privileged, seccomp
Malicious / vulnerable imageYou pull an image that contains malware, or a base image with known CVEs that an attacker can exploit once the container is running.Image scanning, digest pinning, minimal base
Leaked secretsCredentials baked into an image layer, passed via ENV at build time, or visible in docker history. Anyone who pulls the image has your secrets.Secrets management, never bake secrets
Runaway / DoSA single container consumes all CPU, memory, or PIDs on the host, starving other workloads or the host itself.Resource limits (cgroups)
The central insight

Container isolation is weaker than a VM because a VM has a hypervisor between guest and host, so there is no shared kernel. In a container, a kernel exploit crosses the container boundary immediately. This is not a theoretical concern: real container escapes (runc CVE-2019-5736, for example) have exploited this exact path. Defence-in-depth reduces the blast radius when something goes wrong.

2

Run as non-root (the #1 finding)

By default, the process inside a Docker container runs as root (UID 0). That root is not fully the same as host root (namespaces constrain it) but the gap is much narrower than you want it to be. If the container process escapes, it arrives on the host as root. Root can read /etc/shadow, write to /etc/cron.d, and install a kernel module. Every security audit flags this first. When Sam ran the first scan, this was line one of the report.

The fix is a three-line addition to the Dockerfile, which Module 02 introduced. Here is the full pattern with the reasoning visible:

Dockerfile : non-root user pattern
FROM python:3.12-slim

WORKDIR /app

# install deps as root (needs package manager), then create an unprivileged user
RUN apt-get update && apt-get install -y --no-install-recommends libpq5 \
    && rm -rf /var/lib/apt/lists/* \
    && useradd --uid 1001 --create-home appuser

COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt

COPY --chown=appuser:appuser . .

# switch to the unprivileged user for everything that follows
USER appuser

CMD ["python", "app.py"]

You can also override the user at runtime without touching the Dockerfile:

terminal : override user at runtime
docker run --user 1001:1001 myimage
Beware UID 0 in base images

Many official base images run as root by default and do not create an appuser for you. Always check with docker inspect <image> | grep -i user and add the USER instruction yourself if the result is empty or root.

3

Drop Linux capabilities (fine-grained root powers)

Linux does not treat root as a binary on/off switch. Root's power is split into roughly 40 distinct capabilities, and each one grants a specific kernel-level privilege. Examples: CAP_NET_BIND_SERVICE (bind ports below 1024), CAP_SYS_ADMIN (mount filesystems, change namespaces, essentially a root backdoor), CAP_NET_RAW (craft raw packets, useful for attacks).

By default, Docker grants containers a subset of capabilities: not all of them, but more than most apps need. The hardening principle is simple. Drop everything, then add back only what your app actually requires. Sam's app is a plain web service, so the honest answer to "what does it need" turns out to be almost nothing.

terminal : drop all capabilities, add back only what is needed
# most web apps need nothing beyond basic networking
docker run --cap-drop=ALL myimage

# a web server that binds port 80 (below 1024) needs one specific capability
docker run --cap-drop=ALL --cap-add=NET_BIND_SERVICE nginx:alpine

You can inspect what capabilities a running container currently has:

terminal : check capabilities of a running container
# inside the container
docker exec <container> cat /proc/1/status | grep Cap

# decode a hex capability mask (install libcap2-bin if needed)
capsh --decode=00000000a80425fb
CAP_SYS_ADMIN is almost --privileged

CAP_SYS_ADMIN is so broad it effectively grants most of what --privileged grants. Never add it unless you have a very specific, documented reason. If you catch yourself adding it "because something needs permissions," find out what exactly it needs and add only the targeted capability.

4

Read-only filesystem & no-new-privileges

Two flags that together close common post-exploitation paths:

--read-only

Mounts the container's root filesystem as read-only. A compromised process cannot write malware, modify configuration files, or persist state to the container's filesystem. Most well-written apps do not need to write to the root filesystem; they write to specific directories (logs, temp, sockets). Mount those explicitly with --tmpfs for in-memory scratch space, or a named volume for persistent data.

terminal : read-only root fs with tmpfs for writable scratch
docker run \
  --read-only \
  --tmpfs /tmp:rw,noexec,nosuid,size=64m \
  --tmpfs /var/run:rw,noexec,nosuid,size=4m \
  myimage

--security-opt=no-new-privileges

Prevents the container process, and any child process it spawns, from gaining new privileges via setuid/setgid binaries or Linux capabilities bounding set manipulation. Without this flag, a non-root process that executes a setuid binary (like sudo) can escalate. This flag simply locks the privilege ceiling at the level the container started with.

terminal : combining read-only + no-new-privileges
docker run \
  --read-only \
  --tmpfs /tmp \
  --security-opt=no-new-privileges \
  --user 1001:1001 \
  myimage
5

Seccomp & AppArmor / SELinux

Even without capabilities, a process can talk directly to the Linux kernel via system calls (syscalls). There are over 400 syscalls. Most processes use fewer than 50 of them. Seccomp (Secure Computing Mode) filters which syscalls a process is allowed to make. Any call not on the allowlist kills the process immediately.

The default seccomp profile

Docker ships with a default seccomp profile that blocks approximately 44 dangerous syscalls including ptrace (attach to other processes), keyctl (manipulate kernel keyring), add_key, mount, reboot, and others that are almost never needed by applications but are commonly useful in exploits. This profile is active by default, so you get it for free unless you remove it.

terminal : using a custom seccomp profile
# apply a custom profile (JSON file you write defining allowed syscalls)
docker run --security-opt seccomp=/path/to/profile.json myimage

# disable seccomp entirely (DO NOT do this in production)
docker run --security-opt seccomp=unconfined myimage

AppArmor and SELinux

These are Mandatory Access Control (MAC) systems that work at the kernel level, separate from and complementary to capabilities and seccomp. AppArmor (default on Ubuntu/Debian-based hosts, including many CachyOS configurations) attaches profiles to executables, constraining what files, network operations, and capabilities the program can access, regardless of what the process itself requests. Docker applies a default AppArmor profile (docker-default) to all containers on supported hosts.

SELinux plays the same role on RHEL/Fedora/CentOS hosts. If your host has SELinux enforcing mode, Docker labels containers and volumes with appropriate SELinux contexts automatically.

Never disable the default profiles casually

Running with --security-opt seccomp=unconfined or --security-opt apparmor=unconfined removes these protections silently. These flags appear in many "quick fix" Stack Overflow answers when something breaks. Do not use them in production. Investigate which specific syscall or access is being blocked and allow only that. Do not disable the whole layer.

6

Never run --privileged

The --privileged flag does the following in a single pass: grants all Linux capabilities, disables the default seccomp profile, disables AppArmor/SELinux confinement, gives the container access to all host devices, and allows the container process to manipulate namespaces directly. It essentially removes the entire isolation stack Docker provides.

A process running inside a privileged container can mount the host filesystem, load kernel modules, modify iptables rules, and escape to the host in seconds using well-documented techniques. This is not a theoretical risk. It is a documented, trivially-reproduced attack.

--privileged = container escape waiting to happen

If you have --privileged in a production container, you do not have container isolation. You have a process with full host access running inside a container-shaped box. Remove it. If something requires it, that something should run on bare metal or a dedicated VM, not in a shared container environment.

The only legitimate uses of --privileged in practice are: Docker-in-Docker (DinD) build environments in controlled CI where the host is single-purpose, and certain kernel debugging tools. Even in those cases, prefer alternative approaches: Docker socket mounting with explicit access controls, or using --device to grant access to a specific device rather than all of them.

terminal : the safer alternative to --privileged for device access
# instead of --privileged to access a GPU or device, be specific
docker run --device /dev/nvidia0:/dev/nvidia0 myimage

# instead of --privileged for a specific capability only
docker run --cap-add=SYS_PTRACE myimage
7

User namespace remapping & rootless Docker

Module 00 introduced the USER namespace as one of the six Linux namespaces Docker uses. This is where the most powerful isolation improvement lives, and it is often skipped in basic guides.

User namespace remapping (userns-remap)

By default, UID 0 inside the container maps directly to UID 0 on the host, and that is the problem. With userns-remap, Docker remaps container UIDs to a range of unprivileged host UIDs. Container root (UID 0) becomes, for example, host UID 100000. If the container process escapes, it arrives on the host as UID 100000, a non-privileged user that owns nothing important.

/etc/docker/daemon.json : enable userns-remap
{
  "userns-remap": "default"
}

After setting this and restarting the Docker daemon, Docker creates a dockremap user and group and uses the subordinate UID/GID ranges in /etc/subuid and /etc/subgid for the mapping.

Rootless Docker

Rootless mode goes further: the Docker daemon itself runs as a non-root user, not just the containers. The daemon process no longer has host root privileges. This significantly reduces the impact of a daemon-level vulnerability. It uses the same USER namespace mechanism under the hood, but the daemon itself is unprivileged.

terminal : install and use rootless Docker
# install the rootless setup script (requires uidmap package)
dockerd-rootless-setuptool.sh install

# start the rootless daemon for the current user
systemctl --user start docker

# point your CLI at the rootless socket
export DOCKER_HOST=unix://$XDG_RUNTIME_DIR/docker.sock

# confirm the daemon is running as your unprivileged user
docker info | grep "rootless"
8

Secrets (the mistake that doesn't go away)

This is the category that nearly bit Sam. Early on, Sam dropped a .env into the build and only later wondered who could read it. The mechanism is subtle enough to trip up experienced engineers: image layers are permanent and additive. If you add a secret to a layer and delete it in a later layer, the secret still exists in the earlier layer. Anyone with access to the image can extract it with docker history or by exporting the image tarball.

Deleting a secret in a later layer does NOT remove it

This is the most common secrets mistake. The following pattern looks like it cleans up after itself. It does not:

COPY secrets.env .          # baked into layer N
RUN ./setup.sh              # uses it
RUN rm secrets.env          # removes it from layer N+2, but layer N still has it

Anyone who runs docker history --no-trunc myimage or extracts the layer tarball reads the secret in plain text.

What to do instead

ApproachWhen to useHow
Build-time secrets (BuildKit)Secrets needed only during docker build (e.g. private pip index token)RUN --mount=type=secret,id=mysecret ..., never written to any layer
Runtime env from a secrets managerSecrets the running app needs (DB password, API key)Inject at runtime from Vault, AWS Secrets Manager, etc., not baked in
Docker secrets (Swarm)Swarm or Compose production deploymentsMounted at /run/secrets/<name> inside the container, never in the image
Mounted file at runtimeSimple cases, dev/stagingdocker run -v /host/secrets:/run/secrets:ro myimage
ENV and ARG are not secret storage

ARG MY_SECRET=... in a Dockerfile bakes the value into the image metadata. ENV MY_SECRET=... is visible in docker inspect. Neither is safe for credentials. Use build-time secret mounts or runtime injection exclusively for sensitive values.

9

Image hygiene & scanning

Every package in your image is a potential CVE waiting to be exploited. The attack surface of an image is directly proportional to its size. This is why Module 02's advice to use python:3.12-slim instead of python:3.12 was not just about download speed. It was about security posture. Fewer packages means fewer vulnerabilities.

Minimal and distroless bases

The progression from most to least attack surface: full OS image → slim image → Alpine → distroless. Google's distroless images contain only the language runtime and its direct dependencies: no shell, no package manager, no coreutils. A compromised process in a distroless container has no bash to spawn, no curl to exfiltrate data with, no apt to install tools. It is significantly harder to move laterally.

Pinning by digest (from Module 06)

Tagging with :latest or even :3.12-slim means a future pull might give you a different, potentially vulnerable image. Pin by immutable digest:

Dockerfile : pin base image by digest
FROM python:3.12-slim@sha256:3f4e8c8b...  # immutable, always this exact image

Scanning with Docker Scout and Trivy

Image scanning compares every package in your image against a CVE database and reports known vulnerabilities by severity (Critical, High, Medium, Low). This should run in CI on every build, not manually on your laptop occasionally. When Sam wired this into the pipeline, the first run flagged a base-image CVE that nobody would have caught by eye.

terminal : scanning with Docker Scout and Trivy
# Docker Scout (built into Docker CLI, requires login)
docker scout cves myimage:latest

# Docker Scout quickview summary
docker scout quickview myimage:latest

# Trivy (open-source, standalone, great for CI)
trivy image myimage:latest

# Trivy: fail CI if any Critical or High CVEs found
trivy image --exit-code 1 --severity CRITICAL,HIGH myimage:latest
Scanning is not one-and-done

New CVEs are published daily. An image that scanned clean last week may have a new critical vulnerability today. Keep your base images updated, automate scanning in CI, and set up notifications (Docker Scout can alert you) when new CVEs appear in images already in your registry.

10

Resource limits (cgroups in practice)

Module 00 established that cgroups are the kernel mechanism that enforces resource limits on container processes. Without limits, a single misbehaving or compromised container can exhaust the host's memory (causing an OOM kill of other containers), saturate all CPU cores, or fork-bomb the system by spawning unlimited processes. Setting limits is both a security and a reliability control.

terminal : setting memory, CPU, and PID limits
# limit to 256 MB RAM, 0.5 CPU, and 100 processes max
docker run \
  --memory=256m \
  --memory-swap=256m \
  --cpus=0.5 \
  --pids-limit=100 \
  myimage

# check what limits are applied to a running container
docker inspect <container> | grep -A5 "Memory"

Note that --memory-swap set equal to --memory disables swap for the container, which prevents a container from silently using disk as memory and masking a memory leak.

Bridge to Kubernetes

In Kubernetes, these same cgroup controls are expressed as resources.requests and resources.limits in a Pod spec. The concept is identical; Kubernetes simply manages the cgroup configuration for you per container declaration. Understanding --memory and --cpus here makes Kubernetes resource management immediately intuitive.

11

The hardening checklist

Everything covered in this module, as a deployable checklist. Use this before every production deployment.

ControlHow to applyPriority
Run as non-rootUSER appuser in Dockerfile; or --user 1001:1001 at runtimeCritical
Drop all capabilities--cap-drop=ALL, then add back only what's needed with --cap-addCritical
No --privilegedNever in production. Use --device or targeted --cap-add insteadCritical
Secrets never in imageBuildKit secret mounts at build time; runtime injection or Docker secrets at run timeCritical
Read-only root filesystem--read-only with --tmpfs for writable pathsHigh
no-new-privileges--security-opt=no-new-privilegesHigh
Keep default seccomp profileDo not pass --security-opt seccomp=unconfinedHigh
Minimal base image-slim, Alpine, or distroless; remove dev tools from final imageHigh
Pin image digestsFROM image@sha256:... in production DockerfilesHigh
Scan images in CIdocker scout cves or trivy image on every build; fail on Critical/HighHigh
Resource limits--memory, --cpus, --pids-limit on every production containerMedium
User namespace remappinguserns-remap: default in daemon.json, or use rootless DockerMedium
AppArmor / SELinuxLeave default profiles enabled; write custom profiles for high-sensitivity workloadsMedium
12

Hands-on (do this now)

Each of these exercises will either show you a protection working or show you a protection being violated. Both outcomes are learning. This is exactly the loop Sam ran before flipping the app to production, and do not skip it.

  • Build a simple image with a USER appuser instruction. Run it and verify with docker exec <container> whoami that you are not root inside.
  • Run an image with --cap-drop=ALL. Try an operation that requires a capability. For example, run ping (needs CAP_NET_RAW) and watch it fail with a permission error. Then add --cap-add=NET_RAW and watch it succeed.
  • Run a container with --read-only --tmpfs /tmp. Attempt to write a file to /app (should fail) then to /tmp (should succeed). Confirm the filesystem is genuinely immutable.
  • Build an image that (wrongly) copies a fake credentials file into a layer and then deletes it in a later RUN rm. Run docker history --no-trunc <image> and confirm you cannot read it from history directly, then export the image with docker save | tar -xvf and look inside the layer tarballs to see the file still there.
  • Scan an image you have built with docker scout cves <image> or trivy image <image>. Read the output. Find the CVE IDs, understand what package they are in, and look up one CVE on the NVD (nvd.nist.gov). This is not optional. It is a skill you must actually have.
  • Run a container with --memory=64m --memory-swap=64m. Write a small program (or use a shell loop) that allocates memory until it runs out. Watch the OOM kill message from the kernel. Then run docker inspect <container> | grep OOMKilled and see the flag set to true.
  • Run a container that sets --security-opt=no-new-privileges. Inside the container, attempt to use sudo or execute a setuid binary (e.g. passwd). Confirm it fails to escalate privileges.
Write commands from memory

After the first read-through, close this page and try to reproduce each docker run command from memory. If you cannot write --cap-drop=ALL --cap-add=NET_BIND_SERVICE without looking, you do not own it yet. The hands-on is not a checklist to tick. It is a repetition exercise.