Module 07 · Security & Hardening
Sam has built images, run containers, managed volumes, networks, Compose stacks, and pushed their first app to a registry. The demo works. Now comes the part that keeps Sam up at night: making all of that production-worthy. A container that works is not the same as a container that is safe to run in the world, and shipping the first one is where most developers learn that the hard way.
Where we left off
Module 00 stated a truth that everything else in this course traces back to: containers share the host kernel. They are not separate operating systems. The isolation comes from Linux namespaces (PID, NET, MNT, UTS, IPC, USER) and cgroups, both of which are features of the same kernel your host is already running. That single fact defines the entire security picture for Docker.
Module 02 touched the surface: drop root with USER appuser, keep secrets out of the context with .dockerignore. Those two steps were not just best-practice flavour. They were the beginning of a full hardening story, and this module finishes it. Sam already did the easy parts without quite understanding why; here we make the why explicit and add everything else.
Shared kernel = shared risk
A vulnerability in the Linux kernel is a vulnerability in every container on that host, simultaneously.
Root in container ≈ root on host
Module 02 dropped root. This module explains why that matters at a kernel level and adds the rest of the hardening stack.
Defence in depth
No single control is enough. Hardening is layering: non-root + capabilities + seccomp + read-only fs + scanning + secrets hygiene.
This module covers hardening for containers you build and run yourself. For multi-tenant untrusted workloads (running other people's arbitrary code) a VM boundary or a kernel-level sandboxed runtime like gVisor or Kata Containers is the correct tool. Docker's namespace isolation alone is not considered a hard security boundary for that threat model.
The threat model (what you are actually defending against)
Before layering controls, Sam needs to understand what failure looks like. Otherwise hardening is just cargo-culting flags off a blog post. There are four categories of concern with Docker containers:
| Threat | What it means | Primary control |
|---|---|---|
| Container escape | A compromised process inside the container breaks out to the host filesystem, host network, or host process table. The shared kernel is the path. | Non-root, drop caps, no --privileged, seccomp |
| Malicious / vulnerable image | You pull an image that contains malware, or a base image with known CVEs that an attacker can exploit once the container is running. | Image scanning, digest pinning, minimal base |
| Leaked secrets | Credentials baked into an image layer, passed via ENV at build time, or visible in docker history. Anyone who pulls the image has your secrets. | Secrets management, never bake secrets |
| Runaway / DoS | A single container consumes all CPU, memory, or PIDs on the host, starving other workloads or the host itself. | Resource limits (cgroups) |
Container isolation is weaker than a VM because a VM has a hypervisor between guest and host, so there is no shared kernel. In a container, a kernel exploit crosses the container boundary immediately. This is not a theoretical concern: real container escapes (runc CVE-2019-5736, for example) have exploited this exact path. Defence-in-depth reduces the blast radius when something goes wrong.
Run as non-root (the #1 finding)
By default, the process inside a Docker container runs as root (UID 0). That root is not fully the same as host root (namespaces constrain it) but the gap is much narrower than you want it to be. If the container process escapes, it arrives on the host as root. Root can read /etc/shadow, write to /etc/cron.d, and install a kernel module. Every security audit flags this first. When Sam ran the first scan, this was line one of the report.
The fix is a three-line addition to the Dockerfile, which Module 02 introduced. Here is the full pattern with the reasoning visible:
FROM python:3.12-slim
WORKDIR /app
# install deps as root (needs package manager), then create an unprivileged user
RUN apt-get update && apt-get install -y --no-install-recommends libpq5 \
&& rm -rf /var/lib/apt/lists/* \
&& useradd --uid 1001 --create-home appuser
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY --chown=appuser:appuser . .
# switch to the unprivileged user for everything that follows
USER appuser
CMD ["python", "app.py"]
You can also override the user at runtime without touching the Dockerfile:
docker run --user 1001:1001 myimage
Many official base images run as root by default and do not create an appuser for you. Always check with docker inspect <image> | grep -i user and add the USER instruction yourself if the result is empty or root.
Drop Linux capabilities (fine-grained root powers)
Linux does not treat root as a binary on/off switch. Root's power is split into roughly 40 distinct capabilities, and each one grants a specific kernel-level privilege. Examples: CAP_NET_BIND_SERVICE (bind ports below 1024), CAP_SYS_ADMIN (mount filesystems, change namespaces, essentially a root backdoor), CAP_NET_RAW (craft raw packets, useful for attacks).
By default, Docker grants containers a subset of capabilities: not all of them, but more than most apps need. The hardening principle is simple. Drop everything, then add back only what your app actually requires. Sam's app is a plain web service, so the honest answer to "what does it need" turns out to be almost nothing.
# most web apps need nothing beyond basic networking
docker run --cap-drop=ALL myimage
# a web server that binds port 80 (below 1024) needs one specific capability
docker run --cap-drop=ALL --cap-add=NET_BIND_SERVICE nginx:alpine
You can inspect what capabilities a running container currently has:
# inside the container
docker exec <container> cat /proc/1/status | grep Cap
# decode a hex capability mask (install libcap2-bin if needed)
capsh --decode=00000000a80425fb
CAP_SYS_ADMIN is so broad it effectively grants most of what --privileged grants. Never add it unless you have a very specific, documented reason. If you catch yourself adding it "because something needs permissions," find out what exactly it needs and add only the targeted capability.
Read-only filesystem & no-new-privileges
Two flags that together close common post-exploitation paths:
--read-only
Mounts the container's root filesystem as read-only. A compromised process cannot write malware, modify configuration files, or persist state to the container's filesystem. Most well-written apps do not need to write to the root filesystem; they write to specific directories (logs, temp, sockets). Mount those explicitly with --tmpfs for in-memory scratch space, or a named volume for persistent data.
docker run \
--read-only \
--tmpfs /tmp:rw,noexec,nosuid,size=64m \
--tmpfs /var/run:rw,noexec,nosuid,size=4m \
myimage
--security-opt=no-new-privileges
Prevents the container process, and any child process it spawns, from gaining new privileges via setuid/setgid binaries or Linux capabilities bounding set manipulation. Without this flag, a non-root process that executes a setuid binary (like sudo) can escalate. This flag simply locks the privilege ceiling at the level the container started with.
docker run \
--read-only \
--tmpfs /tmp \
--security-opt=no-new-privileges \
--user 1001:1001 \
myimage
Seccomp & AppArmor / SELinux
Even without capabilities, a process can talk directly to the Linux kernel via system calls (syscalls). There are over 400 syscalls. Most processes use fewer than 50 of them. Seccomp (Secure Computing Mode) filters which syscalls a process is allowed to make. Any call not on the allowlist kills the process immediately.
The default seccomp profile
Docker ships with a default seccomp profile that blocks approximately 44 dangerous syscalls including ptrace (attach to other processes), keyctl (manipulate kernel keyring), add_key, mount, reboot, and others that are almost never needed by applications but are commonly useful in exploits. This profile is active by default, so you get it for free unless you remove it.
# apply a custom profile (JSON file you write defining allowed syscalls)
docker run --security-opt seccomp=/path/to/profile.json myimage
# disable seccomp entirely (DO NOT do this in production)
docker run --security-opt seccomp=unconfined myimage
AppArmor and SELinux
These are Mandatory Access Control (MAC) systems that work at the kernel level, separate from and complementary to capabilities and seccomp. AppArmor (default on Ubuntu/Debian-based hosts, including many CachyOS configurations) attaches profiles to executables, constraining what files, network operations, and capabilities the program can access, regardless of what the process itself requests. Docker applies a default AppArmor profile (docker-default) to all containers on supported hosts.
SELinux plays the same role on RHEL/Fedora/CentOS hosts. If your host has SELinux enforcing mode, Docker labels containers and volumes with appropriate SELinux contexts automatically.
Running with --security-opt seccomp=unconfined or --security-opt apparmor=unconfined removes these protections silently. These flags appear in many "quick fix" Stack Overflow answers when something breaks. Do not use them in production. Investigate which specific syscall or access is being blocked and allow only that. Do not disable the whole layer.
Never run --privileged
The --privileged flag does the following in a single pass: grants all Linux capabilities, disables the default seccomp profile, disables AppArmor/SELinux confinement, gives the container access to all host devices, and allows the container process to manipulate namespaces directly. It essentially removes the entire isolation stack Docker provides.
A process running inside a privileged container can mount the host filesystem, load kernel modules, modify iptables rules, and escape to the host in seconds using well-documented techniques. This is not a theoretical risk. It is a documented, trivially-reproduced attack.
If you have --privileged in a production container, you do not have container isolation. You have a process with full host access running inside a container-shaped box. Remove it. If something requires it, that something should run on bare metal or a dedicated VM, not in a shared container environment.
The only legitimate uses of --privileged in practice are: Docker-in-Docker (DinD) build environments in controlled CI where the host is single-purpose, and certain kernel debugging tools. Even in those cases, prefer alternative approaches: Docker socket mounting with explicit access controls, or using --device to grant access to a specific device rather than all of them.
# instead of --privileged to access a GPU or device, be specific
docker run --device /dev/nvidia0:/dev/nvidia0 myimage
# instead of --privileged for a specific capability only
docker run --cap-add=SYS_PTRACE myimage
User namespace remapping & rootless Docker
Module 00 introduced the USER namespace as one of the six Linux namespaces Docker uses. This is where the most powerful isolation improvement lives, and it is often skipped in basic guides.
User namespace remapping (userns-remap)
By default, UID 0 inside the container maps directly to UID 0 on the host, and that is the problem. With userns-remap, Docker remaps container UIDs to a range of unprivileged host UIDs. Container root (UID 0) becomes, for example, host UID 100000. If the container process escapes, it arrives on the host as UID 100000, a non-privileged user that owns nothing important.
{
"userns-remap": "default"
}
After setting this and restarting the Docker daemon, Docker creates a dockremap user and group and uses the subordinate UID/GID ranges in /etc/subuid and /etc/subgid for the mapping.
Rootless Docker
Rootless mode goes further: the Docker daemon itself runs as a non-root user, not just the containers. The daemon process no longer has host root privileges. This significantly reduces the impact of a daemon-level vulnerability. It uses the same USER namespace mechanism under the hood, but the daemon itself is unprivileged.
# install the rootless setup script (requires uidmap package)
dockerd-rootless-setuptool.sh install
# start the rootless daemon for the current user
systemctl --user start docker
# point your CLI at the rootless socket
export DOCKER_HOST=unix://$XDG_RUNTIME_DIR/docker.sock
# confirm the daemon is running as your unprivileged user
docker info | grep "rootless"
Secrets (the mistake that doesn't go away)
This is the category that nearly bit Sam. Early on, Sam dropped a .env into the build and only later wondered who could read it. The mechanism is subtle enough to trip up experienced engineers: image layers are permanent and additive. If you add a secret to a layer and delete it in a later layer, the secret still exists in the earlier layer. Anyone with access to the image can extract it with docker history or by exporting the image tarball.
This is the most common secrets mistake. The following pattern looks like it cleans up after itself. It does not:
COPY secrets.env . # baked into layer N RUN ./setup.sh # uses it RUN rm secrets.env # removes it from layer N+2, but layer N still has it
Anyone who runs docker history --no-trunc myimage or extracts the layer tarball reads the secret in plain text.
What to do instead
| Approach | When to use | How |
|---|---|---|
| Build-time secrets (BuildKit) | Secrets needed only during docker build (e.g. private pip index token) | RUN --mount=type=secret,id=mysecret ..., never written to any layer |
| Runtime env from a secrets manager | Secrets the running app needs (DB password, API key) | Inject at runtime from Vault, AWS Secrets Manager, etc., not baked in |
| Docker secrets (Swarm) | Swarm or Compose production deployments | Mounted at /run/secrets/<name> inside the container, never in the image |
| Mounted file at runtime | Simple cases, dev/staging | docker run -v /host/secrets:/run/secrets:ro myimage |
ARG MY_SECRET=... in a Dockerfile bakes the value into the image metadata. ENV MY_SECRET=... is visible in docker inspect. Neither is safe for credentials. Use build-time secret mounts or runtime injection exclusively for sensitive values.
Image hygiene & scanning
Every package in your image is a potential CVE waiting to be exploited. The attack surface of an image is directly proportional to its size. This is why Module 02's advice to use python:3.12-slim instead of python:3.12 was not just about download speed. It was about security posture. Fewer packages means fewer vulnerabilities.
Minimal and distroless bases
The progression from most to least attack surface: full OS image → slim image → Alpine → distroless. Google's distroless images contain only the language runtime and its direct dependencies: no shell, no package manager, no coreutils. A compromised process in a distroless container has no bash to spawn, no curl to exfiltrate data with, no apt to install tools. It is significantly harder to move laterally.
Pinning by digest (from Module 06)
Tagging with :latest or even :3.12-slim means a future pull might give you a different, potentially vulnerable image. Pin by immutable digest:
FROM python:3.12-slim@sha256:3f4e8c8b... # immutable, always this exact image
Scanning with Docker Scout and Trivy
Image scanning compares every package in your image against a CVE database and reports known vulnerabilities by severity (Critical, High, Medium, Low). This should run in CI on every build, not manually on your laptop occasionally. When Sam wired this into the pipeline, the first run flagged a base-image CVE that nobody would have caught by eye.
# Docker Scout (built into Docker CLI, requires login)
docker scout cves myimage:latest
# Docker Scout quickview summary
docker scout quickview myimage:latest
# Trivy (open-source, standalone, great for CI)
trivy image myimage:latest
# Trivy: fail CI if any Critical or High CVEs found
trivy image --exit-code 1 --severity CRITICAL,HIGH myimage:latest
New CVEs are published daily. An image that scanned clean last week may have a new critical vulnerability today. Keep your base images updated, automate scanning in CI, and set up notifications (Docker Scout can alert you) when new CVEs appear in images already in your registry.
Resource limits (cgroups in practice)
Module 00 established that cgroups are the kernel mechanism that enforces resource limits on container processes. Without limits, a single misbehaving or compromised container can exhaust the host's memory (causing an OOM kill of other containers), saturate all CPU cores, or fork-bomb the system by spawning unlimited processes. Setting limits is both a security and a reliability control.
# limit to 256 MB RAM, 0.5 CPU, and 100 processes max
docker run \
--memory=256m \
--memory-swap=256m \
--cpus=0.5 \
--pids-limit=100 \
myimage
# check what limits are applied to a running container
docker inspect <container> | grep -A5 "Memory"
Note that --memory-swap set equal to --memory disables swap for the container, which prevents a container from silently using disk as memory and masking a memory leak.
In Kubernetes, these same cgroup controls are expressed as resources.requests and resources.limits in a Pod spec. The concept is identical; Kubernetes simply manages the cgroup configuration for you per container declaration. Understanding --memory and --cpus here makes Kubernetes resource management immediately intuitive.
The hardening checklist
Everything covered in this module, as a deployable checklist. Use this before every production deployment.
| Control | How to apply | Priority |
|---|---|---|
| Run as non-root | USER appuser in Dockerfile; or --user 1001:1001 at runtime | Critical |
| Drop all capabilities | --cap-drop=ALL, then add back only what's needed with --cap-add | Critical |
| No --privileged | Never in production. Use --device or targeted --cap-add instead | Critical |
| Secrets never in image | BuildKit secret mounts at build time; runtime injection or Docker secrets at run time | Critical |
| Read-only root filesystem | --read-only with --tmpfs for writable paths | High |
| no-new-privileges | --security-opt=no-new-privileges | High |
| Keep default seccomp profile | Do not pass --security-opt seccomp=unconfined | High |
| Minimal base image | -slim, Alpine, or distroless; remove dev tools from final image | High |
| Pin image digests | FROM image@sha256:... in production Dockerfiles | High |
| Scan images in CI | docker scout cves or trivy image on every build; fail on Critical/High | High |
| Resource limits | --memory, --cpus, --pids-limit on every production container | Medium |
| User namespace remapping | userns-remap: default in daemon.json, or use rootless Docker | Medium |
| AppArmor / SELinux | Leave default profiles enabled; write custom profiles for high-sensitivity workloads | Medium |
Hands-on (do this now)
Each of these exercises will either show you a protection working or show you a protection being violated. Both outcomes are learning. This is exactly the loop Sam ran before flipping the app to production, and do not skip it.
- Build a simple image with a
USER appuserinstruction. Run it and verify withdocker exec <container> whoamithat you are not root inside. - Run an image with
--cap-drop=ALL. Try an operation that requires a capability. For example, runping(needsCAP_NET_RAW) and watch it fail with a permission error. Then add--cap-add=NET_RAWand watch it succeed. - Run a container with
--read-only --tmpfs /tmp. Attempt to write a file to/app(should fail) then to/tmp(should succeed). Confirm the filesystem is genuinely immutable. - Build an image that (wrongly) copies a fake credentials file into a layer and then deletes it in a later
RUN rm. Rundocker history --no-trunc <image>and confirm you cannot read it from history directly, then export the image withdocker save | tar -xvfand look inside the layer tarballs to see the file still there. - Scan an image you have built with
docker scout cves <image>ortrivy image <image>. Read the output. Find the CVE IDs, understand what package they are in, and look up one CVE on the NVD (nvd.nist.gov). This is not optional. It is a skill you must actually have. - Run a container with
--memory=64m --memory-swap=64m. Write a small program (or use a shell loop) that allocates memory until it runs out. Watch the OOM kill message from the kernel. Then rundocker inspect <container> | grep OOMKilledand see the flag set totrue. - Run a container that sets
--security-opt=no-new-privileges. Inside the container, attempt to usesudoor execute a setuid binary (e.g.passwd). Confirm it fails to escalate privileges.
After the first read-through, close this page and try to reproduce each docker run command from memory. If you cannot write --cap-drop=ALL --cap-add=NET_BIND_SERVICE without looking, you do not own it yet. The hands-on is not a checklist to tick. It is a repetition exercise.