SDP Clouds
← All posts
Security·4 min read

Container Runtime Hardening: Rootless, Read-Only, and Fewer Capabilities

Scanning tells you what is inside the image. These four settings decide what it can do once it is running — and what breaks when you turn them on.


Scanning an image answers a question about its contents. It says nothing about what the process is permitted to do after it starts. Two images can contain identical, fully patched packages and still differ enormously in blast radius, because the difference is in the runtime flags.

A container escaping into its host almost never starts as a kernel exploit. It starts as a container that was running as root, with a writable filesystem, holding a capability nobody had a reason to give it.

Four settings close most of that gap.

1. Not root

Declare it in the image and enforce it at the platform:

dockerfile
RUN addgroup --system app && adduser --system --ingroup app app
USER app
yaml
securityContext:
  runAsNonRoot: true
  runAsUser: 10001

runAsNonRoot: true is the enforceable half — the kubelet rejects a pod whose user is 0 rather than trusting the Dockerfile to have done it. Set both, because the Dockerfile alone is a request and the security context is a rule.

Where the platform supports it, user namespaces add a second layer: the container's root maps to an unprivileged uid outside, so even a genuine escape lands as a no-privileged user rather than the host's root. It is the single most valuable setting that most clusters still have switched off.

2. Read-only root filesystem

An attacker who has landed in a container wants to write: a reverse shell, a cron job, an appended line in .bashrc. Make that harder by default:

yaml
securityContext:
  readOnlyRootFilesystem: true

This is the setting that breaks things first, and it is worth knowing why before you roll it out. Applications write to /tmp, log to a directory they own, or build cache under the working directory. All of those now fail — which is useful information about your application, delivered as an error.

The fix is to declare the writable paths explicitly:

yaml
volumeMounts:
  - name: tmp
    mountPath: /tmp
volumes:
  - name: tmp
    emptyDir: {}

Apply it to one deployment at a time and read the errors. In my experience the errors fall into two buckets: something genuinely needs to write, and something was writing only because nobody ever asked whether it should.

3. Capabilities, dropped wholesale

Linux capabilities split the root privilege into named pieces. Containers are granted a short list by default — NET_RAW and NET_ADMIN among them — and almost no application needs any of it.

Start by dropping everything, then add back only what a real error proves you need:

yaml
securityContext:
  allowPrivilegeEscalation: false
  capabilities:
    drop: ["ALL"]
    add: ["NET_BIND_SERVICE"]

NET_BIND_SERVICE is the common one: a process binding port 80 or 443 without being root needs it and nothing else. allowPrivilegeEscalation: false additionally blocks the setuid and ptrace tricks that turn a small foothold into a larger one.

4. seccomp, and AppArmor if your runtime has it

seccomp filters which syscalls a process may make. The default profile already blocks the obviously dangerous ones, and turning it on explicitly means it cannot be quietly skipped:

yaml
securityContext:
  seccompProfile:
    type: RuntimeDefault

RuntimeDefault is the right answer almost always. Writing a custom profile is a project, not a setting, and an inaccurate one produces failures that look like application bugs. If your nodes run AppArmor, apparmor.dock/default on a Debian-based node is a similar low-effort addition.

Apply them in this order

  1. runAsNonRoot — least disruptive, highest value.
  2. Drop all capabilities — usually invisible until something needs NET_BIND_SERVICE.
  3. readOnlyRootFilesystem — expect to add emptyDir mounts.
  4. seccompProfile: RuntimeDefault — should be a non-event.

Doing them in the other order means you debug three problems at once and cannot tell which setting caused which.

What to leave for its own change window

Custom seccomp profiles, mandatory SELinux policies on every node, and rootless container daemons on day one — that last one is valuable but changes how your images and build tooling behave, so it wants its own window.

Summary

Image scanning and runtime configuration answer different questions: one asks what is in the container, the other asks what it can do. Run as a non-root user and enforce it, make the root filesystem read-only with declared writable paths, drop every capability and add back only what an error proves necessary, and enable the default seccomp profile. Four settings, applied one at a time, and the realistic escape routes through a container narrow considerably.

#security#containers#docker#devops

SDP Clouds Team

DevOps and cloud engineers writing practical, battle-tested guides on CI/CD, Kubernetes, infrastructure as code, and production operations — every article is based on real incidents and real pipelines, not docs-page rewrites.

More about us →

Related articles