Runtime Container Scanning: Why Build-Time Isn't Enough
By Doug Edgar ·
Most container security tools solve the wrong half of the problem. They scan images before deployment, flag known CVEs, and declare victory. Then the container starts running, and nobody watches what happens next.
This post explains why runtime scanning matters, how it works at a technical level, and where it fits alongside the build-time tools you probably already have.
The Build-Time Blind Spot
Build-time image scanning tools like Trivy, Snyk Container, Grype, and others, are still valuable. They check your container images against databases of known vulnerabilities and produce reports before anything reaches production. If your base image ships with a vulnerable version of OpenSSL, they'll tell you.
But build-time scanning answers a narrow question: "Does this image contain known-vulnerable packages?" It does not answer:
- Did something download a crypto miner via curl in an entrypoint script after the container started?
- Did an attacker exploit a web application vulnerability and drop a web shell into the container's writable layer?
- Is a container that has been running for 90 days now serving malware through a compromised WordPress plugin?
- Did a process write a binary to a memfd (memory-backed file descriptor) to execute code without touching the filesystem?
These are runtime problems, and no amount of pre-deployment scanning will detect them. Because the malicious content does not exist until after the container is running.
How Containers Store Runtime Changes
Understanding why runtime scanning works requires understanding how container filesystems are structured.
Every running container uses an OverlayFS filesystem. The base image layers are read-only, this is the "immutable" part of container infrastructure. But on top of those layers sits a writable layer called the upperdir. Every file created, modified, or deleted during the container's lifetime is recorded in this directory.
When malware lands in a running container, whether downloaded by an attacker, dropped by an exploit, or created by a post-startup script, it lands in the upperdir. The base image remains clean. A rescan of the original image would show nothing wrong. But the container is compromised, and the evidence is sitting in the writable layer waiting to be found.
This is the fundamental insight behind runtime container scanning: you don't scan the image. You scan the upperdir.
Runtime Scanning Techniques
A runtime scanner operating on a Kubernetes cluster runs as a DaemonSet, one instance per node, with access to the host filesystem. From there, it can employ several detection techniques, each catching a different class of threat.
OverlayFS Upperdir Scanning
The most straightforward approach. The scanner locates each container's upperdir on the host filesystem and runs antivirus signatures against every file in it. This catches any malware that was written to the container's writable layer after startup, like downloaded binaries, web shells, crypto miners, modified configuration files, and staged payloads.
This is analogous to a traditional antivirus scan, but scoped to the runtime delta of each container. Because you're scanning only files that changed since the image was deployed, the scan is fast and the results are immediately actionable. Anything found here was not in the original image.
Real-Time File Monitoring with fanotify
OverlayFS scanning is periodic, you discover malware on the next scan cycle. For near-real-time detection, Linux provides fanotify, a kernel-level file access notification system.
A runtime scanner can place fanotify watches on each container's upperdir. When a process inside the container writes and closes a file, the kernel delivers an event to the scanner. The scanner can then inspect that specific file immediately, rather than waiting for the next full scan.
This is particularly effective against staged attacks where malware is downloaded, written to disk, and executed in rapid succession. With fanotify monitoring, the scanner sees the write event and can trigger a scan before the malware has fully established itself.
The engineering challenge is scale. On a node running dozens or hundreds of containers, each with its own upperdir, the volume of file events can be substantial. Effective implementations use debounced batching, like collecting events over a short window (a few seconds) and scanning the batch, rather than triggering a scan for every individual file write.
Deleted Binary and Library Detection
Sophisticated malware deletes itself from the filesystem after loading into memory. The process continues running, but the file is gone, invisible to any scanner that only looks at files on disk.
A runtime scanner can detect this by examining /proc/<pid>/exe for running processes. If the symlink points to a path marked (deleted), the binary was removed after execution. Similarly, scanning /proc/<pid>/maps reveals shared libraries that have been loaded and then deleted from disk.
This technique catches a specific and dangerous class of evasion: malware that runs entirely in memory after cleaning up its filesystem footprint.
memfd Execution Detection
An even more advanced evasion technique avoids the filesystem entirely. Linux's memfd_create syscall creates anonymous, memory-backed file descriptors. Malware can write a binary to a memfd and execute it directly, and the executable never gets written to disk at all.
These show up in /proc/<pid>/exe as references to /memfd: paths. A runtime scanner that checks process maps can flag these as suspicious, since legitimate applications rarely use memfd for code execution.
Where Runtime Scanning Fits in the Stack
Runtime scanning does not replace build-time scanning. They are complementary:
| Capability | Build-Time | Runtime | |---|---|---| | Known CVEs in base image | Yes | No | | Misconfigured packages | Yes | No | | Post-deployment malware | No | Yes | | Runtime file modifications | No | Yes | | In-memory-only execution | No | Yes | | Container drift detection | No | Yes | | Long-lived container compromise | No | Yes |
The build-time tools are your first line of defense. They prevent known problems from reaching production. Runtime scanning is your final line, it catches everything that gets past the gate or arrives after it.
The Competitive Landscape
Enterprise runtime security tools exist. Sysdig Secure, Aqua Security, and Palo Alto's Prisma Cloud all offer runtime protection capabilities. They are also priced for large enterprises, typically commanding enterprise-tier yearly premiums, with complex licensing models tied to per-host or credit-based pricing.
For a 25-node cluster, that translates to tens of thousands of dollars per year for runtime security alone.
This pricing model effectively locks out startups, small teams, and individual platform operators from runtime container scanning. They get build-time scanning (often through free or low-cost tools like Trivy) but have no affordable path to runtime protection.
Open-source alternatives like Falco provide runtime anomaly detection through syscall monitoring, which is valuable but architecturally different from file-based malware scanning. Falco detects suspicious behavior patterns; it does not scan files for known malware signatures. The two approaches are complementary.
What to Look For in a Runtime Scanner
If you're evaluating runtime scanning for your Kubernetes environment, the capabilities that matter most are:
OverlayFS-aware scanning. The scanner must understand container filesystem structure and scan the writable layer specifically, not just mount points or host paths.
Custom signature support. Generic antivirus signatures catch generic malware, but platform-specific threats, phishing kits, spam tools, and abuse patterns all require custom detection rules. The ability to write and deploy your own signatures is essential for any team operating a multi-tenant platform.
Node-level deployment. The scanner should run as a DaemonSet or equivalent, with one instance per node. This ensures every container on every node is covered without requiring per-pod sidecars or application-level integration.
Real-time monitoring. Periodic scans catch threats on the next cycle. fanotify-based or equivalent real-time monitoring catches threats as they land, closing the window between compromise and detection.
Automated response. Detection without response is just logging. The ability to annotate, quarantine, or generate Kubernetes events for detected threats enables automated remediation workflows, not just alerts that someone has to notice and act on manually.
Closing the Gap
The container security industry has spent years perfecting build-time scanning. The tooling is mature, accessible, and often free. Runtime scanning is where the gap remains, especially for teams that can't justify six-figure enterprise contracts.
If your containers run for more than a few minutes, if your users deploy code you don't fully control, or if you operate in a compliance environment that requires demonstrating runtime security controls, then runtime scanning is not optional. It's the difference between securing what you deploy and securing what actually runs.