"We Don't Have to Worry About Malware in Containers" - And Other Lies We Tell Ourselves
By Doug Edgar ·
Containers are ephemeral. They spin up, do their work, and get replaced. Malware can't persist if nothing persists, right?
I used to hear this at security reviews. The argument was clean and compelling: containers are immutable infrastructure. If something goes wrong, you just redeploy. The blast radius is small, the lifetime is short, and the attack surface is minimal. We run stateless applications, we use emptyDir volumes...
Then I started scanning containers on a production cluster. What I found changed how I think about container security entirely.
The Hundred-Day Container
I was operating a malware scanning system on a large, multi-tenant OpenShift cluster, which had several hundred nodes, serving a public PaaS where users could sign up with a free tier and run their own workloads. The operators half-jokingly called it "code-execution-as-a-service."
Among the many things we discovered was that containers are not nearly as ephemeral as the industry pretends. We routinely found user containers that had been running for over 100 days without a restart. It turns out, users without Highly Available deployments don't appreciate downtime on their applications. At least that part wasn't a surprise.
The reasons were practical. Restarting containers often meant evacuating nodes. Evacuating nodes meant potential downtime for customer deployments that weren't designed for it. Many deployments were single-pod applications with no redundancy, stateful workloads tightly coupled to their local storage or a single PV. For a platform operator, the calculus was simple: the risk of disruption from a restart outweighed the theoretical risk of... what, exactly? Malware in a container? It wasn't worth the churn unless the base host VMs needed to be patched for critical CVEs (e.g. rebooting the VM after updates to kernels or glibc.).
WordPress, the Gift That Keeps On Giving
The most common victim was WordPress. Users would deploy old versions, tightly coupled to a local database. The container was the application, the database, and the storage layer all in one. As long as the application "just worked", there wasn't significant concern for things like separation of concerns, stateless design, and multi-instance redundancy.
These containers would sit on the platform for weeks, sometimes months, exposed to the internet. Eventually, the inevitable happened: known WordPress vulnerabilities would get exploited. Malware would land in the container's writable layer. And because the container never restartedm because restarting it was operationally expensive and nobody was monitoring what was happening inside, the malware would persist for the entire lifetime of the container.
I found compromised WordPress instances that had been silently serving malware for months. The container was still "healthy" by every metric Kubernetes cared about: it responded to health checks, it served HTTP traffic, it wasn't crashing. It just also happened to be hosting a crypto miner, a phishing kit, or a spam relay.
The CVE Race
One pattern that caught my attention was how quickly some users moved after a CVE's exploitation steps were made public. Within days, sometimes hours, of a proof-of-concept being published, I could find users on the platform attempting to reproduce the exploit against the infrastructure itself.
Was this curiosity? Security researchers testing their own environments? Maybe. Was some of it malicious? Almost certainly. The platform had no way to distinguish intent, only behavior. And the behavior was: someone just tried to exploit a freshly-published vulnerability on shared infrastructure.
This is the reality of any multi-tenant platform with a low barrier to entry. You are only one CVE disclosure away from users probing your infrastructure for weaknesses. The ephemeral nature of containers doesn't help you here. The attacker's container may be ephemeral, but the damage to the host or to neighboring tenants is not. Especially if they can pivot to affect the cluster itself.
Why the "Containers Are Ephemeral" Argument Fails
The argument that containers don't need malware scanning rests on three assumptions, all of which are wrong in practice:
Assumption 1: Containers are short-lived.
In theory, yes. In practice, plenty of workloads run for days, weeks, or months. Stateful applications, legacy migrations, and workloads without proper orchestration will all stick around far longer than anyone designed for. And the longer a container runs, the more time an attacker has to compromise it, and the longer that compromise persists undetected.
Assumption 2: Container images are immutable, so malware can't persist.
The image is immutable. The container is not. Every running container has a writable layer. In the case of OCI-compliant containers, the OverlayFS "upperdir" is where runtime changes land. Malware that gets downloaded after startup, files modified during exploitation, and crypto miners installed through a web shell will all land in the writable layer. Totally invisible to any tool that only scans the original image in the container registry.
Assumption 3: Redeployment clears any compromise.
True, if you redeploy. But redeployment requires knowing something is wrong, and knowing something is wrong requires monitoring. Without runtime scanning, there is no signal. The container appears healthy. It passes readiness probes. It serves traffic. The malware persists because nobody has a reason to restart the container, and nobody has a mechanism to detect the compromise. And realisticly, if someone knows that they can repeatedly, reliably compromise your workloads, they'll have the incentive to keep coming back as many times as they can.
The Scanning Gap
Here's the uncomfortable truth about container security today: almost every organization that bothers to scan at all, will only scan at build time. At best, they check the image for known vulnerabilities in the registry before deployment. This is still valuable, as it catches outdated packages, known CVEs, and misconfigured base images.
But it misses everything that happens after the container starts running. And it falls short of anything I would consider "defense in depth".
It misses the WordPress plugin vulnerability exploited three weeks into a container's lifetime. It misses the crypto miner downloaded via curl in a startup script. It misses the obfuscated PHP web shell dropped through a file upload vulnerability. It misses the C2 beacon that only phones home after the container has been running for six hours to avoid sandbox detection.
Build-time scanning answers the question: "Was this image safe when we built it?" Runtime scanning answers the question that actually matters: "Is this container safe right now?"
The Path Forward
Far from being hypothetical, theoretical attack vectors, these are real-world observations from operating a malware scanning system on real infrastructure, serving real users, at real scale. The patterns we've saw, such as long-lived containers, post-deployment compromise, malware persisting for months, are not edge cases. They are the default behavior of any sufficiently large Kubernetes or OpenShift deployment with mixed workloads.
The industry needs to stop treating container security as a build-time problem. Build-time scanning is necessary, but not sufficient by itself. If you are not scanning containers at runtime, inspecting the writable layer, monitoring file changes, and detecting post-deployment compromise, then you have a blind spot large enough to hide a crypto mining operation. Or a phishing kit. Or, much worse.
In the real world, SRE teams cut so many corners they're left with a circle. Containers regularly run with higher privileges than needed. So if someone compromises a web application dependency, they can pivot onto the host filesystem, steal kubelet creds, pivot to the rest of the cluster, and keep going if they can.
Containers are not as ephemeral as we tell ourselves. And the lies we tell ourselves about security are the ones that hurt the most.