Skip to content

Reap the container cgroup when the runc shim is stopped - #14220

Open
yongxiu wants to merge 1 commit into
containerd:mainfrom
yongxiu:shim
Open

yongxiu wants to merge 1 commit into
containerd:mainfrom
yongxiu:shim

Conversation

@yongxiu

@yongxiu yongxiu commented Sep 24, 2026 •

Copy link
Copy Markdown

Reap the container cgroup when the runc shim is stopped

Background

We run containerd on bare metal Kubernetes nodes and found nodes where containerd took 30+ seconds to start, every time it was restarted. Each start logged the same cleanup failures for the same set of old containers:

level=info  msg="cleaning up dead shim" id=<id> namespace=k8s.io
level=warning msg="failed to clean up after shim disconnected" error="unmount rootfs /run/containerd/io.containerd.runtime.v2.task/k8s.io/<id>/rootfs: device or resource busy"
level=error msg="failed to delete dead shim" error="signal: killed" id=<id>
failed to remove ... cri-containerd-<id>.scope: device or resource busy

The bundles of these containers could never be deleted. Every containerd start loaded them again, ran the same failing cleanup (with its unmount retries, until the 5s dead shim cleanup timeout), and gave up. On containerd 1.6/1.7, which loads shims serially, this adds about 5s of boot time per leaked bundle. On 2.x it happens in parallel, but the bundles, cgroups and processes still leak forever.

On such a node, the container cgroup still held processes (in some cases a frozen runc:[2:INIT]). They pinned the rootfs mount, so the unmount failed with EBUSY. Bundle.Delete() then failed, and the dead shim stayed around for the next start.

Root cause

manager.Stop() in the runc shim is what containerd runs to clean up after a shim that has died, and after a task whose creation failed. For tearing the container down, it relies entirely on runc delete --force. That is not enough:

In both cases the processes left in the container cgroup keep the cgroup populated and the rootfs busy, and nothing ever reaps them.

Fix

After runc delete --force, whatever it reported, Stop() now reaps the container cgroup directly: it kills whatever is left in it, waits for it to empty, and removes it.

  • The cgroup path comes from linux.cgroupsPath in the bundle's config.json. containerd writes that file before runc is ever invoked, so this does not depend on any runc state.
  • The runtime's SystemdCgroup option decides whether the path uses systemd slice:prefix:name notation or is a literal cgroupfs path, the same way runc decides.
  • cgroup v2: it writes cgroup.kill (which also kills frozen processes). On kernels older than 5.14 it falls back to the same bounded loop as v1.
  • cgroup v1 (and the v2 fallback): each round freezes the cgroup, lists its processes, SIGKILLs them, and thaws the cgroup before waiting, because frozen tasks do not act on SIGKILL until thawed. The thaw also unsticks a cgroup that a killed runc create left frozen. Signals go through pidfds, re-checked against the cgroup membership, so a reused PID outside the cgroup is never signalled.
  • The whole reap is bounded (2s), since it runs during containerd startup.
  • When runc did its job, the cgroup is already gone and this is a no-op.

This kills nothing that runc delete --force is not already meant to kill: runc treats the container cgroup as owned by the container.

The fix lives in the shim because Stop() is the one place that runs in both cleanup paths: the create failure and the dead shim cleanup at boot. It also does not depend on how runc eventually addresses #4757. It covers both the state.json gap and a plain runc delete failure.

How to reproduce

I wrote a small, deterministic reproducer: https://github.com/yongxiu/leakrepro

It creates a container through the containerd client with an OCI createRuntime hook. runc runs that hook after it has created the cgroup and mounted the rootfs, but before it writes state.json. The hook freezes the container cgroup and SIGKILLs runc create. That leaves exactly the production end state:

  • a frozen runc:[2:INIT], reparented to PID 1, still in the container cgroup and pinning the rootfs;
  • no state.json, so runc delete --force exits 0 and does nothing.

No image or snapshotter is needed. It reproduces 100% of the time.

On a throwaway node (it deliberately leaks a process):

leakrepro create --address /run/containerd/containerd.sock --id leakrepro1
leakrepro inspect --id leakrepro1
systemctl restart containerd
journalctl -u containerd | grep "successfully booted"

The repo also has run-repro.sh. It runs each containerd build in a privileged, throwaway Docker container with its own containerd, so the host is not touched.

Results

runc 1.3.6, cgroup v2, 3 runs each, identical every run:

containerd NewTask fails after boot after each restart leftovers
main (stock) 12.6s 7.56s, on every restart bundle + rootfs mount + frozen runc:[2:INIT] + cgroup
this PR 2.6s 0.02s none

I also verified it on a real cgroup v2 node running containerd v2.2.1, with only this change cherry-picked into the shim. With the stock shim the container leaks. With the patched shim, leakrepro inspect reports CLEAN right after the failed create, and containerd restarts stay fast. After the fix there is no leak at all.

Tests

cgroup_linux_test.go adds unit tests for the path resolution (systemd notation including the root slice -.slice, cgroupfs paths with colons) and the choice of cgroup v1 subsystem. It also runs real-kernel tests against a delegated child cgroup: a live process, a frozen process, empty and already-removed cgroups, the bounded freeze-and-kill fallback, and a check that a listed PID that is no longer in the cgroup is not signalled. Those tests skip themselves when cgroup v2 delegation is not available.

Related: opencontainers/runc#4757, opencontainers/runc#5153, opencontainers/runc#5257

Copilot AI lite review requested due to automatic review settings September 24, 2026 14:56
@github-project-automation github-project-automation Bot moved this to Needs Triage in Pull Request Review Sep 24, 2026

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Critical and moderate cleanup correctness issues remain unresolved.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 1 High severity

Open (1)
What changed in this PR

Adds fallback cgroup cleanup after runc delete --force to reap leaked container processes and remove stuck cgroups.

Changes:

  • Supports cgroup v1/v2 and cgroupfs/systemd paths.
  • Kills, waits for, and removes leftover cgroups.
  • Adds shutdown integration and Linux cleanup tests.
File Summary
cmd/​containerd-shim-runc-v2/​manager/​manager_linux.go Invokes cgroup reaping during shim stop.
cmd/​containerd-shim-runc-v2/​manager/​cgroup_linux.go Implements cleanup. Findings: critical v2 root-slice path handling issue (2 votes); moderate v1 systemd cleanup issue (1 vote); nit for missing v1 test coverage (1 vote).
cmd/​containerd-shim-runc-v2/​manager/​cgroup_linux_test.go Tests path parsing and cgroup reaping. Findings: two moderate assertions can allow tests to hang (1 vote each).

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread cmd/containerd-shim-runc-v2/manager/cgroup_linux.go Outdated

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Critical cgroup cleanup correctness and safety issues remain unresolved.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 4 High severity

Open (4)
Resolved since last review (1)

Comment thread cmd/containerd-shim-runc-v2/manager/cgroup_linux.go Outdated
Comment thread cmd/containerd-shim-runc-v2/manager/cgroup_linux.go Outdated
Comment thread cmd/containerd-shim-runc-v2/manager/cgroup_linux.go Outdated
Comment thread cmd/containerd-shim-runc-v2/manager/cgroup_linux.go Outdated

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Root cgroup handling poses a critical safety risk, and test coverage has blocking reliability gaps.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 1 High severity · 1 Medium severity

Open (2)
Resolved since last review (4)

Comment thread cmd/containerd-shim-runc-v2/manager/cgroup_linux.go Outdated
Comment thread cmd/containerd-shim-runc-v2/manager/cgroup_linux_test.go Outdated

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Relative cgroup paths, v1 deletion timeouts, and pidfd-dependent test behavior remain unresolved.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 1 High severity

Open (1)
Resolved since last review (2)

Comment thread cmd/containerd-shim-runc-v2/manager/cgroup_linux.go
The runc shim's Stop, which containerd runs to clean up after a shim
that is gone or a task whose creation failed, relies entirely on
"runc delete --force" to tear the container down. That is not enough:

- runc can fail to tear the container down, which is only logged;
- runc reports success without doing anything when the container
  state.json was never written, which is what a "runc create" killed
  part way through leaves behind (opencontainers/runc#4757, reverted in
  containerd#5153). The "runc init" left over from such a create can even be
  frozen.

Either way, processes are left running in the container cgroup. They
keep the cgroup from being removed and the rootfs from being unmounted,
so the bundle can never be deleted: containerd reloads the same dead shim
and repeats the same failing cleanup, unmount retries included, on every
start, and the leftover processes are never reaped.

Reap the cgroup directly after runc, whatever runc reported: kill what is
left in it, wait for it to empty, and remove it. The cgroup path comes
from the bundle config.json, which containerd writes before runc is ever
invoked, so this does not depend on any runc state. Both cgroup v1 and
v2 are handled, with either the systemd or the cgroupfs path notation as
selected by the SystemdCgroup option. The reap is bounded, as it runs
while containerd starts up, and it never signals a process outside of
the cgroup: processes are listed while the cgroup is frozen and are
signalled through pidfds. When runc did its job the cgroup is already
gone and this is a no-op.

This kills nothing that "runc delete --force" is not already meant to
kill: runc treats the container cgroup as owned by the container.

Signed-off-by: Yongxiu Cui <cuiyongxiu@gmail.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

Address the three unresolved cleanup timing, context cancellation, and pidfd compatibility findings.

Review effort: Lite
Findings: None

Resolved since last review (1)

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

Status: Needs Triage

Development

Successfully merging this pull request may close these issues.

2 participants