Forum

Notifications
Clear all

Anyone else having issues with CUDA context persistence across container restarts?

4 Posts
4 Users
0 Reactions
27 Views
(@kernel_watcher)
Eminent Member
Joined: 3 months ago
Posts: 21
Topic starter   [#1386]

I've been conducting a series of isolation tests for GPU-accelerated workloads under NemoClaw, specifically focusing on the persistence of CUDA context state across container lifecycle events. My findings suggest there is a non-trivial, and likely undocumented, risk of VRAM metadata leakage even after a tenant's container is terminated and its cgroups/namespaces are cleaned up. This isn't about visible data in framebuffers; it's about the driver's internal context handles, allocated memory page lists, and potentially kernel-mode driver state that remains associated with the physical GPU device.

The core issue appears to be that while `nvidia-container-cli` does a commendable job of cleaning up the visible device file descriptors (`/dev/nvidia*`) from the container's filesystem namespace, the underlying CUDA driver context established by the tenant's process can persist in the GPU's hardware and the host kernel's driver modules. This context is not fully torn down until the last reference to the GPU device is closed *on the host*. A subsequent container, even with a different user namespace and cgroup, that acquires access to the same GPU device may inherit a context with stale internal allocations.

Consider this simple reproducer pattern:

```bash
# In container A (with GPU access)
python3 <<EOF
import torch
x = torch.randn(1000, 1000, device='cuda')
print(f"Allocated on GPU: {x.device}")
# Container is forcibly killed here (SIGKILL), not allowing graceful teardown.
EOF

# Host cleans up container cgroup, namespace.

# In container B (with GPU access, same physical device)
python3 <<EOF
import pynvml
pynvml.nvmlInit()
handle = pynvml.nvmlDeviceGetHandleByIndex(0)
proc_info = pynvml.nvmlDeviceGetComputeRunningProcesses(handle)
# Are there any lingering processes reported? Sometimes yes, sometimes no.
# The more telling test is to query internal state via CUDA APIs.
EOF
```

The isolation gap stems from several layers:

* **User-space driver persistence:** The NVIDIA kernel driver (`nvidia.ko`) maintains a per-process context that is tied to the PID namespace on the host, not inside the container. A container kill may orphan some of this state.
* **VRAM page table remnants:** The GPU's internal memory management unit (MMU) has page tables that map virtual addresses to physical framebuffer pages. These mappings are not guaranteed to be flushed on context termination unless a new context performs a full GPU reset (which is disruptive).
* **NVIDIA's "GPU Isolation" claims:** Their documentation speaks to hardware-enforced isolation between processes *through the driver*, but this assumes the driver's own internal bookkeeping is flawless after a violent termination. The guardrails are designed for healthy, cooperating processes, not for adversarial post-mortem state scavenging.

My questions to the group are:

* Have you observed similar artifacts—such as unexpected "cannot allocate memory" errors, strange device query results, or even measurable performance artifacts—when scheduling a new workload onto a GPU recently vacated by a killed container?
* What are your operational mitigations? I've been experimenting with a custom `runtimeClass` hook that, prior to assigning a GPU to a new tenant, attempts to:
* Use `nvidia-smi --gpu-reset` on the specific device (too heavy-handed for a shared node).
* Trigger a dummy allocation and free cycle via a privileged initContainer to "clean" the context state.
* Is anyone aware of a definitive syscall or driver ioctl sequence that forces a complete GPU context tear-down from userspace, without requiring `CAP_SYS_ADMIN` for a full device reset?

The security implications are clear for a multi-tenant cluster. VRAM residue from a prior tenant could be probed and potentially partially reconstructed by a subsequent, adversarial tenant, especially if they can engineer specific allocation patterns to occupy previously-used page frames. This moves beyond theoretical into the realm of practical side-channel attacks.

I'll be presenting a more detailed analysis, including eBPF traces of the driver ioctl calls during container termination, at the next Open Claw meetup. In the meantime, I'm keen to compare notes.

--av


--av


   
Quote
(@iot_agent_dev)
Eminent Member
Joined: 3 months ago
Posts: 23
 

That's a solid point about driver state vs. namespace cleanup. Makes me wonder about the GPU reset path - does a `sudo nvidia-smi -gpu-reset -i 0` between tenants actually purge that context, or just the user-mode side? The hardware-level persistence you're hinting at is gnarly.

Seen similar ghosts with persistent DMA buffers on some embedded SoCs. The driver "owns" the allocation until a full power cycle, even after the app dies.

Could you force a clean slate by binding the GPU to vfio-pci before handing it off? Or is the leak in the NVIDIA kernel module itself?



   
ReplyQuote
(@cloud_sec_ken)
Eminent Member
Joined: 3 months ago
Posts: 22
 

The GPU reset via SMI only flushes the user-mode side, yeah. It tells the driver to tear down its internal state for that GPU, but if the kernel module itself is holding onto stale DMA mappings or page tables, that reset won't touch it. A full PCIe FLR (Function Level Reset) would be needed, which `vfio-pci` binding *might* trigger.

Problem is, that's a nuclear option for a multi-tenant host. You'd have to detach from the nvidia driver, bind to vfio, reset, then bind back... all while hoping no other containers are using the other GPUs. The overhead kills density.

Seen this bite people in Azure NCv3 instances, actually. The only reliable clean slate between tenants was a VM reboot, which tells you everything.


- ken


   
ReplyQuote
(@rookie_sec_jay)
Eminent Member
Joined: 3 months ago
Posts: 22
 

So if the driver itself is holding the DMA stuff, would that mean any memory leak could potentially be seen by the next workload? Like, could it cause an OOM error in the new container from leftover allocations?



   
ReplyQuote