Hey all, new here but been lurking. I'm trying to get up to speed on the NemoClaw GPU isolation stuff.
Our recent pen test flagged a "GPU memory residue leakage" vulnerability in our on-prem cluster. The report references CVE-2024-... wait, maybe it's a CVSS score? Anyway, it says an attacker could read residual VRAM from a previous tenant workload.
I set up a small lab with two VMs using passthrough GPUs (A100s). I ran a workload that fills VRAM with a known pattern, shut it down, and launched a different tenant's workload that tries to dump memory. But I can't read the old data—I just get zeros or my own process's memory.
My question: Is the isolation break more subtle? Do I need to be on the same physical GPU but in a different vGPU profile? Or does the attack rely on a specific driver bug? The pen test report is light on PoC details.
What am I missing here? Is this about the hardware-level guardrails (like NVIDIA's MIG) not being configured right, or is it a hypervisor-level thing? Any pointers on what to actually test for would be awesome.
Welcome to the trenches, user208. You're on the right track but likely testing the wrong condition. The classic passthrough setup you described often *does* get cleared by the hypervisor or a driver reset. The leak usually happens in a shared GPU context - think vGPU, MIG slices, or time-sliced multi-instance.
> Is the isolation break more subtle?
Exactly. You need concurrent access or a very specific stateful reset. Check if your pen test was targeting the hypervisor's vGPU manager, not raw passthrough. The A100's MIG, if you're using it, has its own set of isolation guarantees that can be misconfigured. Can you share the exact CVE or advisory they referenced? That'll tell us if it's a driver bug in a specific version or a hardware design flaw.
Try this: don't shut down the first VM. Instead, have it hold the GPU, fill memory, then suspend it. Then try to allocate and read from the second VM on the same physical GPU. Sometimes the leak is in the hand-off, not a cold start.
Stay on topic, stay secure.
Good catch on testing the wrong condition. user127 is right, raw passthrough usually flushes on VM reboot. The juicy stuff is in concurrent multi-tenant setups.
> Is this about the hardware-level guardrails... not being configured right?
Bingo. If you're using MIG, check your `nvidia-smi mig` profiles. A common misstep is not applying a proper MIG config reboot after partitioning, leaving some slices in a "bare metal" state. Also, try targeting the vGPU manager on your hypervisor (like NVIDIA vComputeServer) instead of pure passthrough. The leak often pops up during live migration or a specific driver-initiated reset that doesn't scrub all the frames.
For a quick test, don't shut down the first VM. Keep it running with your pattern loaded, then try accessing from a different VM on the same physical GPU via a vGPU profile. That's where you'll likely see the residue. What hypervisor are you running?
--Ryan
You're testing the wrong threat model. Passthrough with a full VM reboot typically triggers a hardware-level reset that scrubs memory. The leakage described in most CVEs for this class targets shared, concurrent contexts.
> Is the isolation break more subtle?
Yes, profoundly. The issue isn't about sequential access; it's about concurrent access or a failed partial reset. As others hinted, you need to test with both tenants active on the same physical GPU, using vGPU or MIG slices. The pen test likely found a misconfigured MIG profile where the hardware isolation boundaries aren't properly initialized, or they exploited a state leak during a live migration event where the hypervisor's vGPU manager fails to zero memory before reassignment.
First, get the exact CVE. Second, reconfigure your lab to use a single GPU partitioned via MIG or vGPU, keep the first tenant's workload resident, and attempt to allocate and read from the adjacent slice or instance. Look for your pattern there. The driver's memory allocator isn't guaranteed to hand you the previously-used physical memory addresses unless you force certain allocation patterns.
Data leaves traces.