I've been conducting a security review of our NemoClaw deployment's NeMo Inference Microservice (NIM) containers, with a specific focus on hardening runtime privileges. As part of this, I've been attempting to enforce a custom AppArmor profile to restrict the container's capabilities beyond the default Docker runtime. However, the container consistently fails to start when the profile is applied, exiting with a non-zero code and sparse logs.
The core of my issue appears to be a conflict between the profile's denials and the NIM container's runtime expectations. I have constructed a profile based on a principle of least privilege, denying writes to most of the filesystem, restricting network access to only necessary ports, and limiting capability sets. The container's `docker run` command is as follows:
```bash
docker run --rm
--name test-nim
--security-opt "apparmor=nim-hardened"
-p 8080:8080
nvcr.io/nvidia/nemo/nemoinfer:latest
```
The container logs, obtained via `docker logs` on the briefly extant container, are not particularly verbose but hint at a failure during initial model loading or internal initialization. The relevant entries from `/var/log/syslog` showing AppArmor denials are:
```
type=AVC msg=audit(1678901234.567:890): apparmor="DENIED" operation="open" profile="nim-hardened" name="/proc/self/status" pid=12345 comm="python3" requested_mask="r" denied_mask="r" fsuid=0 ouid=0
type=AVC msg=audit(1678901234.568:891): apparmor="DENIED" operation="mknod" profile="nim-hardened" name="/dev/urandom" pid=12345 comm="python3" requested_mask="w" denied_mask="w" fsuid=0 ouid=0
```
My primary questions for the community are:
* What are the specific filesystem locations, kernel interfaces (e.g., `/proc`, `/sys`), and special device nodes (e.g., `/dev/`) that a NIM container requires for standard operation? The NVIDIA documentation is silent on mandatory runtime permissions.
* Has anyone successfully deployed NIM containers under a custom, restrictive AppArmor or SELinux policy? If so, what were the critical allow rules?
* Beyond simple filesystem access, which Linux capabilities (e.g., `CAP_NET_BIND_SERVICE` for binding to port 8080, `CAP_SYS_PTRACE` for any internal profiling) are absolutely required? The default Docker profile may grant a broad set.
My immediate goal is to construct a functional baseline profile. The longer-term security objective is to share a hardened template that can be adapted for production deployments, particularly those where the NIM endpoint is exposed via the OpenClaw plugin API and must adhere to strict container isolation standards. Any insights from those who have delved into the runtime behavior of these inference containers would be invaluable.
- Lei
Defense in depth for APIs.
The sparse logs from the container and the syslog denials are the key. You need to correlate them precisely. The `docker logs` output showing a failure during model loading likely masks the actual, earlier denial from AppArmor that prevented a critical setup step, like mounting a tmpfs or accessing a shared library.
Run the container with `apparmor_parser` in complain mode for the profile first, then check `dmesg` or the audit logs. You'll see the exact sequence of denied operations. NIM containers often require specific, non-standard capabilities like `SYS_ADMIN` for GPU memory mapping or `DAC_OVERRIDE` for certain model files. Your principle-of-least-privilege approach is correct, but you have to start from a working baseline, not a fully locked one.
Can you post the specific denials from `/var/log/audit/audit.log`? The profile's network rules might also be too restrictive; the container might need outbound DNS resolution to contact NVIDIA's licensing service on startup, which would cause a silent hang.
~Eli
Start in complain mode is solid advice. I'd add that the NIM containers sometimes need weird permissions on /dev/ devices for the GPU, not just capabilities. My first profile blocked access to /dev/nvidia-uvm and it died silently.
Could you post the denials from dmesg after running in complain mode? I'm trying to lock down my own setup and seeing what you hit would help.