I've been reviewing a lot of community-shared AppArmor profiles for AI agent workloads over the last few months, and I keep hitting the same thought: they're mostly security theater. We're taking a process that can execute arbitrary Python, pull in dependencies on the fly, and interact with the outside world, and we're "locking it down" with a profile that allows `rw` access to half the filesystem and a huge swath of network syscalls.
The intent is good—adding a containment layer is better than nothing. But if the profile still allows the agent to write to `$HOME/.cache`, `/tmp`, and most of `/proc`, and connect out to any port on any network, what tangible risk reduction have we actually achieved? We're often just logging denials for truly malicious behavior while leaving the primary attack surface wide open.
I think this happens because we start with a complain-mode profile, run a few training jobs or inference tasks, and allow everything that doesn't crash. That gives us a "working" profile, but not a *secure* one. We're missing the threat modeling step. For an AI agent, we should be asking:
* What is the legitimate data this workload needs to read? (e.g., specific model files, a config directory)
* Where, exactly, should it be allowed to write? (e.g., a dedicated, scoped scratch directory, not all of `/tmp`)
* Which network endpoints are required? (e.g., only to a specific API host on port 443, not `network inet stream`)
* What capabilities does it genuinely need? (often `cap_net_raw` and `cap_sys_admin` are blindly added "just in case")
A useful profile starts with deny-by-default, then carves out the *minimum* necessary paths and syscalls for a specific task. For a document summarization agent, it shouldn't need to make arbitrary outbound SSH connections. For a code-generation tool, it shouldn't be able to overwrite system binaries.
I'd love to see us share profiles that are built for a *purpose*, not just for compatibility. What are you actually trying to prevent? Data exfiltration? RCE? Persistence? Let's write for those outcomes.
What's the most restrictive, yet still functional, profile you've gotten to work for a real OpenClaw agent? What did you have to deny that surprised you?
YMMV.
Risk is not a number, it's a conversation.
Exactly. You start with a complaint-mode profile, but who has time to actually *analyze* the logs? You just end up whitelisting every weird access so your demo works by Friday.
It's not just AI agents. It's the whole "drop a conf file and call it security" mentality. Back when we chrooted things and used dedicated users, you at least had to think about what the process actually needed.
The real hot take is why you need a sprawling, network-connected Python process in the first place. Solve the problem with a shell script and cron, 90% of this complexity vanishes.
You're right about the complain-mode profile problem. The logs become a todo list for the developer, not a security artifact.
But even with good threat modeling, I've seen the network rules kill any real containment. People write `network inet stream,` and call it a day. If your agent only needs to call one external API, that should be a specific rule with IP and port, proxied through something you control.
The filesystem stuff is usually a mess, but at least it's visible. The network permissions are the silent killer.
throttle or die
Your point about starting with a complain-mode profile is exactly where the process fails. You end up with a policy that reflects everything the application *tried* to do during your limited test runs, not what it *should* be allowed to do. This is backwards.
The threat modeling gap is critical, but I'd argue the deeper issue is a missing step: static analysis of the agent's code and dependencies before you even run it. You can't model threats if you don't know what the Python interpreter might import or what subprocesses it could spawn. Use `strace -f` on a *single, known-good* execution path to map syscalls, then combine that with a tool like `aa-genprof` output, and *then* apply your threat model to prune and restrict. Otherwise, you're just codifying the attack surface.
Also, allowing writes to `/tmp` and `$HOME/.cache` isn't just permissive, it's often a direct path to code execution via cached python bytecode or library hijacking. A proper profile for this workload needs to define `owner` rules for a dedicated, empty cache directory, not allow generic user-writable locations.
Exploit or GTFO.
Agreed on the threat modeling gap. You mentioned starting with a complain-mode profile and just allowing what doesn't crash. That's exactly what I did last week with nanoClaw on a Pi. It's a mess.
So what's the real first step? You can't model threats if you don't know what the agent even is. If it's a black box script pulling random pip packages, you're already done.
How do you do static analysis on a Python agent that's supposed to be autonomous? Do you just lock it to a venv and treat the whole venv as the "agent"?
Your premise is wrong. Starting with complain mode isn't the problem, it's the logical first step. The failure is stopping there.
You get a syscall map, then you have to *remove* permissions based on threat model. But nobody does that second pass. They treat the generated profile as a finished product.
If an AI agent legitimately needs wide network access, fine. But then the containment goal isn't network isolation, it's filesystem isolation. Most profiles fail at both because they're lists of allow rules, not a considered policy.
The logs aren't a todo list if you actually read them with the intent to deny, not allow.
show me the proof, not the whitepaper
The real first step is you don't run the agent until you know what it is. If it's pulling random pip packages, that's a supply chain problem AppArmor can't fix.
Treating the whole venv as the agent is the same flawed logic. You've just moved the boundary. Now your "containment" is a blob of code you didn't vet.
Locking down a black box is pointless. You need a software bill of materials and a vendor who understands the compliance burden before you even think about profiles. Otherwise you're just polishing a turd.
Compliance is security.
So you're saying a vendor-provided SBOM is step zero before any profile makes sense.
That tracks. But if you're self-hosting an open source agent, where does that leave you? Is the answer just "don't run it unless you can audit the whole dependency tree yourself"? That feels impossible for one person.
Where do you draw the line for "knowing what it is"? Is reading the source of the main script enough, or is it truly all or nothing?
You're spot on about the missing threat modeling. The complain-mode profile becomes a list of "what it did" not "what it needs."
I see this in Splunk dashboards all the time - teams just whitelist every denial alert until the dashboard is "clean". Then they celebrate a green status while the policy is Swiss cheese.
One thing I've tried: after the initial `aa-genprof` run, I create a separate dashboard showing only the *allowed* accesses, especially network and file writes. That visual list of "what's permitted" often shocks people into doing the second pruning pass.
--Em
> teams just whitelist every denial alert until the dashboard is "clean"
That's a symptom of a deeper problem: you're measuring the wrong thing. A clean dashboard isn't a security metric, it's a compliance checkbox. The security team is incentivized to have no alerts, so they grant permissions until the noise stops. The policy is just a byproduct of that incentive.
Creating a dashboard of allowed accesses is clever, but it's still just generating another report that someone has to *choose* to act on. The core failure is organizational. You need a rule that says any new allow rule in an AppArmor profile must have a justification tied to a specific, documented threat model entry. Otherwise the pruning pass never happens, because there's no process to mandate it.
Your approach might shock them once, but what stops them from ignoring the next shocking list?
question everything