Forum

Notifications
Clear all

Just released an open-source tool to audit AutoGen agent capabilities

3 Posts
3 Users
0 Reactions
37 Views
(@peter_hardener)
Eminent Member
Joined: 3 months ago
Posts: 20
Topic starter   [#1318]

Just finished a weekend project I wanted to share. I've been digging into AutoGen's security model, specifically around those powerful `UserProxyAgent` and `AssistantAgent` with code execution. The defaults are, frankly, terrifying for any kind of production-adjacent use. 😅

I built a simple static analysis tool, `autogen-audit`, that parses an AutoGen agent configuration and flags high-risk settings. It focuses on the capability model. Here's the core idea:

```python
# Example of what it flags
risky_config = {
"name": "CoderAgent",
"system_message": "You are a helpful assistant.",
"code_execution_config": {
"work_dir": ".",
"use_docker": False, # 🚨 Flagged: Missing Docker isolation
"last_n_messages": 3
},
"human_input_mode": "NEVER" # 🚨 Flagged: No human oversight
}
```

The tool checks for:
- Code execution enabled without Docker isolation.
- Missing `human_input_mode` on code-executing agents (autonomous loops).
- Overly permissive `work_dir` paths (e.g., "/", "~").
- Default `llm_config` allowing unlimited tool use.

It's not a runtime sandbox—that's a separate layer needing seccomp or AppArmor. This is about catching misconfigurations before you deploy. Found it super useful for my own team's setup; we were accidentally running with `use_docker: false` in a staging environment.

You can find it on GitHub under `openclaw-security/autogen-audit`. It's a simple Python script. Would love feedback, especially on what other static checks would be valuable. Anyone else looking at CrewAI's role/permission design? I'm thinking of adding support for their `allow_delegation` and `function_calling_llm` checks next.

-- peter


default deny


   
Quote
(@red_team_ops_ray)
Eminent Member
Joined: 3 months ago
Posts: 15
 

Good focus on the config. Static analysis is a solid first pass, but you're right it's not runtime. The dangerous stuff happens in execution context escapes and tool chaining.

I'd add checks for unrestricted file system access in the code execution config. If `work_dir` is `.`, but the agent can `os.chdir("/")` or call `subprocess.run`, your Docker flag is meaningless.

Also watch for custom function registrations that wrap `eval` or `exec` without validation. Seen that pattern in three projects last month.


--Ray


   
ReplyQuote
(@hype_killer)
Eminent Member
Joined: 3 months ago
Posts: 18
 

A static check for `use_docker: false` misses the point. Docker without a read-only filesystem and network namespace is just a suggestion.

You need to audit the *runtime* permissions, not the config flag. The agent can still call `docker run` itself if the LLM is allowed to execute arbitrary code. Your tool might green-light a config while the process runs with `CAP_SYS_ADMIN`.



   
ReplyQuote