Forum

Notifications
Clear all

How to prevent AutoGen agents from exfiltrating data through the network?

6 Posts
6 Users
0 Reactions
27 Views
(@junior_harden_jay)
Eminent Member
Joined: 3 months ago
Posts: 24
Topic starter   [#1317]

Hey everyone, new to the forum and diving into AutoGen. I've been setting up some multi-agent workflows locally, and a question keeps nagging at me.

I understand that `UserProxyAgent`s with code execution can run `requests.get()` or use other Python modules to make network calls. Even a simple `AssistantAgent` could, in theory, generate a code block that the `UserProxyAgent` would then execute, potentially sending data out.

I want to sandbox these agents to prevent any unauthorized data exfiltration. My goal is to allow them to compute and talk to each other, but block all network egress from the agent's execution environment, unless it's to a specific, allowed internal service (like a local LLM).

I'm thinking about using Docker to containerize the whole AutoGen runtime. What would be the best practice here?

1. Is it enough to run the AutoGen script inside a container with `--network=none`? Or would that break inter-agent communication if they're separate processes?
2. Should I be looking at Linux network namespaces or `iptables` rules on the host instead?
3. How do you handle cases where an agent *needs* to fetch something from a known, safe API? Is a proxy the only secure pattern?

Here's a super basic Docker setup I'm considering, but I'm unsure about the networking part:

```dockerfile
FROM python:3.11-slim
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY . .
# Is --network=none the right flag to use at runtime?
CMD ["python", "my_autogen_crew.py"]
```

I'd really appreciate some step-by-step guidance or examples of how you've locked this down in your own projects. My expertise is more in basic Docker and Linux, so the deeper security mechanics are a bit new to me.

Thanks - Jay



   
Quote
(@threat_model_wizard_ray)
Eminent Member
Joined: 3 months ago
Posts: 20
 

Docker with --network=none is a solid start for the execution environment, but you're right to worry about breaking inter-agent comms. If all your agents are in the same container process, they talk via function calls and queues, not TCP. That network isolation should hold.

For your point about needing a safe API, a proxy is the classic pattern. You could run a separate, tightly-controlled container with network access that acts as a gateway. Your main agent container talks to it over a Docker user-defined bridge or even a Unix socket, which the proxy then forwards out. This creates a clear choke point and trust boundary.

I'd also sketch an attack tree for the proxy itself. What if the agent tries to poison the proxy's request? You need to validate and possibly sanitize the destination URL and parameters at that boundary.


Model it or leave it.


   
ReplyQuote
(@agent_log_watcher)
Eminent Member
Joined: 3 months ago
Posts: 19
 

The network isolation approach is correct, but you must also consider the logging telemetry from the containerized execution. The Docker daemon's audit trail (`dockerd` logs) and container-specific syscalls logged via `auditd` are critical for the verification you mentioned.

If an agent somehow escapes the `--network=none` constraint, you need a forensic record of the attempt. You should configure your container runtime to log all `connect` and `socket` syscalls for that container's PID namespace. This creates the data needed to trace any anomalous network activity back to the specific agent session or code execution event.

Your proxy gateway idea then needs its own independent, immutable audit log for all forwarding decisions, which should correlate with the main container's execution logs. Without that correlation, your choke point is less useful for post-incident analysis.


Log everything, trust nothing.


   
ReplyQuote
(@container_escape_hunter_tina)
Active Member
Joined: 3 months ago
Posts: 13
 

The `--network=none` approach should be fine for inter-agent chatter if they're all in the same Python process. They use in-memory message passing, not localhost sockets.

But you're right to worry about the "needs to fetch" case. A proxy is the standard move, but there's a subtler vector: a determined agent might try to abuse the proxy's own network to pivot. If your proxy container runs with host networking or overly permissive caps, it's game over. The gateway needs its own strict confinement, maybe even a seccomp-bpf profile that only allows `connect` to your approved internal service IPs.

Also, don't forget about DNS. `--network=none` blocks it, but if you use a bridge network for the proxy, an agent could still try to exfil via DNS tunneling. Block all outbound UDP 53 on the gateway except to your internal resolver.


Escape artist.


   
ReplyQuote
(@contrarian_coder)
Eminent Member
Joined: 3 months ago
Posts: 18
 

DNS tunneling, seriously? That's a bit of a stretch for an AutoGen agent that's just trying to execute a `requests.get()` it wrote itself. You're building a threat model for a nation-state actor, not a coding assistant that got a little too creative.

The bigger, dumber hole is that proxy's configuration. Everyone loves to sketch out these perfect seccomp profiles, but who's actually validating them at runtime? If you're letting the agent influence the destination URL at all, a bad allow-list regex is game over. It can just `connect` to `allowed-internal-service.local?exfil=data`. Good luck catching that in a log.

All this complexity for a problem that usually starts with someone setting `code_execution_config={"use_docker": false}` for convenience.


Reality is the only threat model that matters.


   
ReplyQuote
(@hobbyist_hardener_max)
Eminent Member
Joined: 3 months ago
Posts: 22
 

I get your point about DNS tunneling feeling overkill. But the threat model isn't about the agent's intent, it's about the code it's tricked into executing. A single `requests.get("http://evil.com")` is the obvious one, but a library it pulls in could attempt DNS as a fallback.

You're absolutely right about the proxy config being the weak link, though. An allow-list regex mistake is fatal. I've seen people try to use `*.internal-service.com` and forget to anchor the end, allowing `evil.internal-service.com.fake.net`. The safer pattern is an explicit list of literal IP:port combos in the proxy's routing table, no regex at all.

And yeah, `"use_docker": false` is the real root cause in most incidents 😅. That's why my Ansible playbook for our AutoGen nodes forces it to true and drops in a default seccomp profile.


Hardening is a hobby, not a job.


   
ReplyQuote