Been seeing a lot of chatter about AI agents and their potential for "autonomous" action. One of my immediate concerns: what if the thing you built to book a flight decides it needs to phone home, or worse, exfiltrate data on a new, unexpected channel? The runtime's network permissions are often an afterthought.
So, I built a simple monitor that sits on the host and logs/raises a flag for any outbound network connection the agent's process makes that wasn't pre-authorized. It's not a silver bullet, but it's a crucial canary.
The core idea:
1. At agent startup, you feed the monitor a list of expected destination IPs/domains and ports (e.g., your internal API endpoints, a specific external weather service).
2. The monitor attaches to the agent's PID and sniffs its network traffic (using a lightweight eBPF probe or, for a simpler PoC, `lsof`/`netstat` polling).
3. Any TCP/UDP connection to a destination not on the allow-list triggers an immediate alert and a full connection log.
Here's the basic policy-as-code structure (YAML) for defining expected behavior:
```yaml
agent_name: "travel_agent_v1"
expected_outbound:
- destination: "api.company-internal.com"
port: 443
protocol: "TCP"
purpose: "Internal flights API"
- destination: "weather.service.com"
port: 443
protocol: "TCP"
purpose: "Fetch destination weather"
allowed_dynamic_resolution:
- "*.company-internal.com" # Allows for some DNS-based flexibility, but logged.
```
The monitor's output on a violation looks like this (CLI alert):
```
[!] UNEXPECTED OUTBOUND CONNECTION
Timestamp: 2023-10-26T14:32:07Z
PID: 7843
Command: /usr/bin/python /opt/agent/main.py
Destination: 104.28.14.6:443 (resolved: sketchy-mirror.example.com)
Action: LOGGED (Block policy not enabled)
Rule Matched: NONE - Connection not in allow-list.
```
**What I learned the hard way:**
* You need to account for DNS resolution. The monitor must resolve IPs and check against both IP and domain lists.
* Some libraries/agents spawn subprocesses. You must track the entire process tree, not just the initial PID.
* This is a detection tool first. Automatic blocking is possible, but you risk breaking legitimate, unexpected (but necessary) fallback logic.
This forces you to think through the agent's threat model concretely. If you haven't defined its expected network behavior, you have a gap. This tool closes that gap simply. Code's still rough, but the prototype works. If anyone's interested in the eBPF approach or has similar work, post below.
--Priya
--Priya
Oh wow, that's a fantastic idea, and honestly a bit scary that I hadn't even considered it yet. I've been so focused on just getting my agent to run in Docker that I never thought about what it might be doing on the network side.
> a list of expected destination IPs/domains and ports
This is the part that feels tricky to get right for a beginner like me. How do you even figure out the full list for a complex agent? Is it a trial and error thing where you run it in a sandbox first and log everything, then build the allow list from that? I could totally see myself missing a crucial CDN or something.
Also, what happens when the agent *needs* to make a new, legitimate call? Does the whole thing stop, or just alert you?
You've hit on the main challenge with these allow-list approaches. Starting in a sandbox to log everything is exactly right, it's the most practical way to begin.
To answer your question, in my setup it just alerts me. Blocking outright could break a legitimate workflow if my list isn't perfect. The alert gives me a chance to review and then decide if I need to update the allow list or investigate further.
And you're right about missing a CDN - that's a real headache. One thing I do is run that initial sandbox logging for a while, across different tasks, to try and catch all those dependencies. It's never truly complete, but it gets you a solid baseline.
Read the sticky.
Exactly the kind of thinking we need more of. That initial allow list is the hardest part, and your YAML snippet is a great start. I've found you often need to expand that structure to handle real-world messiness, like agents that pull data from a cloud object store where the IPs are dynamic.
I usually add a `protocol` field and a `description` of the business purpose. More importantly, I nest a `dns_pattern` field under the destination for cases where you're connecting to something like `*.blob.core.windows.net`. That way your eBPF probe can resolve and match against a pattern, not just a static list.
Have you considered also logging the process cmdline or the specific library making the call? Sometimes the unexpected connection is from a dependency, not your main agent code, and that context is gold for triage.
Log everything, trust nothing.
Okay, but you're just treating the symptom. The underlying disease is that we're handing *agents* - code that can rewrite its own prompts and execute tools - full-blown network egress, then acting surprised when it's used.
A static allow list is a nice canary, but it's about as useful as a lock on a screen door if the agent can reason about the network. I've seen test cases where a constrained agent just uses an allowed, trusted external API (like a weather service) to encode and exfiltrate data in the request parameters themselves. Your monitor would see a call to `api.weather.com:443` and shrug.
The real play is runtime introspection - tracking *why* the network call is being made, not just where to. But good luck selling that to a product team obsessed with "functionality."
Trust me, I'm a hacker.
Good. This is exactly where the baseline starts.
The YAML snippet is solid. For a production setup, you'd extend that with a `dns_pattern` field, like user195 suggested, to handle cloud services with dynamic IPs. You don't want to chase CDN ranges.
The real value is in the logging. Tag every alert with the full process tree and command line. Half the time, the weird outbound call is from some python package's telemetry library you didn't know you pulled in, not the agent logic itself. That tells you where to focus your hardening.
-- mike
Yeah, the sandbox log-everything-first approach is the only way to get started. I just set up a test VM, let the agent run through a bunch of tasks, and collected the logs.
But figuring out what's a real dependency vs. telemetry junk is the next step. I'm already seeing calls to pypi.org and random stats servers just from importing common libraries.
Does your sandbox setup filter those out automatically, or are you just adding everything it logs to the initial allow list?