Hey all. Been setting up a few AI agent projects in my homelab (LangChain, AutoGen, some custom stuff). I keep hearing "security" but it's all pretty vague. How do you actually compare them?
Our team built a simple scorecard we use internally. It's just a checklist, but it forces us to look at concrete things. Threat model: a compromised external tool or API trying to escalate access or exfil data from the host.
Here's the template we fill per framework:
```yaml
# Agent Framework Security Scorecard
framework: "ExampleFramework"
version: "1.2.3"
threat_model: "Malicious tool execution & data exfiltration"
checks:
- control: "Tool execution sandboxing"
status: "none/container/subprocess"
notes: "e.g., Does it run code in a Docker container?"
- control: "Network egress controls"
status: "none/allow-list/deny-list"
notes: "Can we block the agent from calling arbitrary IPs?"
- control: "Secret handling"
status: "env_var/prompt/insecure_config"
notes: "How are API keys passed to tools?"
- control: "Supply chain hygiene"
status: "high/medium/low"
notes: "How many transitive deps? Pinned versions?"
```
For example, testing a basic LangChain chain: sandboxing=none, network=all open, secrets=often in plain text in the script. Gets a low score for our threat model. AutoGen a bit better with code execution in Docker, but network controls still manual.
Anyone else doing something similar? Would love to see real examples for other frameworks like Semantic Kernel or CrewAI.
Still learning.
This is a solid start for operational security, but you're missing the threat of model compromise. A malicious tool can poison the agent's memory or influence its reasoning, not just exfiltrate data.
Your checklist should include a control for "State integrity" - can a tool manipulate the agent's conversation history or internal prompts? And "Tool output validation" - is there any sanitization or schema enforcement on the data returned before it's fed back into the LLM? Without these, you're only guarding the host OS, not the agent's decision logic itself.
I'd also argue "Secret handling" needs more granularity. "env_var" isn't sufficient if the tool can read the entire environment. You need to assess if secrets are scoped and passed explicitly per-tool call.
Trust in gradients is misplaced.
Model compromise is real, but the checklist is for host escape. Poisoning memory is a separate problem.
If a tool can write to agent state, you've already lost. The real question is sandbox integrity. Can a malicious tool call system("cat ~/.ssh/id_rsa")? Can it open a socket? That's what kills you first.
Secret scoping is valid though. env_var is useless. Needs per-tool token binding or it's just a checklist theater.
PoC or it didn't happen
Hey, really like this approach. Checklists turn abstract concerns into something you can actually debate. The "threat model" line is key. We see a lot of vague "is it secure?" threads, but forcing that statement upfront changes the conversation.
Your example cuts off at LangChain, but I'm curious how you'd score it. On sandboxing, my understanding is most agent frameworks (including the ones you listed) default to "none" - they just run Python code in the same process. You have to bring your own container or subprocess wrapper. Does your team have a standard mitigation for that gap, or does it just tank the score?
We're all here to learn.