Forum

Notifications
Clear all

Walkthrough: Implementing a mandatory approval step for all agent tool calls.

5 Posts
5 Users
0 Reactions
31 Views
(@agent_rookie_mia)
Eminent Member
Joined: 3 months ago
Posts: 23
Topic starter   [#1298]

Hi everyone. I've been trying to set up a local LLM agent on my Raspberry Pi, but I keep getting nervous about letting it run tools on its own. I read a lot here about sandboxing and agent safety, which is great, but sometimes the terms get a bit heavy.

I want a simple, mandatory "yes/no" approval step for *any* tool call before it runs. Think "Agent wants to run `send_email`. Allow?" in my terminal. I'm using a basic Python setup with LangChain.

My threat model is basically: I don't want the agent to accidentally or purposefully modify files, send data out, or execute system commands without me seeing it first. I'm less worried about sophisticated escapes and more about simple mistakes or prompt hijacks.

I managed to override the tool execution with a wrapper that prints the request and waits for my input. It seems to work, but I'm worried I'm missing something obvious. Has anyone else done this? Is there a common pattern or a security pitfall I should look for?



   
Quote
(@policy_writer_axel)
Eminent Member
Joined: 3 months ago
Posts: 17
 

That's a solid, pragmatic approach for your threat model. The big gap people miss is state. Your approval prompt shows "send_email" but does it show the fully rendered content? If the agent decides the content in a previous step, you might be approving an empty shell while a malicious payload is already queued up.

Also, what about speed? If you're stepping through fifty tool calls to write a report, you'll likely start mashing 'y' without reading. Then it's just a noisy NOP gate.

Consider logging every proposed call, approved or not, with a hash of the arguments. If something goes sideways, at least you've got a forensic trail to see what you actually approved versus what it wanted to do.


audit what matters


   
ReplyQuote
(@home_lab_builder_sam)
Eminent Member
Joined: 3 months ago
Posts: 29
 

You're spot on about the state problem. My first attempt at an approval layer had that exact flaw, the agent would generate the email body in thought, then just pass a placeholder like `body=""` to the tool. I had to start serializing the entire agent scratchpad and running a quick diff to see what changed in the last step before presenting the approval. It gets messy fast.

The logging with argument hashes is a great idea. I ended up dumping every interaction, approved or not, into a sqlite db with a timestamp and a hash of the tool args. Then you can at least reconstruct the sequence of *intent*, which is super helpful when you're trying to figure out if a weird output came from a bad approval or a sneaky intermediate step.

But yeah, the alert fatigue is real. After the twentieth "allow read_file?" prompt, you stop reading the args. It feels safer, but it's mostly theatre unless you're truly paranoid for every single call. Maybe that's the point though?


Still learning, still breaking things.


   
ReplyQuote
(@cloud_sec_ken)
Eminent Member
Joined: 3 months ago
Posts: 22
 

Good instinct on the wrapper. The immediate pitfall is what user331 mentioned - you're likely only seeing the tool call signature, not the actual data. If your `send_email` tool just takes a `body` argument, you need to ensure that argument is fully populated and shown to you, not just a variable name the agent prepared earlier.

You also need to intercept *every* tool. LangChain's base classes are leaky; if you wrapped the wrong one, some internal tool might slip through. I'd test by trying to get the agent to do something you'd definitely deny, like `shutdown -h now` via a shell tool, and see if your prompt catches it.

Logging is non-negotiable. Print to screen, but also write to a file with a timestamp. Otherwise, you'll mash 'y' on the fiftieth call and have no record of what you just approved.


- ken


   
ReplyQuote
(@red_team_rookie)
Eminent Member
Joined: 3 months ago
Posts: 22
 

Oh, the logging part is something I almost missed! I set up the print-to-screen part, but writing to a file makes so much sense. I can see myself just hitting 'y' after a while and then having no idea what happened.

How do you handle the log format? Just raw JSON of the tool call, or something more readable? I'm worried about the log file getting huge fast if every thought is logged.



   
ReplyQuote