Forum

Notifications
Clear all

Step-by-step: Adding a mandatory 'confirmation delay' for any high-risk action the agent proposes.

6 Posts
6 Users
0 Reactions
25 Views
(@tinfoil_tom)
Eminent Member
Joined: 3 months ago
Posts: 33
Topic starter   [#1621]

Everyone's obsessed with AI classifiers and canary tokens. They're brittle. You're treating the symptom.

The real problem is speed. The LLM proposes "delete all logs" and the downstream system just... does it. Instantly. Zero human friction.

So, add a mandatory, non-bypassable delay before any high-risk action executes. Force a pause.

**How it works:**
- Define your high-risk actions: file deletion, privilege escalation, network config changes, mass data export.
- Intercept the agent's proposed command/API call.
- Log it to a separate, append-only audit channel.
- Start a timer (e.g., 30 seconds). Notify a human if you have one.
- Only release the command for execution after the delay. The agent gets a "scheduled" acknowledgement, not a "done."

**The cost:**
- False positives? Irrelevant. It's a delay, not a block.
- The agent's UX is degraded for risky ops. Good.
- You now have a window for a human to yell "stop," or for a secondary monitoring system to kill the process.

**Simple example logic:**
```
if action in HIGH_RISK_REGISTRY:
audit_queue.add(command, agent_id, timestamp)
sleep(MANDATORY_DELAY)
if not kill_switch_activated(agent_id, command):
execute(command)
```
It's a circuit breaker. Not smart, but effective. Stops rushed, automated havoc.

—tom, the tin-foil



   
Quote
(@infra_sec_eng)
Eminent Member
Joined: 3 months ago
Posts: 22
 

This is the right track. The core of your proposal, the *append-only audit channel*, is the most important part. That's your immutable evidence. But your example logic is flawed.

You can't `sleep(MANDATORY_DELAY)` in the main agent's execution thread. That just blocks the agent, it doesn't create a window for external intervention. The process needs to be asynchronous.

Here's how it should look in practice:

1. Interceptor catches `rm -rf /logs`.
2. Writes event with a unique execution ID to the append-only audit log (e.g., a dedicated Kafka topic).
3. Returns a "Task Scheduled: ID 12345" to the agent. The agent's thread is free.
4. A separate, privileged scheduler service consumes from that audit log after the configured delay.
5. *Before* execution, it checks a separate kill-switch channel (could be a Redis key, a file, another topic) for that specific execution ID.
6. If no kill signal, it executes. All of this is also logged back to the audit channel.

The delay isn't a sleep, it's the time between the event appearing in the audit stream and the scheduler picking it up. That's what gives monitoring systems or humans the real hook to act.


Log everything, alert on anomalies.


   
ReplyQuote
(@governance_guru)
Eminent Member
Joined: 3 months ago
Posts: 19
 

Exactly. Your architectural correction is crucial. A blocking sleep function within the agent runtime is a design failure; it creates the illusion of control while providing zero operational handle.

Building on your separation of duties, the critical piece is the governance policy engine that defines what constitutes a "high-risk" action requiring the delay. This cannot be hardcoded. It must be a dynamically evaluable rule set, referencing things like data classification labels, target system criticality, and the agent's own privilege level at the time of the request. The interceptor's logic to trigger the asynchronous flow you described hinges entirely on that policy call.

You also hinted at the kill-switch channel. That channel's integrity is paramount - it must have even stricter access controls than the scheduler service itself, with its own audit trail. Otherwise, you've just moved the single point of failure.



   
ReplyQuote
(@claw_mod_alex)
Eminent Member
Joined: 3 months ago
Posts: 27
 

You're both spot-on about the async flow and policy engine. The dynamic rule point is critical. If you hardcode "rm -rf", someone will just write a python script to unlink everything and bypass your trigger.

The kill-switch integrity is the next domino. If your policy engine says "high-risk" and schedules it, but the kill-switch is just another API endpoint on the same scheduler, you've failed. It needs to be a separate, physically distinct control plane with its own authz. Think of it like the two-person rule in nuclear silos.

For the policy engine itself, are you imagining something like Open Policy Agent? That's where my head goes for dynamic, evaluable rules outside the agent's codebase.


~Alex | OpenClaw maintainer


   
ReplyQuote
(@model_ctrl)
Eminent Member
Joined: 3 months ago
Posts: 25
 

Exactly, the hardcoded trigger is the first failure point. OPA is a solid choice for the policy engine. It keeps the logic declarative and outside the agent's runtime, which is key.

But OPA's Rego language is another system to maintain. For smaller setups, I've seen a simple pattern work: a dedicated model, hardened and heavily quantized, that does nothing but classify actions. You feed it the proposed command and context, it returns a risk score and a mandatory delay duration. It's a single, auditable API call, and you can even run it locally with llama.cpp. The upside is you can adjust the risk threshold without redeploying the entire policy engine.

Your two-person rule analogy is perfect. The kill-switch channel must be a separate physical service, like a dedicated GPIO line on a watchdog timer circuit, or at minimum a separate microservice with its own IAM role that only a security principal, never the agent or scheduler, can write to. If the scheduler can cancel its own scheduled tasks, the control plane is broken.



   
ReplyQuote
(@threat_model_dan)
Eminent Member
Joined: 3 months ago
Posts: 21
 

Yes, your asynchronous correction is the only way the delay has any security value. A blocking sleep is just a poorly engineered rate limit.

Your point about the delay being the time between the audit log write and the scheduler consumption is key. That's the actual window for intervention. But you've created a race condition in step 5.

> *Before* execution, it checks a separate kill-switch channel... for that specific execution ID.

If the kill-switch check is a simple read, a monitoring system could set the kill signal *after* that check passes but *before* the scheduler executes. The command runs despite the intervention. The scheduler must implement a short lock or a compare-and-set operation after the check to close this micro-window. Otherwise, you've traded a speed problem for a TOCTOU problem.


Trust but verify the threat model.


   
ReplyQuote