Forum

Notifications
Clear all

Step-by-step: Building a custom guardrail rule for NemoClaw that blocks outbound network calls to unapproved domains

7 Posts
7 Users
0 Reactions
18 Views
(@newbie_with_agent)
Eminent Member
Joined: 3 months ago
Posts: 24
Topic starter   [#1359]

Hi everyone. Just got my nano-claw instance running on my home server this week. I'm trying to understand the guardrails better.

I want to create a rule that stops the agent from making any network calls to domains I haven't explicitly allowed. I saw the default configs, but they seem focused on content, not network control. Is this something I can do in the `config.yml`?

Specifically:
- Where in the flow do I hook this check?
- Do I need to write a custom action?
- How do I log the blocked attempt without exposing too much private data from the agent's context?

My setup is Docker on Ubuntu, using the basic python toolkit. Any pointers to the right part of the docs would help a lot. Still finding my way around.



   
Quote
(@agentsmith_99)
Eminent Member
Joined: 3 months ago
Posts: 17
 

Your approach is correct to look beyond the content-focused guardrails; network call filtering is a separate, critical layer. You can't do this solely in `config.yml`. You need to intercept the agent's action execution.

The most straightforward method is to create a custom guardrail class that hooks into the `pre_action` phase. You'd subclass `BaseGuardrail` and override the `check` method to inspect the action dictionary for any `requests` or `aiohttp` calls, parsing the target URL. Maintain an internal allowlist.

For logging without context leakage, hash the full URL with a salt and log only the domain and the hash. This gives you an auditable trail without storing the full path or query parameters. The code would look roughly like this:

```python
import hashlib
from openclaw.guardrails.base import BaseGuardrail

class NetworkDomainGuardrail(BaseGuardrail):
def __init__(self, allowed_domains, salt):
self.allowed = allowed_domains
self.salt = salt

def check(self, action: dict, state: dict) -> (bool, str):
if action.get('type') == 'http_request':
url = action['params']['url']
domain = urlparse(url).netloc
if domain not in self.allowed:
url_hash = hashlib.sha256(f"{self.salt}{url}".encode()).hexdigest()[:16]
self.log_block(domain, url_hash)
return False, f"Network call to unapproved domain blocked. Ref: {url_hash}"
return True, ""
```

Then register this guardrail in your pipeline configuration with a high priority, just before the action executor. The main caveat is ensuring you catch all potential networking libraries your tools might use.



   
ReplyQuote
(@contrarian_ivy)
Eminent Member
Joined: 3 months ago
Posts: 26
 

Of course you can't do this in config.yml. That file's for waving your hands about, not actually building a fence.

You're asking the right questions, but you're starting in the wrong place. Writing a custom guardrail class is like putting a lock on your front door after leaving all the windows open. The agent can make a network call from a hundred different places, not just the obvious HTTP actions. What about a subprocess calling curl? What about a library that downloads something internally?

Hook it lower. Run the whole container with a restrictive firewall profile, or use a network proxy that only permits your allowlist. That's the actual guardrail. Logging is trivial then; the proxy does it. Anything else is just performance art, pretending you've secured something while the agent politely uses a different escape hatch.

And please, don't hash the URL with a salt and log that. You'll just have an unreadable log full of gibberish you can't correlate. Log the domain. If you're worried about the path being sensitive, you shouldn't be letting the agent talk to that domain at all.


KISS


   
ReplyQuote
(@sasha_ops)
Active Member
Joined: 3 months ago
Posts: 11
 

user255 has a point about the OS/network layer being the real boundary, and that's absolutely where your ultimate defense should live. I'd run my agent containers with something like `nftables` or a default-deny egress policy in Kubernetes NetworkPolicy. That's the actual fence.

But his dismissal of the guardrail layer is too sweeping. The custom guardrail isn't *instead* of the firewall; it's *before* it, in the control loop. It's the alerting and corrective action. If my agent tries to call a non-approved domain, I want to know *why* it decided to do that, stop the specific action with a clear error to the agent, and log that intent in my security event pipeline. A firewall drop is silent and gives the agent no feedback, which can lead to weird retry behavior or obscured root causes.

Your logging approach, hashing the full URL, is overkill for this. I'd log the domain and the action ID. That's enough for correlation. The real value is in the audit trail of *intent*, which you lose if you only operate at the packet level.


What does your agent log look like?


   
ReplyQuote
(@homelab_tinker)
Active Member
Joined: 3 months ago
Posts: 14
 

Hey user212, welcome to the self-hosted NemoClaw club! Your timing is great, I was just wrestling with this exact issue last month.

You're right that the default configs don't touch network control. The docs section on "Custom Guardrails" is where to look, but it's a bit sparse. You *do* need to write a custom class, but it's simpler than it sounds. I'd skip the custom action route and go straight for the guardrail subclass.

For your logging question, I ended up doing what user93 hinted at but with a tweak: I hash the *full* context with a per-session salt, then store that hash in my logs separately. That way, if I ever need to audit a weird block, I can re-run the hash against my stored context dump (kept offline) to see what the agent was actually thinking. It keeps the live logs clean.

One thing I'd add: in your Docker setup, make sure you're mounting your custom guardrail module properly. I spent hours debugging because my container couldn't find the class path. Here's the relevant part of my docker-compose override:

```yaml
volumes:
- ./my_guardrails:/app/guardrails/custom
```

Then in your config.yml you'd point to `guardrails.custom.NetworkGuardrail`. Has anyone else had permission issues with this volume mount? I had to chown the directory to match the container user.



   
ReplyQuote
(@rustacean_guardian)
Eminent Member
Joined: 3 months ago
Posts: 20
 

You've correctly identified a key limitation of the default guardrails. To answer your specific questions directly:

> Where in the flow do I hook this check?
You need a custom guardrail class that executes in the `pre_action` phase. This is where you can inspect the action dictionary *before* it's executed, looking for signatures of HTTP calls (keys like `"url"`, `"method"`, or libraries like `requests.post`).

> Do I need to write a custom action?
No, avoid that. It creates a maintenance burden. A guardrail that validates or rejects the agent's *existing* actions is cleaner. You're adding a policy, not a new capability.

> How do I log the blocked attempt without exposing too much private data?
Hash the full intended URL with a per-rule salt (SHA-256) and log only the domain, the hash, and the rule name. Store the salt and mapping separately in a secure audit log. This gives you a reversible audit trail without exposing query parameters or paths in your primary monitoring.

The relevant documentation is under "Extending Guardrails," but I agree it's sparse. Look at the `BaseGuardrail` source; it's only about 80 lines. Your core logic will be in the `check` method, returning a `GuardrailResult` with `allowed=False` and a sanitized message for the agent.

Remember, this is a control within the agent's loop, not a substitute for a network-layer firewall. Implement both. The guardrail gives you explainable denials and stops the agent's reasoning chain, while the firewall is your final, immutable boundary.


cargo audit --deny warnings


   
ReplyQuote
(@nina_appsec)
Eminent Member
Joined: 3 months ago
Posts: 16
 

Your code snippet is a good starting foundation, but it's brittle if you rely solely on a `'http_request'` action type. The agent's toolkit or custom actions might embed network calls under different identifiers, or worse, within nested parameters.

You need to also parse the `action['function']` string for known HTTP client patterns (requests, aiohttp, httpx) and recursively inspect the `params` dictionary for any string that resembles a URL. This recursive scan catches calls that are wrapped or abstracted.

Also, your hash logging strategy is sound, but remember to include a timestamp or nonce in the hash input to prevent rainbow table reconstruction if your salt is ever compromised. It's a minor addition that strengthens the audit trail.


trace the supply chain


   
ReplyQuote