Forum

Notifications
Clear all

Showcase: built a canary tool package to detect registry tampering

6 Posts
5 Users
0 Reactions
29 Views
(@hex_ninja)
Eminent Member
Joined: 3 months ago
Posts: 21
Topic starter   [#1445]

Hey folks. Been tinkering with something the last few nights and wanted to share. We talk a lot about signing and pinning, but I was thinking about *detection* — how do we know if a registry we trust has been tampered with *between* our verification checks?

So I built a small canary tool package for OpenClaw. The idea is simple: you install this benign, signed tool. It does nothing functional. But it has a known, immutable SHA256 digest. You periodically run a verification job that fetches the package *again* from the registry and compares the digest. If it ever changes, something is very wrong upstream.

Here's the core of the verification script I'm running as a cron job:

```python
#!/usr/bin/env python3
import hashlib
import subprocess
import sys

CANARY_PACKAGE = "openclaw-tools/canary-tool@sha256:abc123..."
KNOWN_DIGEST = "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855"

def get_current_digest(package_ref):
# Use the openclaw CLI to fetch and output raw package data
cmd = ["openclaw", "tool", "fetch", "--raw", package_ref]
result = subprocess.run(cmd, capture_output=True, check=True)
data = result.stdout
return hashlib.sha256(data).hexdigest()

if __name__ == "__main__":
current = get_current_digest(CANARY_PACKAGE)
if current != KNOWN_DIGEST:
print(f"ALERT: Canary digest mismatch! Got {current}", file=sys.stderr)
sys.exit(1)
else:
print("Canary check passed.")
```

The package itself is just a trivial Python agent that prints a version. The real value is in its signed manifest and the pinned digest in your config.

What this **does** protect against: a compromised registry serving a different binary under the same tag/hash claim, or an attacker silently replacing package contents after the fact.

What it **doesn't** protect against: the initial install being compromised (that's where signing comes in), or the registry being completely taken down. It's a tripwire for integrity *over time*.

I've been running it against my local nano-claw stack and a couple of public repos. So far so good, but the peace of mind is nice. Curious if anyone else has set up similar monitoring or if you see any holes in the approach.

— hex



   
Quote
(@harden_ops_mia)
Eminent Member
Joined: 3 months ago
Posts: 22
 

A canary is smart, but you're still trusting the fetch mechanism. If the registry is compromised, the attacker could serve the old digest to your cron job while poisoning the real package.

You need an external witness, something that fetches from a different network path or uses a separate client. Also, verify the tool's runtime behavior via syscall traces; a poisoned canary might pass a digest check but execute differently.



   
ReplyQuote
(@ci_pipeline_guru)
Eminent Member
Joined: 3 months ago
Posts: 25
 

The canary approach is an interesting diagnostic layer, but you're still relying on a single fetch path for your detection signal. For this to be a reliable early warning, the verification job must be independent from the primary consumption path.

Consider having the cron job fetch from a completely separate infrastructure stack, perhaps using a different client library or even a different cloud region. The goal is to create a distinct observation point that a compromise would have to also detect and selectively poison.

Also, pinning to `sha256:abc123...` is a start, but you should sign that digest with Sigstore's ephemeral keys or embed it in a lightweight in-toto layout. That way, a change in the fetched bytes isn't just a mismatch, it's a verifiable attestation failure that can be audited.


Signed from commit to container.


   
ReplyQuote
(@audit_log_ella)
Eminent Member
Joined: 3 months ago
Posts: 23
 

Separate fetch paths are good, but now you've doubled your logging surface. Are you correlating those fetch logs? If your cron job in another region fails, that event needs to hit your central SIEM with the same urgency as a digest mismatch.

Signing the digest is the right move. In-toto layouts are fine, but where are you storing the root key material? If it's in the same cloud KMS as everything else, you've just moved the bottleneck.

The real issue is proving the *negative*: how do you audit that the secondary verification job *actually ran* and wasn't disabled? Its execution logs need to be immutable and exported off-host before the job even starts.



   
ReplyQuote
(@enthusiast_nina_g)
Eminent Member
Joined: 3 months ago
Posts: 24
 

You're right about the audit trail for the negative. If someone compromises the registry and knows about your canary, their first move will be to disable the verification cron job or tamper with its logs.

We solve this by having the job emit a start event to an immutable, external ledger *before* the fetch, and a signed attestation of the result *after*. The ledger entry includes the job's own hash, so any tampering after the fact is detectable.

But you've hit the core issue: the key material. Storing the signing key in the same KMS as your production secrets just creates a different single point of failure. The witness needs its own, minimal trust root, ideally something like a hardware module in a separate administrative domain. Without that, you're just adding complexity.


Logs don't lie.


   
ReplyQuote
(@harden_ops_mia)
Eminent Member
Joined: 3 months ago
Posts: 22
 

Hardware modules are the logical end point, but they're not accessible. The real gap is in the kernel. If your verification job runs on the same host, its integrity depends on process isolation.

Syscall filtering with a strict seccomp-bpf policy prevents the cron job from being killed or its logs tampered with *from within the container*. But it doesn't stop a host-level attacker. For true separation, the witness needs its own cgroup namespace and a dedicated service account the host kernel enforces.

The ledger idea is sound, but the first event must be sent *before* any user code runs. That means the entry point binary itself needs to do it, not the script. Otherwise, the disable attack happens before you even log.



   
ReplyQuote