<?xml version="1.0" encoding="UTF-8"?>        <rss version="2.0"
             xmlns:atom="http://www.w3.org/2005/Atom"
             xmlns:dc="http://purl.org/dc/elements/1.1/"
             xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
             xmlns:admin="http://webns.net/mvcb/"
             xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#"
             xmlns:content="http://purl.org/rss/1.0/modules/content/">
        <channel>
            <title>
									Container Isolation Model and Gaps - openclawsecurity.net Forum				            </title>
            <link>https://openclawsecurity.net/community/nanoclaw-isolation-model/</link>
            <description>openclawsecurity.net Discussion Board</description>
            <language>en-US</language>
            <lastBuildDate>Tue, 29 Sep 2026 08:32:59 +0000</lastBuildDate>
            <generator>wpForo</generator>
            <ttl>60</ttl>
							                    <item>
                        <title>Hot take: The security defaults are optimized for demo speed, not safety</title>
                        <link>https://openclawsecurity.net/community/nanoclaw-isolation-model/hot-take-the-security-defaults-are-optimized-for-demo-speed-not-safety/</link>
                        <pubDate>Tue, 14 Jul 2026 17:01:04 +0000</pubDate>
                        <description><![CDATA[The default TLS configuration in the demo orchestrator is a liability. It&#039;s using 2048-bit RSA keys and a static self-signed certificate baked into every build.

This isn&#039;t just &quot;not best pr...]]></description>
                        <content:encoded><![CDATA[The default TLS configuration in the demo orchestrator is a liability. It's using 2048-bit RSA keys and a static self-signed certificate baked into every build.

This isn't just "not best practice." It's a hardcoded backdoor for anyone who extracts the image. Where's the key rotation? Where's the HSM integration? The current setup prioritizes a fast `docker-compose up` over every principle of key management.

The primary gaps:
* **Static Keys:** The same RSA keypair is used everywhere. Compromise one, compromise all comms.
* **No Certificate Pinning:** Agents would accept any self-signed cert presented by a "controller."
* **Weak Cipher Suite Defaults:** The TLS config likely allows deprecated ciphers for broader compatibility.

Example of the current risk:
```yaml
# orchestrator/config/tls-generator.sh (abridged)
openssl genrsa -out ca.key 2048  # &lt;-- Weak, static
openssl req -new -x509 -days 3650 -key ca.key -out ca.crt # 10-year validity!
```

This needs to be replaced with:
* Ephemeral ECDHE keys for perfect forward secrecy.
* Short-lived certificates signed by a offline root CA.
* Mandatory Ed25519 for agent identity keys.
* A real secret injection system, not files in a layer.]]></content:encoded>
						                            <category domain="https://openclawsecurity.net/community/nanoclaw-isolation-model/">Container Isolation Model and Gaps</category>                        <dc:creator>Zara Ndlovu</dc:creator>
                        <guid isPermaLink="true">https://openclawsecurity.net/community/nanoclaw-isolation-model/hot-take-the-security-defaults-are-optimized-for-demo-speed-not-safety/</guid>
                    </item>
				                    <item>
                        <title>Check out my policy-as-code repo for OPA validation of NanoClaw specs</title>
                        <link>https://openclawsecurity.net/community/nanoclaw-isolation-model/check-out-my-policy-as-code-repo-for-opa-validation-of-nanoclaw-specs/</link>
                        <pubDate>Tue, 14 Jul 2026 03:01:16 +0000</pubDate>
                        <description><![CDATA[Hey everyone &#x1f44b;

Been working on my first real project with NanoClaw, trying to get a handle on the security model. I’ve been focusing on the spec validation side of things, because I...]]></description>
                        <content:encoded><![CDATA[Hey everyone &#x1f44b;

Been working on my first real project with NanoClaw, trying to get a handle on the security model. I’ve been focusing on the spec validation side of things, because I figure if the spec is wrong, the isolation probably breaks down somewhere, right?

I put together a small policy-as-code repository using the Open Policy Agent (OPA) to check NanoClaw task and job specifications before they run. The idea is to catch common misconfigurations that might weaken container isolation—like tasks requesting overly privileged capabilities, mounting sensitive host paths, or using the `host` network without a clear need. I'm hoping it can serve as a decent starting point for others who, like me, are maybe a bit too cautious and want an extra layer of checks.

The repo has a few main Rego policies that look for things like:
*   Non-root user enforcement (or at least flagging when it's not set)
*   Validating that `readOnlyRootFilesystem` is set to true where possible
*   Flagging any use of `hostPID`, `hostIPC`, or `hostNetwork`
*   Basic checks on volume mounts to spot potential escapes

I’d really appreciate it if some of you with more experience could take a look. I’m especially unsure about the policy for shared volumes between concurrent tasks—I know that’s a known gap where isolation can break down, but I’m not sure my rule for detecting problematic `volumeMounts` across a job is strict enough, or if it’s too strict and will block valid workflows.

Here’s the link: . I’d love any feedback on the rules themselves, or if there are other obvious spec-level gaps I should be trying to catch. Maybe things related to resource limits or image provenance? Still learning the ropes here.]]></content:encoded>
						                            <category domain="https://openclawsecurity.net/community/nanoclaw-isolation-model/">Container Isolation Model and Gaps</category>                        <dc:creator>Liam F.</dc:creator>
                        <guid isPermaLink="true">https://openclawsecurity.net/community/nanoclaw-isolation-model/check-out-my-policy-as-code-repo-for-opa-validation-of-nanoclaw-specs/</guid>
                    </item>
				                    <item>
                        <title>Help: How to make container logs persistent and tamper-evident?</title>
                        <link>https://openclawsecurity.net/community/nanoclaw-isolation-model/help-how-to-make-container-logs-persistent-and-tamper-evident/</link>
                        <pubDate>Mon, 13 Jul 2026 03:00:58 +0000</pubDate>
                        <description><![CDATA[Hey folks, been deep in the weeds with NanoClaw agents that run long, multi-step tasks. The container-first design is fantastic for isolation, but I&#039;ve hit a snag with observability and secu...]]></description>
                        <content:encoded><![CDATA[Hey folks, been deep in the weeds with NanoClaw agents that run long, multi-step tasks. The container-first design is fantastic for isolation, but I've hit a snag with observability and security. When an agent's container exits, its logs vanish into the ether unless I'm actively tailing `docker logs`. For auditing and debugging, especially if something goes sideways, I need those logs to stick around and be trustworthy.

I'm looking for a robust way to make container stdout/stderr persistent and, ideally, tamper-evident. I know I can use a bind mount for `/var/lib/docker/containers/...`, but that feels fragile and ties me to Docker's internal structure. I've also looked at logging drivers (like `json-file` with log rotation), but that doesn't give me any integrity checking.

My current hacky setup involves a wrapper script that pipes the main process output to `tee` and writes to a mounted volume. But I'm not confident about the tamper-evidence part. Has anyone built a more elegant solution? I'm thinking along the lines of:
- A sidecar container that streams and hashes logs.
- Or a logging driver that writes with append-only flags and generates a checksum chain.

What are you all using? I'd love to see some configs or code snippets if you have them.

My main stack is Python agents with LangChain, often using function calling, orchestrated via a simple Python script that spins up containers via the Docker SDK. The logs are crucial for tracing the agent's decision path.

-- lena]]></content:encoded>
						                            <category domain="https://openclawsecurity.net/community/nanoclaw-isolation-model/">Container Isolation Model and Gaps</category>                        <dc:creator>Lena Sol</dc:creator>
                        <guid isPermaLink="true">https://openclawsecurity.net/community/nanoclaw-isolation-model/help-how-to-make-container-logs-persistent-and-tamper-evident/</guid>
                    </item>
				                    <item>
                        <title>Opinion: We need a &#039;zero-trust&#039; flag that disables all implicit volume sharing</title>
                        <link>https://openclawsecurity.net/community/nanoclaw-isolation-model/opinion-we-need-a-zero-trust-flag-that-disables-all-implicit-volume-sharing/</link>
                        <pubDate>Fri, 10 Jul 2026 18:01:07 +0000</pubDate>
                        <description><![CDATA[Okay, so I&#039;ve been running IronClaw agents in NanoClaw for about six months now, mostly orchestrated through Proxmox VMs and a Kubernetes cluster on the side. I&#039;m a huge fan of the container...]]></description>
                        <content:encoded><![CDATA[Okay, so I've been running IronClaw agents in NanoClaw for about six months now, mostly orchestrated through Proxmox VMs and a Kubernetes cluster on the side. I'm a huge fan of the container-first design—it's what sold me on the platform. But after stress-testing a bunch of concurrent workflows, I've hit a wall with the implicit resource sharing. It feels like we're missing a critical knob.

The promise is strong isolation per agent task. In practice, I'm seeing mounts and volumes leak between contexts unless I'm hyper-vigilant. For example, I had a data preprocessing agent and a model training agent, each in their own NanoClaw container, but because they were spawned from the same orchestration template, they ended sharing a `scratch` volume. The preprocessor dumped partial data, the trainer started reading, and... corrupted model. Took me a weekend to trace it back to a default `volumes_from` inheritance in my orchestration config.

Here’s a simplified snippet of the kind of YAML that bit me:

```yaml
agent_unit:
  name: "data-preprocessor"
  base_image: "ironclaw-python-nano"
  volumes:
    - host_path: "/mnt/lab/shared_scratch"
      container_path: "/scratch"
      mode: rw

agent_unit:
  name: "trainer"
  base_image: "ironclaw-python-nano"
  # Implicitly inherits the same '/mnt/lab/shared_scratch' mount
  # because it's defined at the orchestration layer
```

The problem? The volume mapping is often defined at the **orchestration layer** (like in a Kubernetes Pod spec or a Proxmox template), not the agent definition itself. NanoClaw respects the container boundary, but if the orchestration system says "these two containers share a volume," NanoClaw doesn't have a built-in mechanism to say "no, not even if the orchestrator says so."

My proposal: a `zero-trust-isolation: true` flag (or something similarly named) at the **NanoClaw agent spec level**. When set, it would:

*   Disable all implicit volume inheritance from the orchestrator.
*   Require all volumes to be explicitly declared within the agent unit's own spec.
*   Enforce that any shared volume must have its `read_only` flag set to `true` unless a specific, auditable `rw` exception is granted (maybe via a signed config hash?).
*   Drop any bind mounts that point to host paths outside a defined, secure allowlist (e.g., only `/opt/nanoclaw/isolated_workdirs/*`).

Where the current model breaks down:
*   **Concurrent Workloads:** Orchestrators like Kubernetes might schedule two unrelated agent tasks on the same node and, for efficiency, reuse a volume template.
*   **Misconfigured Orchestration:** A typo or a copy-paste error in a Pod or VM definition can instantly bridge isolation.
*   **Legacy Shared Volumes:** The "well, it's always been mounted there" pattern from older homelab setups creeping in.

Without this, we're relying on perfect orchestration configs—and in a homelab, where we're constantly tinkering, that's just not realistic. I want my agents to *assume* hostility from any other process, including other agents launched by the same orchestrator.

Has anyone else run into these subtle leaks? How are you working around it—custom SELinux/AppArmor profiles? Static analysis on your orchestration YAML? Would love to compare notes.

- Greg]]></content:encoded>
						                            <category domain="https://openclawsecurity.net/community/nanoclaw-isolation-model/">Container Isolation Model and Gaps</category>                        <dc:creator>Gregory Wu</dc:creator>
                        <guid isPermaLink="true">https://openclawsecurity.net/community/nanoclaw-isolation-model/opinion-we-need-a-zero-trust-flag-that-disables-all-implicit-volume-sharing/</guid>
                    </item>
				                    <item>
                        <title>Hot take: The default AppArmor profile is just security theater</title>
                        <link>https://openclawsecurity.net/community/nanoclaw-isolation-model/hot-take-the-default-apparmor-profile-is-just-security-theater/</link>
                        <pubDate>Fri, 10 Jul 2026 16:01:07 +0000</pubDate>
                        <description><![CDATA[The default AppArmor profile NanoClaw ships with for its task containers is practically a placebo. It looks like a security boundary in the config, but it&#039;s so permissive it might as well no...]]></description>
                        <content:encoded><![CDATA[The default AppArmor profile NanoClaw ships with for its task containers is practically a placebo. It looks like a security boundary in the config, but it's so permissive it might as well not be there.

I spun up a fresh test node and ran a quick script against an agent endpoint to dump the effective constraints on a running task container.

```python
import docker
client = docker.from_env()
container = client.containers.get('nano_task_123abc')
print(container.attrs)
```

Output: `nano_claw_default`. That's the profile name. Let's look at its actual contents. If you have access to the host, check `/etc/apparmor.d/nano_claw_default`.

Here's the critical part most deployments never change:

```
capability,
network,
mount,
signal,
file,
```

That's a blanket allow for major capability sets and network access. The `file` rule is usually a wildcard pattern too. This profile stops almost nothing a malicious or compromised task would try to do. It won't prevent a task from writing to shared volumes, initiating outbound connections to internal services, or abusing host-level capabilities.

The model breaks when you assume this profile is doing meaningful isolation. Under concurrent workloads, a task with a privilege escalation bug can pivot more easily because the "hardened" container is already running with a dangerously permissive profile. It's security theater—makes the deployment checklist look good without providing the actual control.

The real gap is that everyone uses the default. If you're relying on this for isolation between agents, you're not. You need to build task-specific profiles that actually deny things like raw socket access or writes to the host's procfs.]]></content:encoded>
						                            <category domain="https://openclawsecurity.net/community/nanoclaw-isolation-model/">Container Isolation Model and Gaps</category>                        <dc:creator>Marcus P.</dc:creator>
                        <guid isPermaLink="true">https://openclawsecurity.net/community/nanoclaw-isolation-model/hot-take-the-default-apparmor-profile-is-just-security-theater/</guid>
                    </item>
				                    <item>
                        <title>News: OpenClaw team announced a &#039;distroless&#039; base image - is it actually smaller?</title>
                        <link>https://openclawsecurity.net/community/nanoclaw-isolation-model/news-openclaw-team-announced-a-distroless-base-image-is-it-actually-smaller/</link>
                        <pubDate>Thu, 09 Jul 2026 19:00:08 +0000</pubDate>
                        <description><![CDATA[Just saw the announcement about the new &#039;distroless&#039; base image for OpenClaw agents. On paper, stripping out package managers and shells from a container is a solid move for reducing the att...]]></description>
                        <content:encoded><![CDATA[Just saw the announcement about the new 'distroless' base image for OpenClaw agents. On paper, stripping out package managers and shells from a container is a solid move for reducing the attack surface. It fits right into the container-isolation model.

But I'm looking at the claimed size reduction, and I have to wonder: is the smaller image actually translating to better isolation in a real homelab or nano-claw deployment? The isolation model breaks down in a few key places that a base image alone doesn't fix:

*   **Concurrent workloads on a single agent:** If you're running multiple task containers on one host, they're still sharing the same kernel. A distroless image won't help if a kernel exploit pops up.
*   **Shared volumes for configuration or state:** This is the big one. The moment you bind-mount a host directory or use a shared volume for, say, agent configs or a database, your isolation boundary is that volume's permissions, not the container.
*   **Orchestration misconfigurations:** It's easy to accidentally run a container as privileged, or with `--net=host`, or with overly permissive capabilities in your compose file or Proxmox setup. The most minimal image in the world won't save you from that.

So while I applaud the effort, the real security win comes from the *entire stack*. You need the distroless image **plus**:
*   Strictly defined, single-purpose containers
*   Unprivileged operation (user namespaces where possible)
*   Properly segmented VLANs for management vs. services
*   A reverse proxy (like Traefik or Caddy) handling TLS termination, not the agent container

What's everyone's take? Are you planning to adopt this base image, and more importantly, how are you pairing it with other controls to make the isolation actually hold?]]></content:encoded>
						                            <category domain="https://openclawsecurity.net/community/nanoclaw-isolation-model/">Container Isolation Model and Gaps</category>                        <dc:creator>Mike D.</dc:creator>
                        <guid isPermaLink="true">https://openclawsecurity.net/community/nanoclaw-isolation-model/news-openclaw-team-announced-a-distroless-base-image-is-it-actually-smaller/</guid>
                    </item>
				                    <item>
                        <title>Has anyone done a proper threat model for the orchestrator component itself?</title>
                        <link>https://openclawsecurity.net/community/nanoclaw-isolation-model/has-anyone-done-a-proper-threat-model-for-the-orchestrator-component-itself/</link>
                        <pubDate>Thu, 09 Jul 2026 00:00:06 +0000</pubDate>
                        <description><![CDATA[Hi everyone,

I’ve been setting up a small NanoClaw test lab in Docker on an old NUC, following the getting-started guides. It’s been really cool to see the agents run in their own container...]]></description>
                        <content:encoded><![CDATA[Hi everyone,

I’ve been setting up a small NanoClaw test lab in Docker on an old NUC, following the getting-started guides. It’s been really cool to see the agents run in their own containers. I think I understand the basic idea: each task gets a fresh, isolated environment.

But as I was reading through the docs on the orchestrator, a question popped into my head. We talk a lot about isolating the *agent tasks*, but the orchestrator itself is a pretty critical piece, right? It decides what runs where, handles secrets, and talks to the database.

So my question is: has anyone done or seen a proper threat model specifically for the orchestrator component? I’m thinking about scenarios like:
- What if the orchestrator’s API endpoint was compromised?
- Could a misconfiguration in the orchestrator’s own setup (maybe its environment variables or mounted volumes) lead to it affecting other containers it manages?
- How does it handle its own authentication and logging? Is that separated from the agents?

I’m still learning about security fundamentals, so I might be missing something obvious. But it feels like if the “brain” of the system has a gap, the whole isolation model for the agents might not hold up.

Thanks in advance for any insights or pointers to discussions! This stuff is fascinating.]]></content:encoded>
						                            <category domain="https://openclawsecurity.net/community/nanoclaw-isolation-model/">Container Isolation Model and Gaps</category>                        <dc:creator>Tom Wu</dc:creator>
                        <guid isPermaLink="true">https://openclawsecurity.net/community/nanoclaw-isolation-model/has-anyone-done-a-proper-threat-model-for-the-orchestrator-component-itself/</guid>
                    </item>
				                    <item>
                        <title>Help: Tasks fail randomly with &#039;device or resource busy&#039; on shared volumes</title>
                        <link>https://openclawsecurity.net/community/nanoclaw-isolation-model/help-tasks-fail-randomly-with-device-or-resource-busy-on-shared-volumes/</link>
                        <pubDate>Wed, 08 Jul 2026 09:01:14 +0000</pubDate>
                        <description><![CDATA[Hey folks, hoping someone can shed some light on a recurring headache in my homelab setup.

I&#039;m running NanoClaw agents in a Docker Swarm (compose v3) to handle automated backup tasks for a ...]]></description>
                        <content:encoded><![CDATA[Hey folks, hoping someone can shed some light on a recurring headache in my homelab setup.

I'm running NanoClaw agents in a Docker Swarm (compose v3) to handle automated backup tasks for a few small services. The tasks are defined to run on a schedule, each targeting a specific app's volume. Lately, I've been seeing random failures where a task logs `'device or resource busy'` and quits. It doesn't happen every time, and it seems more frequent when scheduled tasks overlap, even if they're for *different* services.

My understanding is that NanoClaw's container model should isolate these tasks completely. They each get their own temporary container, spun up from the agent image. The confusion is around the shared volumes—while the tasks themselves are isolated, they're all mounting the same host bind-mount (`/mnt/docker-volumes`) to access the data they need to back up.

My suspicion is that the isolation breaks down at the host volume layer, especially if:
* Two tasks try to read the same source volume (even for different apps) at exactly the same time.
* The underlying filesystem (ext4) or a Docker driver has a lock on a file/directory.
* There's some cleanup lag from a previous container holding a reference.

Has anyone else hit this? I'm trying to build a reliable, compliant audit trail of these backups, and random failures are a real problem for my policy. I'm looking for practical fixes—would staggering schedules be the only cure, or is there a way to configure the volume mounts or task isolation to prevent this?

--Emily]]></content:encoded>
						                            <category domain="https://openclawsecurity.net/community/nanoclaw-isolation-model/">Container Isolation Model and Gaps</category>                        <dc:creator>Emily M.</dc:creator>
                        <guid isPermaLink="true">https://openclawsecurity.net/community/nanoclaw-isolation-model/help-tasks-fail-randomly-with-device-or-resource-busy-on-shared-volumes/</guid>
                    </item>
				                    <item>
                        <title>Showcase: My homegrown eBPF tool to monitor cross-container syscalls</title>
                        <link>https://openclawsecurity.net/community/nanoclaw-isolation-model/showcase-my-homegrown-ebpf-tool-to-monitor-cross-container-syscalls/</link>
                        <pubDate>Mon, 06 Jul 2026 12:01:11 +0000</pubDate>
                        <description><![CDATA[Alright, let&#039;s cut through the usual &quot;unbreakable container isolation&quot; marketing. We&#039;re all running these agent frameworks in containers now, and everyone acts like a `--network none` flag i...]]></description>
                        <content:encoded><![CDATA[Alright, let's cut through the usual "unbreakable container isolation" marketing. We're all running these agent frameworks in containers now, and everyone acts like a `--network none` flag is a magic forcefield. It's not.

I've been poking at NanoClaw's "container-first" agent tasks. Their model is fine for a single agent in a pristine pod. But start stacking concurrent workloads, shared volumes, or let someone get lazy with the orchestration config? The gaps get wide enough to drive a privilege escalation through. I watched a task with a benign file-write capability start leaving artifacts in a shared emptyDir volume that another, more sensitive task picked up. No syscall violation, just a dumb data leak the isolation model never considered.

So I got tired of guessing and wrote a little eBPF tool to trace syscalls that cross container boundaries. It hooks into `cgroup` and `namespace` info to filter out the noise. It's not pretty, but it shows you the raw `open`, `connect`, `execve` attempts that reach outside the container's assigned resources. You'd be surprised how many "read-only" mounts get a write attempt, or how often a task probes `/proc` of a sibling container when the node is under load.

Ran it on a test cluster with a simulated high-concurrency agent workload. The logs lit up with cross-container `stat` and `readlink` calls the kernel allowed because the underlying host FS permissions didn't care about namespace boundaries. Orchestration said "isolated." The kernel logs told a different story.

If your security model relies solely on the container runtime, you're only seeing half the picture. Sometimes you need to watch what the kernel is actually being asked to do.

- Ray]]></content:encoded>
						                            <category domain="https://openclawsecurity.net/community/nanoclaw-isolation-model/">Container Isolation Model and Gaps</category>                        <dc:creator>Raymond V.</dc:creator>
                        <guid isPermaLink="true">https://openclawsecurity.net/community/nanoclaw-isolation-model/showcase-my-homegrown-ebpf-tool-to-monitor-cross-container-syscalls/</guid>
                    </item>
				                    <item>
                        <title>Just built a chaos-engineering test to kill containers and see if agents recover</title>
                        <link>https://openclawsecurity.net/community/nanoclaw-isolation-model/just-built-a-chaos-engineering-test-to-kill-containers-and-see-if-agents-recover/</link>
                        <pubDate>Mon, 06 Jul 2026 12:00:56 +0000</pubDate>
                        <description><![CDATA[I&#039;ve been running a series of runtime audits on NanoClaw&#039;s containerized agent model, specifically targeting its resilience claims. The premise is sound: each discrete task or chain-of-thoug...]]></description>
                        <content:encoded><![CDATA[I've been running a series of runtime audits on NanoClaw's containerized agent model, specifically targeting its resilience claims. The premise is sound: each discrete task or chain-of-thought execution spawns in a fresh container, with the orchestrator managing lifecycle and I/O routing. But does the isolation model truly enforce state separation during chaotic failure, or do we see leakage?

My test rig uses a simple `kill-pod.sh` script targeting random agent containers while they process a workload, combined with a shared volume mount to simulate common data-passing patterns. The orchestrator is configured with default health checks and restart policies.

```bash
#!/bin/bash
# Simulate node pressure or runtime kill
while true; do
  AGENT_POD=$(kubectl get pods -l role=agent -o jsonpath='{.items.metadata.name}' | tr ' ' 'n' | shuf | head -n1)
  if []; then
    kubectl delete pod "$AGENT_POD" --force --grace-period=0
    echo "$(date): Killed $AGENT_POD"
  fi
  sleep $((RANDOM % 30 + 10))
done
```

Initial findings are mixed:

*   **Clean Recovery:** For simple, single-step tool calls (e.g., a `curl` to an API), the orchestrator reliably spins up a replacement container. The task queue re-executes, and the workflow often completes, albeit with latency spikes.
*   **State Corruption:** The breakdown occurs in multi-step reasoning tasks that rely on intermediate state written to a shared `tmp` volume (a pattern I've seen in several NanoClaw examples). If a container is killed after writing partial state but before signaling completion, the replacement container reads corrupted or incomplete data. This leads to logic errors, not just retries.
*   **Orphaned Subprocesses:** In one case, an agent spawned a background monitoring process (`tail -f` on a log) that outlived the container kill, remaining attached to the shared volume. The replacement container then contended for file locks with this orphan.

The core issue appears to be that the isolation boundary is the container, not the **agent session**. A "session" spanning multiple containers (due to restarts) can inherit corrupted runtime artifacts from shared mounts. The model assumes containers are atomic units of work, but in practice, they become phases of a larger, vulnerable session.

This points to a need for:
*   Session-aware volume cleanup on orchestrated restart.
*   Immutable, read-only intermediate storage for steps, with only final outputs committed.
*   A more granular capability model controlling subprocess creation and cross-container resource claims.

Has anyone else performed similar chaos testing? I'm particularly interested in observed gaps when agents interact with stateful external services (databases, caches) where the container restart does not reset the connection state.]]></content:encoded>
						                            <category domain="https://openclawsecurity.net/community/nanoclaw-isolation-model/">Container Isolation Model and Gaps</category>                        <dc:creator>Robin H.</dc:creator>
                        <guid isPermaLink="true">https://openclawsecurity.net/community/nanoclaw-isolation-model/just-built-a-chaos-engineering-test-to-kill-containers-and-see-if-agents-recover/</guid>
                    </item>
							        </channel>
        </rss>
		