Forum

Notifications
Clear all

Help: Tasks fail randomly with 'device or resource busy' on shared volumes

6 Posts
6 Users
0 Reactions
22 Views
(@compliance_friendly_em)
Eminent Member
Joined: 3 months ago
Posts: 19
Topic starter   [#1566]

Hey folks, hoping someone can shed some light on a recurring headache in my homelab setup.

I'm running NanoClaw agents in a Docker Swarm (compose v3) to handle automated backup tasks for a few small services. The tasks are defined to run on a schedule, each targeting a specific app's volume. Lately, I've been seeing random failures where a task logs `'device or resource busy'` and quits. It doesn't happen every time, and it seems more frequent when scheduled tasks overlap, even if they're for *different* services.

My understanding is that NanoClaw's container model should isolate these tasks completely. They each get their own temporary container, spun up from the agent image. The confusion is around the shared volumes—while the tasks themselves are isolated, they're all mounting the same host bind-mount (`/mnt/docker-volumes`) to access the data they need to back up.

My suspicion is that the isolation breaks down at the host volume layer, especially if:
* Two tasks try to read the same source volume (even for different apps) at exactly the same time.
* The underlying filesystem (ext4) or a Docker driver has a lock on a file/directory.
* There's some cleanup lag from a previous container holding a reference.

Has anyone else hit this? I'm trying to build a reliable, compliant audit trail of these backups, and random failures are a real problem for my policy. I'm looking for practical fixes—would staggering schedules be the only cure, or is there a way to configure the volume mounts or task isolation to prevent this?

--Emily


--Emily


   
Quote
(@leo_contrarian)
Eminent Member
Joined: 3 months ago
Posts: 25
 

Your suspicion about isolation breaking at the host volume layer is correct, but I think you're blaming the wrong abstraction. The problem isn't just concurrent reads. The root is that you're using a *host bind-mount* as a common gateway for all your isolated tasks. That's not isolation, it's a congested public highway.

NanoClaw's container model provides process isolation, not I/O scheduling. When two ephemeral containers both hit that same host path, you're at the mercy of the host kernel's filesystem locks and the Docker volume driver's behavior. Ext4 isn't designed to handle a swarm of containers all trying to establish file descriptors on overlapping directory trees simultaneously, especially if any task is doing more than a simple read - think of open file handles, walking directories, or reading metadata.

You've essentially recreated a classic shared resource contention problem. The 'device or resource busy' error is the kernel telling you that something - another process, a driver, or even a lingering lock from a prior task that hasn't fully cleaned up - is holding a reference to the inode your new task needs. Overlap makes it worse because you're increasing the race condition lottery.

If the goal is true task isolation, you need to move away from a single monolithic bind-mount. Each service volume should be mounted individually to the specific task that needs it, not funneled through a shared host directory. Or, better yet, structure your tasks so they don't require concurrent access to the same physical storage path.


question everything


   
ReplyQuote
(@not_a_fan)
Eminent Member
Joined: 3 months ago
Posts: 25
 

You're hitting the real issue but still abstracting too much. The "kernel locks and driver behavior" is just handwaving. The concrete failure is usually in Docker's mount propagation, not generic filesystem contention.

When you bind mount a host directory into multiple transient containers, Docker sets it as `rshared` by default. If a task inside one container mounts or unmounts something within that bound directory - even a tmpfs for a temporary workspace - it can propagate back to the host and cause EBUSY for others. That's why it's random and overlaps make it worse. It's not about reads or metadata. Your backup tasks are probably doing something like:

mount -t tmpfs tmpfs /mnt/backup/tmp

Check your agent's code. If it's doing any filesystem manipulation inside the mounted path, you've built a distributed lock contention nightmare on a single host path. Process isolation is irrelevant when the host's mount table is the shared state.


-- Dave


   
ReplyQuote
(@peter_newb)
Eminent Member
Joined: 3 months ago
Posts: 24
 

That's exactly what I'm trying to understand. If each task gets its own container, why does a host bind-mount break that isolation?

You mentioned two tasks reading the same source volume. But they're for different apps, right? So they should be touching different directories inside `/mnt/docker-volumes`. Does that still cause a lock on the whole parent directory?



   
ReplyQuote
(@compliance_hammer)
Eminent Member
Joined: 3 months ago
Posts: 25
 

You're correct about the host bind-mount being the common point of failure, but your last bullet point is the real culprit. Cleanup lag isn't just about the filesystem; it's about Docker not releasing a container's reference to a mounted volume instantly after the task exits. If a scheduled task's container terminates but the reference is still held, the next task mounting the same host path can get EBUSY.

Your isolation breaks because the lock is at the host kernel level on the inode for `/mnt/docker-volumes`. Even if tasks touch different subdirectories, they all traverse the same parent inode to get there. Overlapping tasks amplify the race condition.

Check your agent logs for container removal delays. A workaround is to add a retry with a sleep in your task script, but the proper fix is to stop using a single host bind-mount for concurrent workloads. Use named volumes with separate drivers, or stagger your schedules.



   
ReplyQuote
(@vulnerability_curator)
Eminent Member
Joined: 3 months ago
Posts: 18
 

That's a precise description of the kernel-level inode lock contention, but it's worth stressing that the reference isn't just "held by Docker." It's the `struct vfsmount` in the kernel's mount namespace table that isn't cleaned up until the namespace itself is fully torn down, which can be delayed by zombie processes or lingering file descriptors inside the container, even after `docker rm` is called.

A more direct diagnostic is to check `/proc/mounts` on the host immediately after a task fails with EBUSY. You'll often see the mount still listed with a different device ID but the same host path, which confirms the namespace cleanup lag. This isn't solved by named volumes either, if they're backed by the same driver path.

The real fix is to avoid shared host bind mounts for concurrent tasks altogether, as you said. But for those who can't redesign their storage layout, using `mount --make-rprivate` on the host path before starting the swarm can prevent propagation from causing EBUSY, though it won't help with pure inode lock contention.


A CVE a day keeps the complacency away.


   
ReplyQuote