problem
KVM agent reports successful unmount of primary storage when umount fails (device is busy), leaving an orphaned NFS mount
versions
ACS 4.23.x
ACS 4.22.x
ACS 4.20.X
2 kvm host
2 pirmary storages ( nfs and cluster scoped)
The steps to reproduce the bug
- Confirm the pool is mounted on the KVM host and tracked in the DB:
KVM host
df -hT | grep <pool-uuid>
virsh pool-list --all | grep <pool-uuid>
management server DB
mysql> SELECT r.host_id, r.pool_id FROM storage_pool_host_ref r JOIN storage_pool p ON p.id = r.pool_id WHERE p.uuid = '<pool-uuid>';
-
A running VM or system VM with a disk on the pool.
-
Make the management server disconnect the host from the pool. Use either option:
Option A: disable the pool and reconnect the host
cmk update configuration name=mount.disabled.storage.pool value=false clusterid=<cluster-id>
cmk update storagepool id=<pool-id> enabled=false
cmk reconnect host id=<host-id>
Option B: storage access group mismatch
cmk configure storageaccess storageid=<pool-id> storageaccessgroups=groupA
cmk configure storageaccess hostid=<host-id> storageaccessgroups=groupB
cmk reconnect host id=<host-id>
-
Check the agent log on the KVM host
grep -E "DeleteStoragePoolCommand|umount|device is busy|Succeeded in unmounting|Failed to unmount"
/var/log/cloudstack/agent/agent.log | grep -E "|busy" | tail -20
You see umount.nfs4: ... device is busy twice, followed by Succeeded in unmounting /mnt/<pool-uuid>.
Logs
2026-09-29 05:23:06,828 DEBUG [cloud.agent.Agent] (AgentRequest-Handler-1:null) (logid:50042c6e) Processing command: com.cloud.agent.api.DeleteStoragePoolCommand
2026-09-29 05:23:06,828 DEBUG [utils.script.Script] (AgentRequest-Handler-1:null) (logid:50042c6e) Executing: /bin/bash -c umount /mnt/0861c431-b02f-3a66-9be5-10587d4f38e5
2026-09-29 05:23:06,844 DEBUG [utils.script.Script] (AgentRequest-Handler-1:null) (logid:50042c6e) umount.nfs4: /mnt/0861c431-b02f-3a66-9be5-10587d4f38e5: device is busy
2026-09-29 05:23:06,844 INFO [kvm.storage.LibvirtStorageAdaptor] (AgentRequest-Handler-1:null) (logid:50042c6e) Attempting to remove storage pool 0861c431-b02f-3a66-9be5-10587d4f38e5 from libvirt
2026-09-29 05:23:06,846 INFO [kvm.storage.LibvirtStorageAdaptor] (AgentRequest-Handler-1:null) (logid:50042c6e) Storage pool 0861c431-b02f-3a66-9be5-10587d4f38e5 has no corresponding secret. Not removing any secret.
2026-09-29 05:23:06,875 ERROR [kvm.storage.LibvirtStorageAdaptor] (AgentRequest-Handler-1:null) (logid:50042c6e) deleteStoragePool removed pool from libvirt, but libvirt had trouble unmounting the pool. Trying umount location /mnt/0861c431-b02f-3a66-9be5-10587d4f38e5again in a few seconds
2026-09-29 05:23:06,875 DEBUG [utils.script.Script] (AgentRequest-Handler-1:null) (logid:50042c6e) Executing: /bin/bash -c sleep 5 && umount /mnt/0861c431-b02f-3a66-9be5-10587d4f38e5
2026-09-29 05:23:11,896 DEBUG [utils.script.Script] (AgentRequest-Handler-1:null) (logid:50042c6e) umount.nfs4: /mnt/0861c431-b02f-3a66-9be5-10587d4f38e5: device is busy
2026-09-29 05:23:11,897 ERROR [kvm.storage.LibvirtStorageAdaptor] (AgentRequest-Handler-1:null) (logid:50042c6e) Succeeded in unmounting /mnt/0861c431-b02f-3a66-9be5-10587d4f38e5
- Check the host state:
df -hT | grep <pool-uuid> # still mounted
virsh pool-list --all | grep <pool-uuid> # gone from libvirt
- Check the DB (same query as step 1). The
storage_pool_host_ref row for this host and pool has been deleted, and no "Unable to detach storage pool" alert was raised (cmk list alerts).
[root@ref-trl-12433-k-Mol8-kiran-chavala-kvm1 ~]# fuser -vm /mnt/0861c431-b02f-3a66-9be5-10587d4f38e5
USER PID ACCESS COMMAND
/mnt/0861c431-b02f-3a66-9be5-10587d4f38e5:
root kernel mount /mnt/0861c431-b02f-3a66-9be5-10587d4f38e5
root 28287 F.... qemu-kvm
[root@ref-trl-12433-k-Mol8-kiran-chavala-kvm1 ~]# for d in $(virsh list --name); do virsh domblklist $d | grep -q 0861c431 && echo "$d uses PS2"; done
v-1-VM uses PS2
The log says "Succeeded in unmounting" after the second unmount also failed. In LibvirtStorageAdaptor.java:821-824, Script.runSimpleBashScript("sleep 5 && umount ...") returns only stdout. umount writes its error to stderr, so the result is null and the code treats that as success and returns true.
As a result, the management server believes the disconnect worked and deletes the storage_pool_host_ref row. Meanwhile the libvirt pool is gone but the NFS mount is still there. That's why df still shows PS2. This is worth filing as its own issue.
What to do about it?
Expected
- The agent checks the
umount exit code, e.g. with Script.runSimpleBashScriptForExitValue(...) == 0.
- On failure,
DeleteStoragePoolCommand returns a failed answer, so the management server keeps the storage_pool_host_ref row and raises the existing "Unable to detach storage pool ... from the host ..." alert.
Actual
- The agent logs "Succeeded in unmounting" and returns success.
- The DB row is removed, the libvirt pool is removed, and the NFS mount stays on the host.
problem
KVM agent reports successful unmount of primary storage when umount fails (device is busy), leaving an orphaned NFS mount
versions
ACS 4.23.x
ACS 4.22.x
ACS 4.20.X
2 kvm host
2 pirmary storages ( nfs and cluster scoped)
The steps to reproduce the bug
A running VM or system VM with a disk on the pool.
Make the management server disconnect the host from the pool. Use either option:
Option A: disable the pool and reconnect the host
Option B: storage access group mismatch
Check the agent log on the KVM host
grep -E "DeleteStoragePoolCommand|umount|device is busy|Succeeded in unmounting|Failed to unmount"
/var/log/cloudstack/agent/agent.log | grep -E "|busy" | tail -20
You see
umount.nfs4: ... device is busytwice, followed bySucceeded in unmounting /mnt/<pool-uuid>.Logs
storage_pool_host_refrow for this host and pool has been deleted, and no "Unable to detach storage pool" alert was raised (cmk list alerts).The log says "Succeeded in unmounting" after the second unmount also failed. In LibvirtStorageAdaptor.java:821-824, Script.runSimpleBashScript("sleep 5 && umount ...") returns only stdout. umount writes its error to stderr, so the result is null and the code treats that as success and returns true.
As a result, the management server believes the disconnect worked and deletes the storage_pool_host_ref row. Meanwhile the libvirt pool is gone but the NFS mount is still there. That's why df still shows PS2. This is worth filing as its own issue.
What to do about it?
Expected
umountexit code, e.g. withScript.runSimpleBashScriptForExitValue(...) == 0.DeleteStoragePoolCommandreturns a failed answer, so the management server keeps thestorage_pool_host_refrow and raises the existing "Unable to detach storage pool ... from the host ..." alert.Actual