Skip to content

NAS backup: four bugs around guest filesystem quiescing #14295

Description

@MitchDrage

problem

I encountered an issue where a NAS backup of a quiesced database instance failed.
During investigations I found four issues. Issues 1 and 3 apply to any failed backup whatever the cause; issues 2 and 4 are specific to quiescing.
I've listed these four issues together as they are somewhat related. If you'd prefer them broken out into separate issues, just let me know.

The log messages from the failed backup:

ERROR [o.a.c.b.NASBackupProvider] (API-Job-Executor-7:[ctx-5f7ccff5, job-44506, ctx-f6a4a8d4]) (logid:3a91f756) Failed to take backup for VM i-58-3123-VM: Failed to thaw the filesystem for vm i-58-3123-VM: error: guest agent command timed out: guest agent didn't respond to command within '5' seconds
ERROR [o.a.c.b.NASBackupProvider] (API-Job-Executor-7:[ctx-5f7ccff5, job-44506, ctx-f6a4a8d4]) (logid:3a91f756) Backup cleanup failed for VM i-58-3123-VM. Leaving the backup in Error state.

1. A failed scheduled backup is reported as manual

The failed backup mentioned above is shown in the backups list as MANUAL, despite it being a scheduled backup.

Image

From my investigation, I've found that BackupManagerImpl resolves the schedule id before the backup is attempted (line 794) and writes it to the row only after a successful result (line 835). On failure the throw at 830 runs first, so the row keeps backup_schedule_id = NULL. Two effects:

  • listBackups defaults to intervaltype: MANUAL.
  • Retention calls listBySchedule, which filters on backup_schedule_id, so the row is never counted toward maxbackups and never rotated.

Fix: persist the schedule id when the backup row is created, so the linkage survives any outcome.

2. The guest agent timeout is not configurable

The freeze and thaw calls in nasbackup.sh don't pass --timeout to virsh. There doesn't appear to be a way to change it: there are timeouts for Veeam and Networker only, and there is nothing in agent.properties or libvirt's qemu.conf. On a write-heavy instance a thaw exceeds the default (5 seconds) and fails the backup. In my case virsh reached timeout but the thaw completed, so a successful operation was recorded as a failure.

Fix: a zone-scoped config key, backup.quiesce.agent.timeout, giving a global default and a per-zone override (minimum), and ideally a per-VM override.

3. cleanup() tears down the backup destination while the libvirt job is still running

When the thaw failed, the script called cleanup.
That gave me 17 minutes of errors on the hypervisor after the script had already exited:

Timed out during operation: cannot acquire state change lock (held by monitor=remoteDispatchDomainBlockStats)
Path '/tmp/csbackup.zyIrC/i-58-3123-VM/2026.10.01.16.01.20/datadisk.680baabd-859d-4982-9d33-5cdbe1b12e76.qcow2' is not accessible: No such file or directory
End of file while reading data: Input/output error

cleanup() does rm -rf $dest and umount with no virsh domjobabort, while the job started by backup-begin was still writing there.

Fix: abort the libvirt backup job before removing the destination.

4. A failed freeze is never thawed, and its error is discarded

I came across this one unexpectedly - I didn't hit this specifically, so it's simply a code review and a strong suspicion based on how it reads.

Current code:

local thaw=0
if [[ ${QUIESCE} == "true" ]]; then
  if virsh ... '{"execute":"guest-fsfreeze-freeze"}' > /dev/null 2>/dev/null; then
    thaw=1
  fi
fi

thaw is set only when the freeze returns success, and the freeze's stderr goes to /dev/null. A freeze that fails and then completes inside the guest leaves the filesystem frozen with nothing logged.
I've checked the rest of the path and I couldn't find any code later that would unfreeze it.

Fix:

local thaw=0
if [[ ${QUIESCE} == "true" ]]; then
  thaw=1
  if ! freeze_err=$(virsh -c qemu:///system qemu-agent-command "$VM" '{"execute":"guest-fsfreeze-freeze"}' 2>&1 > /dev/null); then
    echo "Failed to freeze the filesystem for vm $VM: $freeze_err"
  fi
fi

versions

CloudStack 4.22.1.1, KVM on AlmaLinux 9, libvirt/virsh 11.10.0, QEMU 10.1.0, Ceph RBD primary storage, NFS backup repository.

The steps to reproduce the bug

As above.

What to do about it?

As above.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions