August 30: Stuck background job queue blocked Sprite deletions

August 30: Stuck background job queue blocked Sprite deletions

Incident start: 22:10UTC

An Oban background job queue used by Sprites got stuck when repeated attempts to place it on a Sprites API Machine failed. This caused Sprite deletions and some other background operations to fail. The queue recovered automatically about an hour later, and Sprite deletion requests began succeeding again.

Investigation traced the placement failures to a Sprites API Machine with a corrupted ext4 volume. The Machine had been migrated between hosts several days earlier, which led us to investigate the migration path.

When we migrate a volume, we use Linux’s dm-clone to allow the destination Machine start before the entire volume has been copied: reads for uncopied blocks come from the source host, while writes and already-copied blocks use the destination disk. After the copy completes, the running Machine continues using the clone device until its next restart.

Although the clone and its underlying destination ultimately access the same storage, Linux treats them as separate block devices and maintains an independent page cache for each. A nightly volume check read filesystem metadata directly from the destination device and cached it there. The running Machine continued updating that metadata through the clone, but those writes did not invalidate the destination device’s cached copy. When the Machine was eventually restarted, it read a mixture of stale metadata and current disk contents, corrupting the filesystem.

We now ensure the kernel has invalidated the underlying device’s cache before allowing a Machine to start. The volume check and other maintenance commands have also been updated to read through the clone device when it exists, preventing the incoherent cache from being created in the first place.