Corrective action: make WAL archiving independent of the primary's boot disk (INC-14518)

Corrective action: make WAL archiving independent of the primary’s boot disk (INC-14518)

Macro item 2 from the INC-14518 incident review: WAL archiving resilience (P0)

Problem

On 2026-09-24 the boot disk of the gprd main primary, patroni-main-v17-01, stopped completing I/O from about 22:43 to 23:23 UTC. The data disk was healthy, so Postgres kept committing, but WAL archiving to GCS stopped for the whole 40 minutes.

This mattered because archiving was the only other copy of 01’s WAL:

  • the replicas stopped streaming from 01 at 22:49:56
  • 03 was promoted at 23:20:29 from the last WAL it had (22:49:55)
  • 01 kept committing until 23:23:14

After streaming stopped at 22:49:56, 03’s only other source of WAL was GCS, which had nothing newer than 22:43. So 03 could not get past 22:49:55. Everything 01 committed in the last ~33 minutes stayed on 01 only and did not make it into production.

If archiving had kept running, those writes would have been in GCS within about a minute of being committed (archive_timeout = 1min).

Root cause

1. Everything the archive command needs to start is on the boot disk.

All gprd Patroni clusters (main, ci, sec, registry) archive with the same command, wal-g 3.0.5:

archive_command = /opt/wal-g/bin/archive-walg.sh %p

which runs:

/usr/bin/envdir /etc/wal-g.d/env /opt/wal-g/bin/wal-g wal-push $1 &>>/var/log/wal-g/wal-g.log
Needed by the archive commandPathDisk on 01
Script/opt/wal-g/bin/archive-walg.shboot (/dev/nvme0n1p1)
envdir/usr/bin/envdirboot
wal-g binary/opt/wal-g/bin/wal-gboot
Config + GCS key/etc/wal-g.d/boot
wal-g log/var/log/wal-g/wal-g.loglog disk (/dev/nvme2n1)
WAL files/var/opt/gitlab/postgresql/data17/pg_wal/data disk (/dev/nvme1n1)

Sources: archive-walg.sh, gitlab-walg attributes, gprd-base-db-patroni-main-v17.json, df on 01.

2. One archive call hung and blocked all archiving.

WAL is written in 16 MB files, and each file is archived by one call of the command above. Postgres runs one archive call at a time and only starts the next one when the previous call returns.

  • 22:43:07: last line in the wal-g log before the gap. The file it archived was 00000002001B755500000055, and it reached GCS.
  • The next file, 00000002001B755500000056, was not archived until after the disk recovered. The wal-g log has no lines at all between 22:43:07 and 23:23:19, although the log disk was healthy. wal-g logs Files will be uploaded to storage early in every run, right after loading its config (wal_uploader.go#L87), so no archive call got that far during the stall.
  • 22:47:10: kernel hung-task report for the boot disk journal: task jbd2/nvme0n1p1-:1566 blocked for more than 122 seconds.
  • 23:23:20: Postgres log, the only archiver line that day: FATAL archive command was terminated by signal 3: Quit The failed archive command was: /opt/wal-g/bin/archive-walg.sh pg_wal/00000002001B755500000056
  • 23:23:31: first successful upload of …56 in the wal-g log.

Pending WAL on 01 (pg_archiver_pending_wal_count) went from 5 at 22:43 to 5,931 at 23:24, the last sample before Postgres on 01 was stopped.

3. The alert for this fired but didn’t page.

walgBackupDelayed (no WAL archived for 15 minutes) fired at 23:04 and cleared at 23:22. It is severity: s3 and only posts to Slack (rule).

pg_archiver_pending_wal_count kept reporting from 01 throughout the stall, so an alert on pending WAL growth would also have worked.

Still to confirm

Which exact step the archive call hung in. Candidates, all needing the boot disk:

  • starting the script or envdir
  • loading the wal-g binary
  • reading /etc/wal-g.d
  • the first GCS request (DNS and TLS files under /etc)

The logs can’t narrow this down. No process stack was captured on the day, and the kernel stopped logging hung tasks after reaching its limit of 10. This decides what an option like A (below) would need to move, so it will be confirmed with a reproduction.

Plan

1. Confirm the root cause

  • Map the archive path and the disks it depends on
  • Collect evidence from 01 (wal-g.log, postgresql.csv, kern.log)
  • Save a copy of 01’s Sep 24 wal-g.log.1 before logrotate removes it
  • Reproduce on patroni-amtest-v17 (db-benchmarking) and capture the stack of the stuck archive process
  • Post the confirmed root cause here

2. Evaluate options

Each option is tested with the same boot-disk freeze test, under write load. It passes if WAL keeps reaching GCS (or another copy outside the primary) up to the last commit on the primary, while the primary’s / is frozen.

OptionWhat it changesTo check
A. Archive command independent of the boot diskMove everything archive-walg.sh needs (script, wal-g binary, config, GCS key) off /Does it keep archiving with / frozen? Does anything else it uses (DNS, TLS certs, envdir) still depend on /?
B. archive_mode = always on replicasReplicas also archive the WAL they receiveOn Sep 24 the replicas stopped receiving WAL from 01 at 22:49:56. Would they have archived anything past that point? Duplicate uploads to the same bucket?
C. pg_receivewal --synchronous on a separate nodeA process outside Patroni that streams and stores WAL from the primaryDoes it keep streaming when Patroni/Consul on the primary freeze? How does it follow a failover? Replication slot and max_slot_wal_keep_size impact on the primary?
  • Run the freeze test for each option on patroni-amtest-v17 and record the results here
  • Review the results with DBRE and agree on the option (or combination)

3. Implement the chosen option

  • Open follow-up MRs / issues for the agreed option
  • Validate with the freeze test in gstg
  • Roll out to gprd: main, ci, sec, registry

Test plan

On patroni-amtest-v17, primary node 03. It has the same archive setup as gprd: wal-g 3.0.5, the same archive_command, and separate /var/log and data disks. Run a 60-second dry run first, then a 10-minute run.

  1. Disable chef and pause Patroni.
  2. Start a write load in the postgres database that forces a new WAL file every 5 seconds, so the archiver is always busy.
  3. Remount / with strictatime, so that reading a file also writes to the disk, as on Sep 24.
  4. Start a watcher that records the state, kernel stack and open files of any archive process every 2 seconds, into /var/log.
  5. Freeze /, and thaw it automatically from the same process after the timeout.
  6. Watch new WAL files arrive in GCS, from outside the node.
  7. Afterwards: remount relatime, collect the results, drop the test objects, resume Patroni and enable chef.

The exact commands will be posted here after the dry run.

Differences from Sep 24:

  • fsfreeze blocks writes at the filesystem level, while Sep 24 was stuck disk I/O
  • pd-ssd data disk
  • kernel 6.8.0-1058-gcp instead of 6.8.0-1045-gcp

Acceptance criteria

  • The exact step where archiving blocked on Sep 24 is confirmed and posted
  • Options A, B and C tested with the freeze test, results posted
  • An option agreed with DBRE and documented here
  • The agreed option rolled out to all gprd Patroni clusters, and passing the freeze test