Corrective action: make WAL archiving independent of the primary’s boot disk (INC-14518)
Macro item 2 from the INC-14518 incident review: WAL archiving resilience (P0)
Problem
On 2026-09-24 the boot disk of the gprd main primary, patroni-main-v17-01, stopped completing I/O from about 22:43 to 23:23 UTC. The data disk was healthy, so Postgres kept committing, but WAL archiving to GCS stopped for the whole 40 minutes.
This mattered because archiving was the only other copy of 01’s WAL:
- the replicas stopped streaming from 01 at 22:49:56
- 03 was promoted at 23:20:29 from the last WAL it had (22:49:55)
- 01 kept committing until 23:23:14
After streaming stopped at 22:49:56, 03’s only other source of WAL was GCS, which had nothing newer than 22:43. So 03 could not get past 22:49:55. Everything 01 committed in the last ~33 minutes stayed on 01 only and did not make it into production.
If archiving had kept running, those writes would have been in GCS within about a minute of being committed (archive_timeout = 1min).
Root cause
1. Everything the archive command needs to start is on the boot disk.
All gprd Patroni clusters (main, ci, sec, registry) archive with the same command, wal-g 3.0.5:
archive_command = /opt/wal-g/bin/archive-walg.sh %p
which runs:
/usr/bin/envdir /etc/wal-g.d/env /opt/wal-g/bin/wal-g wal-push $1 &>>/var/log/wal-g/wal-g.log
| Needed by the archive command | Path | Disk on 01 |
|---|---|---|
| Script | /opt/wal-g/bin/archive-walg.sh | boot (/dev/nvme0n1p1) |
envdir | /usr/bin/envdir | boot |
| wal-g binary | /opt/wal-g/bin/wal-g | boot |
| Config + GCS key | /etc/wal-g.d/ | boot |
| wal-g log | /var/log/wal-g/wal-g.log | log disk (/dev/nvme2n1) |
| WAL files | /var/opt/gitlab/postgresql/data17/pg_wal/ | data disk (/dev/nvme1n1) |
Sources: archive-walg.sh, gitlab-walg attributes, gprd-base-db-patroni-main-v17.json, df on 01.
2. One archive call hung and blocked all archiving.
WAL is written in 16 MB files, and each file is archived by one call of the command above. Postgres runs one archive call at a time and only starts the next one when the previous call returns.
- 22:43:07: last line in the wal-g log before the gap. The file it archived was
00000002001B755500000055, and it reached GCS. - The next file,
00000002001B755500000056, was not archived until after the disk recovered. The wal-g log has no lines at all between 22:43:07 and 23:23:19, although the log disk was healthy. wal-g logsFiles will be uploaded to storageearly in every run, right after loading its config (wal_uploader.go#L87), so no archive call got that far during the stall. - 22:47:10: kernel hung-task report for the boot disk journal:
task jbd2/nvme0n1p1-:1566 blocked for more than 122 seconds. - 23:23:20: Postgres log, the only archiver line that day:
FATAL archive command was terminated by signal 3: QuitThe failed archive command was: /opt/wal-g/bin/archive-walg.sh pg_wal/00000002001B755500000056 - 23:23:31: first successful upload of
…56in the wal-g log.
Pending WAL on 01 (pg_archiver_pending_wal_count) went from 5 at 22:43 to 5,931 at 23:24, the last sample before Postgres on 01 was stopped.
3. The alert for this fired but didn’t page.
walgBackupDelayed (no WAL archived for 15 minutes) fired at 23:04 and cleared at 23:22. It is severity: s3 and only posts to Slack (rule).
pg_archiver_pending_wal_count kept reporting from 01 throughout the stall, so an alert on pending WAL growth would also have worked.
Still to confirm
Which exact step the archive call hung in. Candidates, all needing the boot disk:
- starting the script or
envdir - loading the wal-g binary
- reading
/etc/wal-g.d - the first GCS request (DNS and TLS files under
/etc)
The logs can’t narrow this down. No process stack was captured on the day, and the kernel stopped logging hung tasks after reaching its limit of 10. This decides what an option like A (below) would need to move, so it will be confirmed with a reproduction.
Plan
1. Confirm the root cause
- Map the archive path and the disks it depends on
- Collect evidence from 01 (
wal-g.log,postgresql.csv,kern.log) - Save a copy of 01’s Sep 24
wal-g.log.1before logrotate removes it - Reproduce on
patroni-amtest-v17(db-benchmarking) and capture the stack of the stuck archive process - Post the confirmed root cause here
2. Evaluate options
Each option is tested with the same boot-disk freeze test, under write load. It passes if WAL keeps reaching GCS (or another copy outside the primary) up to the last commit on the primary, while the primary’s / is frozen.
| Option | What it changes | To check |
|---|---|---|
| A. Archive command independent of the boot disk | Move everything archive-walg.sh needs (script, wal-g binary, config, GCS key) off / | Does it keep archiving with / frozen? Does anything else it uses (DNS, TLS certs, envdir) still depend on /? |
B. archive_mode = always on replicas | Replicas also archive the WAL they receive | On Sep 24 the replicas stopped receiving WAL from 01 at 22:49:56. Would they have archived anything past that point? Duplicate uploads to the same bucket? |
C. pg_receivewal --synchronous on a separate node | A process outside Patroni that streams and stores WAL from the primary | Does it keep streaming when Patroni/Consul on the primary freeze? How does it follow a failover? Replication slot and max_slot_wal_keep_size impact on the primary? |
- Run the freeze test for each option on
patroni-amtest-v17and record the results here - Review the results with DBRE and agree on the option (or combination)
3. Implement the chosen option
- Open follow-up MRs / issues for the agreed option
- Validate with the freeze test in gstg
- Roll out to gprd: main, ci, sec, registry
Test plan
On patroni-amtest-v17, primary node 03. It has the same archive setup as gprd: wal-g 3.0.5, the same archive_command, and separate /var/log and data disks. Run a 60-second dry run first, then a 10-minute run.
- Disable chef and pause Patroni.
- Start a write load in the
postgresdatabase that forces a new WAL file every 5 seconds, so the archiver is always busy. - Remount
/withstrictatime, so that reading a file also writes to the disk, as on Sep 24. - Start a watcher that records the state, kernel stack and open files of any archive process every 2 seconds, into
/var/log. - Freeze
/, and thaw it automatically from the same process after the timeout. - Watch new WAL files arrive in GCS, from outside the node.
- Afterwards: remount
relatime, collect the results, drop the test objects, resume Patroni and enable chef.
The exact commands will be posted here after the dry run.
Differences from Sep 24:
fsfreezeblocks writes at the filesystem level, while Sep 24 was stuck disk I/Opd-ssddata disk- kernel
6.8.0-1058-gcpinstead of6.8.0-1045-gcp
Acceptance criteria
- The exact step where archiving blocked on Sep 24 is confirmed and posted
- Options A, B and C tested with the freeze test, results posted
- An option agreed with DBRE and documented here
- The agreed option rolled out to all gprd Patroni clusters, and passing the freeze test
Related
- Incident: https://gitlab.com/gitlab-com/gl-infra/production/-/work_items/23018
- Incident review: https://gitlab.com/gitlab-com/gl-infra/production/-/work_items/23026
- Macro item 1, watchdog: https://gitlab.com/gitlab-com/gl-infra/production/-/work_items/23046
- Macro item 4, alerts: https://gitlab.com/gitlab-com/gl-infra/production/-/work_items/23045.
walgBackupDelayedfired at 23:04 but is s3;pg_archiver_pending_wal_countkept reporting from 01 during the stall. - Half-dead primary runbook: https://gitlab.com/gitlab-com/gl-infra/production-engineering/-/work_items/29936