WireGuard gateways couldn't reach recently created or moved Machines via 6PN
WireGuard gateways couldn’t reach recently created or moved Machines via 6PN
September 19 (14:52UTC)
s6pnforwarderd, the daemon that sets up 6PN (our internal IPv6 private networking) forwarding on WireGuard gateways, was crash-looping on some gateway hosts. 6PN on our gateways is configured slightly differently than on workers, and some files expected by the daemon were not found after a recent rollout. We are in the middle of reworking our 6PN setup for better performance and reliability, hence the rollout – gateways were not touched for a long time before this, and some of their peculiarities were forgotten.
This incident was not helped by the fact that we also recently switched our alerting setup, which somehow missed a number of checks on gateways during the conversion. This means that we didn’t notice this for a while (because all existing machines’ 6PN kept working on gateways). Once we noticed though, the fix was quick and we just had to place some files at the expected locations and restart s6pnforwarderd.
We have since carried out an audit for our new alerting setup to catch any missing alerts and reinstate them. For gateways, our plan is to rework their setup to match what’s on workers, so that we do not need to always keep in mind that files are at slightly different locations based on host type. All hosts that run the 6PN stack should share more or less a similar setup. We are also planning to merge some functionalities of gateways with Anycast edges, instead of keeping them on an obscure host type that we seldom work on or notice.