September 15: Stuck DNS broke some Depot builds
Incident start: 03:59UTC
A number of build attempts failed over a roughly 24-hour period, with affected builds hanging for about five minutes before timing out. The failures were concentrated on two ingress machines in SYD and JNB that had stopped properly routing connections to builders.
The underlying cause was a bug in corro-dns, the internal DNS resolver we use for service discovery. corro-dns deduplicates concurrent lookups for the same record: if a refresh is already in flight, new callers subscribe to that same refresh rather than starting another. If the caller’s patience runs out (about one second), it falls back to the cached answer. The problem is that when the background refresh itself hangs, it never completes, and every subsequent lookup for that record subscribes to the same stuck refresh, waits a second, and gets the same stale cached answer, forever. For most records the stale data was still close enough to correct that nothing visibly broke. But _instances.internal, a heavy query used for builder discovery, went stale in a way that left the affected ingress machines unable to find running builders. The machines kept accepting connections but couldn’t route them, so clients saw TLS handshakes fail or time out.
Once we became aware of the issue we cordoned the two affected ingress machines, which immediately restored builds for affected users. We then restarted corro-dns on those hosts, confirmed the DNS results were fresh again, uncordoned the machines, and verified traffic was flowing.
The short-term answer is that we’ll add a timeout to the refresh path in corro-dns so a single hung query can’t permanently poison the cache for that record. Also, equally importantly, queries shouldn’t hang. We’re also planning a bigger overhaul of our internal DNS setup – stay tuned!