Meta, October 2021
The Meta outage had no bad decision in it
- When
- 4 October 2021, from about 15:39 UTC. Roughly six hours to full restoration.
- Scale
- Facebook, Instagram and WhatsApp unreachable worldwide. The servers ran the whole time.
- Trigger
- A routine command to assess backbone capacity took down every backbone connection. A bug in the audit tool that should have stopped it did not.
- Why it became global
- Edge DNS servers withdraw their BGP routes when they cannot reach the data centres, by design. With the backbone gone, every one withdrew at once.
- Why it took six hours
- Remote access needed the network. Internal tools needed DNS. The data centres are hard to enter and modify by design. Every recovery path ran through the thing that was broken.
- Primary sources
- Meta Engineering, 4 and 5 October 2021.
On 4 October 2021, Facebook, Instagram and WhatsApp did not go down. They disappeared. For six hours the rest of the internet could not find out where they were, and the servers were running the whole time.
Nobody attacked them. And more uncomfortably, nobody made a mistake that looks like a mistake.
Two pieces of infrastructure
Meta’s backbone is the private network connecting their data centres to each other and out to the smaller facilities at the edge. Every internal system rides on it.
BGP is how networks tell the rest of the internet which addresses they can reach. It is not a lookup service. It is continuous advertisement: send traffic for these addresses to me. If you stop advertising, you stop existing, as far as everyone else is concerned.
Meta’s edge locations advertise the routes to their DNS servers, and DNS is what
turns facebook.com into an address.
Hold that shape. The backbone carries everything internal, and BGP tells the world where to find the front door.
The command, and the guardrail that had a bug
Routine maintenance. In Meta’s words, a command issued “with the intention to assess the availability of global backbone capacity.” A capacity check. Read-only in intent. The kind of command that runs constantly at that scale.
It “unintentionally took down all the connections in our backbone network.”
Meta had anticipated exactly this, and it is the part most retellings skip. They had a system whose entire job was to catch commands like this before they executed: “Our systems are designed to audit commands like these to prevent mistakes like this.”
And then: “a bug in that audit tool prevented it from properly stopping the command.”
The guardrail existed. It was built for precisely this scenario. It had a bug.
So the command ran, and every data centre disconnected from every other data centre, globally, at once. That is already a serious outage. It is not yet a six hour disappearance from the internet.
The safety mechanism that worked perfectly
Meta’s edge locations run DNS servers, and those servers have a safety mechanism.
If a DNS server cannot reach the data centres, it might answer with stale or wrong information. Answering wrongly is worse than not answering. So the design is: if you cannot talk to the data centres, declare yourself unhealthy and withdraw your BGP advertisement. Stop attracting traffic you cannot serve properly.
That is good engineering. If one edge location loses connectivity, it removes itself and traffic goes elsewhere. It is the correct behaviour, and if you were reviewing the design you would approve it.
Except the backbone was gone. So every edge location asked the same question at the same moment, and every one of them got the same answer.
Meta’s words: “the entire backbone was removed from operation, making these locations declare themselves unhealthy and withdraw those BGP advertisements.”
Every DNS server withdrew. Simultaneously. Worldwide.
With that, Meta’s name servers stopped being reachable from the internet. Not down. Unreachable. “Our DNS servers became unreachable even though they were still operational.”
The machines were fine. The data was fine. There was simply no longer any route that led to them.
Then the retries started: billions of devices asking again, and asking harder, piling load onto DNS infrastructure across the whole internet for a name that could no longer be answered. If your own service felt slow that afternoon and you never worked out why, that is why.
Nothing here malfunctioned. The health check did precisely what it was designed to do. It was designed for one edge location failing, and it was handed all of them.
Every way back in was already gone
This is where it becomes genuinely uncomfortable.
They could not reach the data centres remotely: “it was not possible to access our data centers through our normal means because their networks were down.”
They could not use their tooling: “the total loss of DNS broke many of the internal tools we’d normally use to investigate and resolve outages.”
Read that twice. The tools for diagnosing an outage were themselves resolved by DNS. When DNS went, so did the ability to see what was wrong.
So, physical access. And Meta’s data centres are “hard to get into, and once you’re inside, the hardware and routers are designed to be difficult to modify even when you have physical access to them.”
A correction, since this is the most repeated detail of the whole incident. You will read everywhere that engineers could not badge into the building. Meta’s post-mortem does not say that. It says the facilities are hard to enter and the hardware hard to modify, by design. That is enough to make the same point, and it has the advantage of being what they actually wrote.
Every one of those properties is a security control working correctly. Hardened facilities, hardened hardware, no easy console access. On any other day, that is exactly what you want.
Engineers travelled to data centres, debugged on site, and restarted systems by hand. Then brought things back gradually, because flipping everything on at once risked “a new round of crashes due to a surge in traffic.” There was a second reason too: individual data centres were “reporting dips in power usage in the range of tens of megawatts,” and suddenly reversing a dip that size could put “everything from electrical systems to caches at risk.”
Six hours.
What actually failed
Not the command. Not the audit tool bug. That is the trigger, and triggers are interchangeable. Something will always eventually get through.
The failure was circular dependency. Every path to fixing the problem ran through the thing that was broken. Remote access needed the network. The tools needed DNS. The monitoring that would tell you what was wrong was inside the system that was wrong.
A second failure sits alongside it: an automated safety mechanism with no floor. Withdrawing routes when unhealthy is right. Withdrawing every route globally, with nothing that says if every location is failing at the same moment then this is not a local fault, hold the last known good state and escalate to a human, is what turned a bad internal day into a global disappearance.
Three controls, and the order matters.
Out-of-band access. A management path that does not depend on the production network. Separate connectivity, separate credentials, separate DNS or none at all. Expensive, boring, unused for years at a time. Also the only thing that works when the primary path is the thing that failed.
A floor on automated withdrawal. If every health check in the world fails at once, the problem is probably not the world.
Recovery tooling with no dependency on production. Including the runbook. If your incident documentation lives behind a single sign-on that depends on the data centre that is down, you do not have documentation.
Why this one is harder than the others
Meta published a technical post-mortem the next day naming their own guardrail failure. Almost everything above comes from their document, and that is unusual enough to be worth saying.
But the reason this incident matters more than its scale suggests is that there is no villain in it.
The command was routine. The audit tool existed and was the right idea. The DNS health check was correct design. The hardened data centres were correct security. Every individual decision was defensible, and several were best practice.
They composed into six hours of a company not existing.
That is the version of a trust boundary failure that is hardest to find before it happens, because there is no bad decision to go looking for. There is only a set of good ones that share a dependency nobody drew on the same diagram.
So the question for your environment is not what could go wrong. It is this:
What would I need in order to fix it, and does that thing depend on what broke?
Your VPN. Your single sign-on. Your runbook. Your monitoring. Your password manager.
Go and check. Most people find at least one loop.
Sources
- Meta Engineering, More details about the October 4 outage, 5 October 2021. https://engineering.fb.com/2021/10/05/networking-traffic/outage-details/
- Meta Engineering, Update about the October 4th outage, 4 October 2021. https://engineering.fb.com/2021/10/04/networking-traffic/outage/
Outage timings are corroborated by third-party network telemetry. No revenue figure appears here because none of the widely quoted estimates are primary.
Corrections are welcome, and any made are listed, dated, at the end of this article.
Questions this answers
Was Facebook hacked on 4 October 2021?
No. A routine maintenance command intended to assess backbone capacity took down all the connections in Meta's backbone network, and a bug in the audit tool that should have blocked the command did not. No attack was involved, and Meta's servers were operational throughout.
Why did Facebook disappear from the internet rather than just go down?
Meta's edge DNS servers are designed to withdraw their BGP advertisements when they cannot reach the data centres, so they never answer with stale information. With the backbone gone, every edge location withdrew at the same moment, and the rest of the internet no longer had a route to Meta's name servers. Unreachable, not down.
Why did the Facebook outage take six hours to fix?
Every recovery path depended on the system that had failed. Remote access needed the network that was down. The internal tools needed DNS, which was unreachable. The data centres are hard to get into and the hardware hard to modify by design. Engineers had to travel on site, restart systems by hand, and restore gradually to avoid a fresh round of crashes.
Could engineers not badge into the buildings?
That claim is widely repeated but is not in Meta's post-mortem. What Meta wrote is that the facilities are hard to get into and, once inside, the hardware and routers are designed to be difficult to modify even with physical access. That is enough to make the same point, and it is what they actually said.