I feel obligated to acknowledge the Facebook outage yesterday. Big outage. I’m sure you noticed – it seemed most of the planet did. It was a BGP and DNS issue.
It’s always DNS.
I don’t normally care about outages, as I outlined last week. Two reasons to highlight this.
First, for international users, the impact to WhatsApp was material, as that’s the basic communications tool for a good portion of the world. That makes this a significant communications outage.
And that reason leads to the second. I’m linking this to the story about the VoIP outages from last week. Those VoIP outages are longer, and as I noted last time, it’s the length there that matters. What I missed, which was pointed out to me by a vendor in this space, was that with the move to VoIP, these outages are also impact to critical infrastructure. Phone service, 911 service and the like… and importantly, that most customers (and their providers) may have extensive disaster recovery plans for data and then not for phone service. What’s the failover plan?
I noted another story that New Zealand is decommissioning its plain old telephone system next year. Failing over to POTS lines may soon be a thing of the past, and what of the failure of WhatsApp as its replacement, particularly in non US countries.
So there’s a space here for continuity planning beyond data into this critical service.

