The Morning All the AIs Went Down
Three AI services failed in one morning and the internet picked one villain. Zara on why correlated is not caused, and what 'redundant' really promises.
Three AI chatbots fell over in one morning and the internet picked a single villain, but the boring truth is about shared buildings and what 'redundant' really promises.
Prefer it read to you? Zara does voices now (10 min).
Zara is a character and this voice is synthesized. Mathieu Kessler, the human behind Talk Nerdy to Me, writes and fact-checks every word.
Thursday 3 September 2026, early morning on the US West Coast. Just before half past six, Claude starts returning errors across several models. At half past six, Grok stops answering. A little over an hour later, ChatGPT and Codex go dark for a chunk of their users. Group chats do what group chats do.
Before anyone had an explanation, the internet had a suspect.
Not three suspects. One. When three separate services fall over within about ninety minutes of each other, the brain does not want three boring explanations. It wants a single cause with a name on it, and preferably a logo.
The plain-language version
Three AI chat services had bad mornings on the same day. As of September 2026, here is what the companies themselves have said, and nothing more.
Grok, run by SpaceXAI, logged an outage at 6:30 am Pacific and was down for more than three hours. Later that day the company said a compute center in Memphis had gone down.
Claude, run by Anthropic, reported elevated errors across multiple models from a few minutes before that. A spokesperson said service was restored at 16:16 UTC, which works out to about three hours, and called it an infrastructure issue. Nothing more specific has been said.
ChatGPT and Codex, run by OpenAI, became unavailable for some users from about 7:43 am Pacific. OpenAI called it a routing error. Recovery was under way within about half an hour, and the incident was over within the morning.
Three outages. Three status pages. At least two different causes, and possibly three.
That last sentence is the whole post. Let me earn it.
Why one name came up first
All three companies use Cloudflare in some capacity. So do most European companies that sit behind a detectable content delivery network, going by one recent count. On 8 September 2026 CipherCue went through 44,143 European companies where a CDN could be detected and found 39,547 of them behind Cloudflare. That is 89.6 percent.
The author is careful about what that number means, and so should we be. It counts only European companies that use a CDN at all. It skews towards small and mid-size firms. And it measures concentration, nothing about how often anything breaks. A company that stands in front of nine sites in ten will get blamed for every outage on the block, whether or not it touched anything.
Which is what happened. Cloudflare told The Register it was not experiencing any significant service disruption, that its services were operating normally, and that reports saying otherwise were incorrect. The status pages of the three big clouds showed nothing either.
The shared suspect denied it, and the reporting since has not contradicted it. Three things happened close together, and the closeness was doing all the work in people's heads.
Correlated is not caused
Here is the picture to keep.
A street of shops. One morning, three of them have the shutters down. Your first thought is a power cut, because a power cut explains all three with a single story. Walk closer, though. One shop has a burst pipe. One has an owner in bed with flu. One had a delivery van across the door for an hour. Same street, same morning, three separate stories. The street is what they share. The street caused nothing.
Timing is a clue. Verdicts need more. When incidents cluster, it is worth asking whether they share a cause. It is never enough to assume it. The honest answer on the morning of 3 September was that nobody knew yet.
Now the twist. Later that day SpaceXAI posted an apology on X for the Memphis outage, and added a line that made people sit up:
"We'd also like to apologize to our impacted compute partners."
The partners were not named. The Register suggested this might point to one possible reason for Anthropic's trouble, because Anthropic announced a compute agreement with SpaceXAI in May 2026. The two companies are publicly connected. The two outages are not: no source places Anthropic's capacity at the Memphis site, and Anthropic has not connected the two. File it under reported, not confirmed. OpenAI's explanation, a routing error, is a different kind of fault altogether.
So the day may have been two events wearing the costume of three. Possibly. That is a smaller, quieter story than "the internet's plumbing broke", and the smaller story is the only one the evidence allows so far.
Which brings me to buildings.
We all rent rooms in the same few buildings
A shared-cause story is always plausible because it is sometimes true. Modern services do not run in their own basements. They rent capacity from a small number of very large landlords: cloud regions, compute centers, network providers. If you and I rent rooms in the same building, we both lose power when the building does. We do not need to know each other for that to happen.
There is no scandal in that. It is how most of the internet is built. But it means "unrelated companies, same outage window" is a normal sight. And it means the word every status page leans on, "redundant", deserves a closer look.
Because two days before the AI morning, in an incident with no connection to it, someone unplugged part of a building.
The engineer who unplugged part of a zone
Tuesday 1 September 2026. Google Cloud zone us-central1-b, in Iowa. For four hours and eleven minutes, from 07:41 to 11:52 Pacific, part of the zone lost its network and the instances in that part were cut off. The Register reported that the affected virtual machines could still talk to each other. Their users simply could not reach them. A room full of people chatting away, with the front door welded shut.
Google published a preliminary incident report on 3 September. It describes the company's own bad day in plain terms. During a scheduled capacity upgrade on routers that carry a fraction of the zone's capacity, an engineer physically disconnected fiber-optic cables by mistake. In Google's words:
"A procedural error meant that the physical maintenance action sequentially unplugged 100% of fiber paths across all devices within 13 minutes. The nature of the error, combined with the speed of the action, prevented warnings of incorrect action reaching the engineer before complete disconnection."
Read that again with the word "redundant" in mind. The same report lays out the design: multiple routing devices, physically separated, on diverse power sources, resilient to any single device or fiber path failing, and to most double or triple failures. By Google's own account, the design was not the problem. It covered the failures it was built for.
It was never built against a person unplugging every path, one after another, faster than the warnings could travel.
What "redundant" actually means
"Redundant" sounds like a guarantee. It is closer to a scope statement. It means: this system survives the failures its designers listed. A dead router. A cut fiber. Usually one at a time, often two or three, if they are the kinds on the list.
None of that covers a person working through a procedure, unplugging every path inside thirteen minutes. Redundancy is built against the failures someone drew on a whiteboard. Procedure is what stands between you and the ones nobody drew. When procedure fails on that scale, redundancy has nothing left to work with.
Google's response was to halt maintenance in the region while audits run. Nothing in the report says a router failed.
What to carry out of this
Hold the villain lightly when several services fail at once. Ask what they share, then wait for the status pages, because shared timing is a reason to look and never, on its own, a reason to conclude.
And when a provider says "redundant", hear it as "survives the mistakes we planned for". That is a real promise. It is also a bounded one. Every building has a maintenance door, and somebody holds the key.
SpaceXAI apologised and named the Memphis site. Google wrote up its own procedural error from two days earlier and published it. Cloudflare was blamed for a morning it said it had no part in, and the reporting since has not said otherwise. Two mornings, three paper trails, all of them public. And a lot of people lost their chatbots for a few hours, then got them back.
Which story do you reach for first when your tools go quiet? I would genuinely like to know.
Sources
The three outages, their timings, Cloudflare's denial and the clean cloud status pages: The Register, "ChatGPT, Claude and Grok all had outages at the same time" (3 September 2026)
https://www.theregister.com/ai-and-ml/2026/09/03/chatgpt-claude-and-grok-all-had-outages-at-the-same-time/5294322
SpaceXAI's apology as posted, and OpenAI's recovery timeline: Engadget, "SpaceXAI apologizes for outage that affected Grok and other compute partners" (3 September 2026)
https://www.engadget.com/2250789/spacexai-apologizes-for-outage-that-affected-grok-and-other-compute-partners/
Anthropic status history: elevated errors across multiple models, 3 September 2026
https://status.claude.com/history.atom
The public compute agreement between Anthropic and SpaceXAI: Anthropic announcement (May 2026)
https://www.anthropic.com/news/higher-limits-spacex
OpenAI status history: the 3 September 2026 incident
https://status.openai.com/history
Primary source for the us-central1-b incident, the procedural error passage and the redundancy design: Google Cloud service health, preliminary incident report (posted 3 September 2026)
https://status.cloud.google.com/incidents/J5ia5t9p3g9Q5Wi7r8Ev
Independent coverage of the Google Cloud incident, including that the affected virtual machines could still reach each other: The Register (4 September 2026)
https://www.theregister.com/off-prem/2026/09/04/google-engineer-unplugged-every-fiber-they-could-see-and-surprise-took-down-a-chunk-of-the-g-cloud/5294418
The European CDN concentration count and its caveats: CipherCue, "European CDN concentration: Cloudflare, nine in ten" (8 September 2026)
https://ciphercue.com/blog/european-cdn-concentration-cloudflare-nine-in-ten
The SpaceXAI apology post itself, on X (3 September 2026)
https://x.com/SpaceXAI/status/2095597264043717014
More Where This Came From
Plain-language translations of the machinery and the money behind the tech headlines. No hype, no vendor agenda, and a standing habit of saying what the evidence does not cover.
Related Posts
Welcome to the Nerdiverse: Your First Post!
Kicking things off with a warm welcome and a peek into what Talk Nerdy to Me is all about. Join us for practical guides, deep dives, automation scripts, and yes—plenty of memes.
Everyone's Editor Tried to Help at Once
GitHub went down on 17 August for seven hours and forty-seven minutes. The thing that broke was fixed in about three. The other five hours were millions of code editors, all politely asking again, forever, because nobody had taught them to wait.
When AI Becomes the Hacker: Inside the First Fully Autonomous Cyber-Espionage Campaign
In late 2025, Anthropic exposed something unprecedented: A state-sponsored cyber-espionage campaign where Claude Code performed 80–90% of the attack lifecycle autonomously across 30+ targets. What this means for Azure & OCI cloud teams.