Companies

ChatGPT, Claude and Grok Down the Same Morning, No Shared Cause

On September 3 Claude was down for three hours, Grok for a Memphis compute outage, ChatGPT for 34 minutes. Different explanations, clean cloud providers, and a lesson about fallbacks that share infrastructure.

Muhammet Fatih BatmanSeptember 5, 20263 min read2 views
ChatGPT, Claude and Grok Down the Same Morning, No Shared Cause

Picture the operations lead at a mid-sized logistics company on Thursday, September 3. Her dispatch assistant runs on Claude, with ChatGPT wired in as the fallback. At 13:26 UTC the primary starts throwing errors. Seventy-seven minutes later the fallback does too. Grok, which she does not use, is down in the same window. The scene is invented; the timestamps are not. Three companies gave three different explanations, and none of them pointed to a common cause.

The timeline

Anthropic's status page logged "elevated errors for multiple models" from 13:26 UTC to 16:16 UTC: three hours and six minutes. Affected surfaces were claude.ai, the Claude API, Claude Code and Claude Cowork; affected models were Mythos 5.1, Fable 5.1, Opus 5, Opus 4.8 and Opus 4.6. The root cause was marked "identified" but never described publicly beyond "an infrastructure issue."

xAI opened an investigation into Grok around 13:30 UTC. Its later statement was more specific: "We are sorry for the issues you may have experienced with Grok following an outage at our Memphis compute center this morning," followed by an apology to "our impacted compute partners." That phrase drew attention because Anthropic agreed in May to lease the entire capacity of the Colossus 1 facility in Memphis. Neither company confirmed whether the two incidents shared the site.

OpenAI's was the shortest: a "routing error" starting at 14:43 UTC made ChatGPT and Codex unavailable for some users, resolved by 15:17. Thirty-four minutes. An OpenAI infrastructure engineer wrote on Hacker News that it was an error inside their own stack.

What it was not

Cloudflare denied involvement outright: "Cloudflare is not experiencing any significant service disruptions at this time." AWS, Google Cloud and Azure status pages showed nothing relevant. Wired ran the story on September 4 under the headline that nobody is saying why, and the shared-cause theories circulating on Hacker News remain theories.

What an IT lead actually saw

From the operations chair, three things stand out. First, "we have two providers, we are covered" held for 77 minutes. Second, outages do not have to be simultaneous to overlap, and inside that overlap agents can quietly drop queued work. Third, no provider tells you the root cause in time to act on it; Anthropic's "root cause identified" line is still empty.

Forrester analyst Charlie Dai put it to ITPro this way: enterprises should treat AI availability as a resilience issue and build "multi-model strategies, fallback workflows, and business continuity plans instead of assuming frontier AI services will always be available."

A short checklist

  • Ask whether your fallback provider shares physical infrastructure with your primary. Memphis shows the question is not theoretical.
  • Write the "model does not respond" branch into every agent workflow: hold in queue, route to a human, but never drop silently.
  • Subscribe to status pages, but trigger alerts from your own error rate, not the vendor's announcement. OpenAI's 34 minutes ended before some users saw it acknowledged.
  • If your contract has no SLA, price the outage yourself and compare it with the cost of a second integration. For most teams a second API key is cheaper than the first incident.

Sources: The Register, Anthropic status page, ITPro: Forrester comment, Hacker News thread

Share This Article

Muhammet Fatih Batman

Written by

Muhammet Fatih Batman

Founder & Editor

Founder of YZ Uzman, with 20+ years of experience in web design and software development.

More news

Want to put this technology to work in your business?

Let's talk