AI & Technology
Grok and Claude Failed at the Same Hour and Lease Compute From the Same Firm
Published: 2026-09-05
In short
Grok and Claude failed minutes apart on September 3. SpaceX blamed its Memphis compute center and apologized to its compute partners.
Mr. Latte's take
You pay two vendors to avoid a single point of failure, but the money underneath both invoices can land with the same landlord who owns the racks. When your primary and your fallback rent from one building, you bought the redundancy line item without buying redundancy. The leasing chain stays out of the contract, which does not make it cheap, it just means the buyer funds it blind.
Grok and Claude Broke Within Minutes of Each Other at 6:30 AM PT
Grok stopped answering first on the morning of September 3. SpaceX opened a models outage on its status page at 6:30 AM PT, and the service stayed down for about three and a half hours. Grok on X and the mobile apps went with it.
Anthropic’s status page lit up at almost the same moment. The incident opened at 13:26 UTC and closed at 16:23 UTC, just under three hours. Claude Mythos 5.1, Fable 5.1, Opus 5, Opus 4.8 and Opus 4.6 were all affected. Anthropic had already logged a separate Sonnet 5 incident earlier that day and cleared it in nineteen minutes.
ChatGPT joined roughly an hour later. OpenAI opened an incident for elevated errors across ChatGPT and Codex at 14:58 UTC and marked it resolved at 16:55 UTC. The company attributed it to a routing error.
The dominoes followed. Cursor posted its own notice saying its service was degraded because of the Grok and Claude outages. Gemini drew a spike in user reports, but Google never confirmed an outage of its own.
Anthropic Signed a Deal This Year to Lease Compute From Musk’s AI Company
SpaceX posted an apology that afternoon. An outage at its Memphis compute center had taken Grok down, the statement said. Then it added one more line: an apology to its impacted compute partners.
It did not name them. Engadget reported that the two companies signed a deal earlier this year for the maker of Claude to lease compute from Musk’s AI firm. Claude’s outage began around the same 6:30 AM PT mark as Grok’s. Anthropic did not disclose a cause, and neither company has said the two failures were one event.
There is no need to write an unconfirmed link as fact. What is confirmed is sharp enough. Two model companies that compete head to head are joined by the same compute, and a supplier publicly acknowledged that when one of its buildings fails, another company’s customers get an apology too.
Cloudflare and the Big Three Clouds Were Fine That Day
When three services wobble together, the shared path is the first suspect. Cloudflare, used in different capacities by all three AI companies, was named first. A spokesperson told The Register that its systems were running fine. AWS, Google Cloud and Microsoft Azure showed nothing relevant on their status pages. Azure drew a spike in outage reports of its own, and Microsoft said that was not the cause.
So the overlap sat below the edge of the internet, down where the GPUs actually run, in a building. The layers a founder counts when designing for redundancy are usually regions, CDNs and model vendors. Who rents compute from whom is not on that list. It does not appear on a status page and it does not appear in your contract.
The severity labels are not reliable either. For the same window on the same day, Anthropic recorded Major impact and OpenAI recorded Minor. The scales are set company by company. Reading your own blast radius off someone else’s dashboard color will mislead you.
The Incident That Hit Asian Working Hours Came the Next Day
For teams in the Americas, September 3 landed in the middle of the workday. For everyone in Asia it ran from 22:26 local time to 01:55 the next morning in Seoul and Tokyo. It passed while they slept, which makes it easy to read as someone else’s outage.
The one that actually took Asian working hours down came the following day. OpenAI’s status page carries a separate APAC incident on September 4, from 07:00 to 10:46 UTC, covering elevated errors in ChatGPT, Work, image generation, file upload, Voice and Codex Cloud. In Seoul and Tokyo that is 16:00 to 19:46, three hours and forty six minutes straight through the afternoon.
Almost no English coverage picked it up. It was labeled Minor and it was regional. If you sell into Asia, the order is reversed: the famous outage happened overnight and the one that stopped work went unreported.
The Question Now Is Where Your Vendor Rents Its Compute
Two model vendors is close to standard practice now. Route to A, fail over to B. September 3 showed how far that arrangement carries. When both vendors sit under the same lease, the switch has nowhere to send traffic.
So the thing to check is not the vendor name but the layer underneath it. Does the model company you use own its GPUs or rent them, and if it rents, from whom. Often you cannot tell unless a deal was announced, which is a reason to put the question on the list when you negotiate an enterprise agreement. When you pick a fallback, put one option in the mix whose rental chain is different.
Then decide what your own screen shows during a three hour outage. Cursor had to publish a notice on September 3 because two models died together. Your users do not know which model stopped. The name on the screen is yours.
Sources
- ChatGPT, Claude, and Grok all had outages at the same time The Register
- SpaceXAI apologizes for outage that affected Grok and other 'compute partners' Engadget
- Elevated errors across ChatGPT and Codex OpenAI Status
- Elevated errors for multiple models Claude Status
- It's not just you; ChatGPT, Claude, and Grok were all down in confirmed outages 9to5Google
- 챗GPT·클로드·그록이 같은 날 잇달아 멈췄다 Platum