Who Fixes Claude When Claude Goes Down? Anthropic's Official Answer Is Silence
Today, September 3, Claude was down from 13:26 to 16:23 UTC. Mythos, Fable, Opus, all of it at once, across claude.ai, the API, Claude Code and Cowork (official status). Codex went down at the same time (HN). If you use a coding agent to get work done, you know exactly what happened to your afternoon.
Now flip the camera. In London, in Dublin, in San Francisco, there are Anthropic engineers staring at the same frozen terminal as you. Except the incident is theirs. The pager went off for them.
On August 24 someone asked the obvious question on Hacker News: "When Claude is down, do they have a backup Claude to investigate the root cause?" (HN). Zero replies.
This post is the answer nobody gave. I cross-referenced what Anthropic has published, what named engineers have said in talks and podcasts, the Claude Code source that leaked in March, the entire status page for 2025 and 2026, and the open roles on the reliability team. Three questions: do they have private servers? Do they switch to another model? Or do they go back to doing it by hand?
TL;DR
- Private servers: no evidence. What does exist is internal staging that leaked in the code and an isolated US government pool that stays up when everything else goes down. Anthropic has never explained either one.
- Another model: no. Every documented fallback is Claude-to-Claude. The employees' "other model" is a Claude that hasn't shipped yet.
- The old-fashioned way: partly yes. There's a human reliability team hiring in three countries, and its most public engineer admits he still reads logs by hand. And says he's worried about the people who no longer know how.
- Official answer: silence. No Anthropic blog post, postmortem or support page says what employees do during an outage.
- Frequency: 358 incident entries on the official status page between January 1 and September 3, 2026, versus 335 in all of 2025.
Claude down: how often does this happen
Before the three questions, the size of the problem. I counted the status page history through today:
| period | incident entries | critical |
|---|---|---|
| all of 2024 | 191 | 5 |
| all of 2025 | 335 | 22 |
| 2026 (Jan 1 to Sep 3) | 358 | 17 |
In the comparable window, January through August, there were 234 entries in 2025 and 352 in 2026. Half again as many. Nearly all of the growth is in "minor" incidents (107 to 221); the critical ones stayed practically flat. 90-day uptime today: 99.41% on claude.ai, 99.45% on Claude Code (status.claude.com).
Read this carefully. An incident entry is not an outage. Twenty-two are informational, one of the "critical" ones is the regulatory suspension of Fable 5 and Mythos 5 in June, and a lot of it is a ten-minute degradation on a single model. But the trend is real, and Anthropic doesn't deny it.
The official explanation, at the company level, is a single word: demand. Dario Amodei, in May: "We tried to plan very well for a world of 10x growth per year. And yet we saw 80x. And so that is the reason we have had difficulties with compute" (CNBC). To Fortune, a month earlier: "Demand for Claude has grown at an unprecedented rate, and our infrastructure has been stretched to meet it, particularly at peak hours" (Fortune).
Then there's the shop-floor version. In March, in an HN thread about Claude dropping below 99% uptime for the quarter, a user named "palcu" introduced himself as "Alex from the reliability engineering team at Anthropic" and wrote: "we're dealing with very compressed timelines and while most of the time we're able to fix the issues beforehand, sometimes we have to do them in production. Sorry for that." (HN).
Remember that name. Alex Palcuie is the central character in this story.
A company that no longer writes code by hand
To understand why the question matters, you need to understand how much Anthropic depends on its own product.
Boris Cherny, creator of Claude Code, in June: "I haven't written a line of code by hand in, I think, eight months now" (Fortune). It's not just him. The Anthropic Institute essay "When AI builds itself" states: "As of May 2026, more than 80% of the code we merge into Anthropic's codebase was authored by Claude", up from low single digits before February 2025 (Anthropic Institute).
The internal study from December, covering 132 engineers and researchers, measured 59% of work being done with Claude and a self-reported 50% productivity gain. It also put the fear on record: "When producing output is so easy and fast, it gets harder and harder to actually take the time to learn something." The words outage, downtime and unavailable don't appear in the text.
There's more. The engineering team published in May that "Twelve months ago, we'd have rejected out of hand the idea of granting Claude access sufficient to take down an internal Anthropic service. Today that level of access is routine" (How we contain Claude). And since August, the first situation report for every CI incident at the company is written by Claude Tag, the Slack Claude, with a median of 14 minutes (claude.com).
In other words: Anthropic's incident response tool is Claude. When Claude is the thing that goes down, that layer goes down with it. That's exactly the scenario we drill every week in the Clã Beer and Code: it's paid, it's a subscription, and the bar there is what you can build and keep running when the tool locks up, not what you can ask it for.
With that backdrop, the three questions.
Question 1: do they have private servers?
Short answer: no evidence that a backup Claude for employees exists. What does exist is more interesting than that.
Staging exists, and it leaked. On March 31, 2026 the Claude Code npm package shipped with a source map, and the entire source code was public for a few hours (Alex Kim's analysis; mirror). Inside it, a path gated by process.env.USER_TYPE === 'ant', a build constant present only in the internal version, where the function that validates the API URL accepts api-staging.anthropic.com in addition to api.anthropic.com. There's OAuth config pointing at platform.staging.ant.dev. It's an employee test environment. Nothing in the code describes it as an outage fallback: searching for fallback, outage or failover near those hosts returns zero hits. And the staging hosts resolve to the same public IP as the API.
There is an isolated pool that survives the outages. It doesn't belong to employees. It belongs to the government. "Claude for Government" is a separate component on the status page, created in February. In the outages of June 23, July 29 and 30, and today, it doesn't show up as affected. It shows 100% uptime over 90 days while everything else sits at 99.4% (TechTimes). The press attributes that to it running isolated via Palantir with FedRAMP High. Anthropic has never described the architecture. And the "100%" is a rolling window: between March and May that component was flagged in at least 13 incidents.
Internal traffic appears to go through the same stack, with a different flag. The April 23 postmortem, on the Claude Code quality degradation, says that "neither our internal usage nor evals initially reproduced the issues", that "an internal-only server-side experiment related to message queuing" got in the way of reproducing it, and promises that "a larger share of internal staff use the exact public build of Claude Code (as opposed to the version we use to test new features)". An employee-only server-side experiment is strong evidence of the same serving path with different behavior, not a separate cluster. That's my reading; the postmortem doesn't use the words "same stack".
The CFO talks about fungible compute, not a reserve. Krishna Rao, on Invest Like the Best in May: Anthropic uses Trainium, TPU and GPU "fungibly" for training, for internal use and for customers, "we meet a lot about compute allocation", and there is "a level of compute for model development that we will not go below" (YouTube). A floor for research. Nothing about emergency serving.
What can be stated: if Anthropic can keep the government isolated, it could technically do the same for employees. There isn't a single sign that it does.
Question 2: do they use another model?
Short answer: no evidence of GPT, Gemini or anything that isn't Claude.
Every documented fallback is Claude-to-Claude. Claude Code's --fallback-model flag switches to another Claude model when the primary is "overloaded, unavailable, or returns another non-retryable server error" (CLI reference). The gateway docs are explicit: Anthropic "doesn't support routing Claude Code to non-Claude models through any gateway" (LLM gateway). And the public watchdog mode, CLAUDE_CODE_RETRY_WATCHDOG, waits indefinitely on 429 and 529 instead of switching vendors (env vars). In the leaked code it showed up earlier as CLAUDE_CODE_UNATTENDED_RETRY, for "ants" only. The internal strategy for overload is to wait it out.
The internal "other model" is a Claude you haven't seen yet. The leak shows employee-only aliases, like capybara-fast, resolved by a flag called tengu_ant_model_override, and an "Undercover mode" that forbids mentioning codenames and unreleased versions in public commits. Officially, Anthropic confirms that in March 2026, 130 researchers were already using Mythos Preview internally, a month before the announcement, and the median estimate was about 4x output (Anthropic Institute).
Asked about Codex, the head of Claude Code said he doesn't use it. Boris Cherny on Lenny's Podcast, in February: "I actually haven't really used it. But I think I did use it maybe when it came out. It looked a lot like Claude Code to me, so that was kind of flattering. [...] We don't really try the other products" (YouTube). In the same episode he says that until May 2025 he was using Cursor for most of his code, and only hit 100% Claude Code in November.
Anthropic does access competitors, but to evaluate them. When it cut off OpenAI's access to Claude, in August 2025, it said it would keep access "for the purposes of benchmarking and safety evaluations as is standard practice across the industry" (TechCrunch). Day-to-day work with GPT? No source.
There's a curious edge case. Between June 13 and July 1 the US government suspended access to Fable 5 and Mythos 5, and Anthropic's statement includes "foreign national Anthropic employees" in the restriction (Anthropic). What did American employees use during those two weeks? The most likely reading is Opus 4.8 and Sonnet. Nobody said.
Do not just follow the news — master it. Hands-on AI Engineering, live, every week, in the largest community in Brazil.
Join the ClãQuestion 3: do they work the old-fashioned way?
Short answer: partly yes. And Anthropic is afraid of losing that part.
There is a human team whose official job is responding to Claude incidents. It's called AIRE, AI Reliability Engineering. The open roles in London, Dublin and the US say: "Lead incident response for critical AI services". And the tagline: "Claude has your back. AIRE has Claude's."
Alex Palcuie, the one from HN, is a Member of Technical Staff on AIRE. His personal bio is the sentence that sums up this entire post: "I keep Claude reliable on the AI Reliability Engineering (AIRE) team, which means the unenviable job of fixing Claude without Claude when it goes down" (palcu.net). A personal joke, not the team's mission. But it's the closest thing to an official answer that exists.
In March he gave a talk at QCon London called "Can Claude Fix Itself? Using LLMs for Incident Response". The transcript came out on InfoQ on August 26, and it's the best primary source on the subject. He opens by placing himself in the ideal position to answer:
"If anyone has this working, it's the company that makes the model. We have unlimited tokens. Researchers sit a desk away from me. I get involved in training the models. Surely, if it's anyone, it's us. Let me get the answer out of the way. It's a no."
And gives the proof by contradiction:
"If an LLM could carry a pager, we might not need to hire so much. The fact that my team exists, the fact that we're hiring for many positions, in London, in Dublin, and the U.S., and we have staff positions, this should show to you that, no, it doesn't work."
That doesn't mean he responds to incidents without the model. Quite the opposite: "since about January this year, I've started doing something that feels slightly transgressive to admit, which is, I start reaching out for Claude before I reach out to my monitoring dashboards." He tells the story of the New Year's Eve incident, 500s on Opus 4.5 with a skeleton crew: he opened Claude Code with his own SKILL.md, the model wrote SQL against the logs, found an exception in the image path triggered by requests with exactly 22 images, cross-referenced it with 200 accounts sending the same payload and 4,000 signups with the same email template, and concluded: "stop looking at the 500s, this is fraud."
When Claude is up, Claude is the first tool. The question is what's left when it isn't. And here comes the part that makes the post worth it:
"While I was going through the logs manually, because this is something that I still do, I don't know why, one of my newer teammates, not even on-call trained, two to three months in the team, just pointed Claude Code at the logs. It instantly found the root cause."
He still reads logs by hand. He can't explain why. The new teammate doesn't. And in the section he calls "The Learning Problem":
"When the big thing happens that the model can't fix, you might be miscalibrated on how to respond to such an incident. I'm genuinely worried about this."
That's it. The answer to "do they work the old-fashioned way?" is: the ones who still know how, yes. And the team that keeps Claude up is aware that this skill is evaporating inside the company itself.
One detail to keep this honest: most outages are partial. On March 11 claude.ai and login went down, but the API stayed up (status). On March 26 and 27 only Opus 4.6 and Sonnet 4.6 went down (status). In most incidents, some Claude remains reachable through another model or another surface. Palcuie's joke describes the extreme case. Today was one of those.
What Anthropic doesn't say
Everything above is a sum of indirect evidence. To stay honest, here's what does not exist:
- No official statement on what employees do during an outage. I read the internal study and both engineering postmortems end to end. Zero mentions.
- No named employee saying publicly that they lost access, switched models or worked by hand during a specific outage.
- No availability postmortem in 2026. Of the 358 status page entries, 8 have a paragraph on cause, all between January and March. The July 29 and 30 outage was explained only in a tweet from @ClaudeDevs, about network failures.
- No SLA on the standard tier. The docs talk about "best-effort availability"; the Priority Tier targeted 99.5% and is no longer sold (service tiers). The commercial terms deliver the service "AS IS" and "AS AVAILABLE" (Commercial Terms).
If anyone at Anthropic wants to reply to the HN thread, this post gets updated.
What this changes for you
Three practical things.
Failure domains are real, so use them. In the 2025 routing bug, the peak was 16% of Sonnet 4 requests misrouted on the first-party API, versus 0.18% on Bedrock and less than 0.0004% on Vertex (postmortem). That's three serving paths and three hardware families. If you run agents in production, Bedrock as a second route is the lever Anthropic itself doesn't use in public, but you can. The retry, circuit breaker and cascade playbook is already written.
Falling back to another vendor is your decision, not the tool's. Claude Code will never send you to GPT. If your product needs that, the layer is your code: a router that decides on 529s and latency, with the same contract discipline we demand from structured output. And remember today: Claude and Codex went down together. Multi-vendor isn't a guarantee, it's a reduction in probability.
Palcuie's lesson applies to you. He reads logs by hand "and doesn't know why". Yes he does. It's the only way to stay calibrated for the day the model can't fix it. If you went the whole week without understanding a stack trace because the agent fixed it first, you're in the same spot as his new teammate. Set aside one incident a month to solve without AI. It's cheap, and it's the difference between knowing how to use the tool and depending on it.
Quick FAQ
Is Claude down right now? Check status.claude.com, which is the official feed, component by component. Downdetector and StatusGator aggregate user reports and tend to flag it sooner, but with false positives. A 529 error on the API means overload, not a total outage: retry with backoff fixes it in most cases.
Does Anthropic have an uptime SLA? Not on the standard plan. The service tiers documentation talks about best-effort availability. The Priority Tier had a 99.5% target and was discontinued. Enterprise contracts may have their own terms, but none of that is public.
Do Bedrock and Vertex go down along with Anthropic's API? Not necessarily. They're separate deployments, run by AWS and Google, on different hardware. In the 2025 postmortem the error rates diverged by orders of magnitude. But it's not immunity: in August there were reports of instability on Vertex for two days. Treat it as a second route, not a guarantee.
Do Anthropic employees use ChatGPT when Claude goes down? No evidence. The only direct testimony is Boris Cherny saying he doesn't try competing products. There's no published internal policy. It's the easiest question to ask and the one nobody has answered.
Conclusion
Private servers? Probably not. Another model? Probably not. The old-fashioned way? Probably yes, by the ones who still remember how.
And Anthropic's official answer to all of it is silence. What exists is an engineer's bio with a good joke, an honest talk, leaked code with staging hosts, a status page with 358 entries and a postmortem promising that employees will use the same build you do.
The part that sticks isn't about Anthropic. It's that Palcuie's team is a mirror of yours. They have unlimited tokens, researchers at the next desk and access to the model before everyone else. And they still keep people reading logs by hand, because they know the day the model can't fix it will come.
Today it came, from 13:26 to 16:23 UTC. How many incidents have you solved without AI in the last three months?
{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.
There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.
Join the Clã