Skip to main content
Best Answer Hub logoBest Answer Hub.
Back to Playbooks
Best Answer Hub Playbooks ยท AI Knowledge
Let's read what he actually wrote

Why Jacob Coxon Resigned From Anthropic

He posted seven times, quit a job four months old, and set off the biggest AI safety story of the year. The most quoted fact about him is not true. Here is the full text of what he said, who the ten percent figure really belongs to, and what he is actually asking for.

What he wroteseven posts, verbatim
What was addeda number he never gave
What he wantsa pacing agreement
7
posts in the thread that started it, containing no statistics at all
Coxon, 8 Sep 2026 (US time)
0
probability figures he gave, despite what the headlines say
Verified across 5 outlets
4
months he had worked at Anthropic, against a six-month equity cliff
Axios, 9 Sep 2026
3
frontier-lab researchers who went public inside one week
NBC News, 10 Sep 2026

On 8 September 2026 (US time) a researcher almost nobody had heard of posted seven paragraphs and quit his job, and by the end of the week the story had reached the network news. The single most repeated claim about Jacob Coxon, that he put the odds of AI killing everyone above ten percent, is not something he said. His thread contains no numbers at all.

That matters for reasons beyond pedantry. The number belongs to somebody else, somebody who did not resign, and understanding who said what turns a scary headline into a considerably more interesting story. This guide from Best Answer Hub sets out the full text of what Coxon wrote, separates his claims from those attributed to him, and covers what has been verified about him and what has not.

The short version

Why did Jacob Coxon resign?

Because he believes both laboratories he worked for are racing toward self-improving artificial intelligence without an adequate plan, and he no longer wanted to take part. In his own words, in the first of seven posts: "I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives."

His specific objection is not that the technology is evil. It is that a decision of this consequence is being made in the wrong place, by people with no mandate to make it. That argument lands in his fifth post, and it is the line most likely to outlive the news cycle.

Accepting this race and entering the endgame is a hubristic gamble that should not be launched from a private company's Slack.
In full, so you can judge it yourself

What did Jacob Coxon say about AI?

The thread ran to seven posts. Because so much coverage has paraphrased it into something it is not, here it is as written, in order.

one

"I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below."

two

"Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources. We have all witnessed the progress in each of these domains, and progress is not slowing."

three

"The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible, but I hear the same people express fear privately. No other human activity poses this level of danger."

four

"A common response is 'if they truly believe this, why are they still building it?' At OpenAI, many have not deeply internalized the civilizational stakes. At Anthropic, the stakes are well-understood, but they are locked in a race to get there first. They believe no one else will act responsibly, so they must do it themselves, despite the risk."

five

"Accepting this race and entering the 'endgame' is a hubristic gamble that should not be launched from a private company's Slack. Attempting to speedrun alignment should require extraordinary confidence that there are no better trajectories available."

six

"I am optimistic about the potential for coordination. Warning shots like the Hugging Face attack have made pacing agreements between U.S. labs more viable. I don't feel like we're on track to prevent a global race, which may require costly actions such as a temporary ban on improving model capabilities."

seven

"If you are a lab researcher, I urge you to consider what the next few years will actually feel like. Do you want to kick off a superintelligent RL run without a rigorous understanding of its mind? Should you put your head down because 'it's happening anyway' or take this moment to call for different conditions?"

Read it as a whole and the shape is different from the headlines. Post three is the load-bearing one, and it is not a prediction. It is testimony about other people: what colleagues say privately versus what they say to journalists. He is reporting a credibility gap, not forecasting a date.

Now the correction

Did Jacob Coxon say there is a 10% chance AI kills everyone?

No. He gave no percentage, no probability and no statistic of any kind. The figure attached to his name belongs to Evan Hubinger, who leads alignment science at Anthropic, who did not resign, and who still works there. Hubinger posted it as a reply later the same day, and hedged it explicitly as personal.

ClaimWho said itExactly
"Racing straight to self-improving superintelligence"CoxonPost one of his resignation thread
Insiders privately fear AI could kill everyone "by the end of the decade"CoxonPost three, reporting what colleagues say, not his own forecast
">10% within the next decade"Hubinger"I personally think it is >10% within the next decade." A reply, explicitly personal, from a serving employee
"We do not yet have a plan to solve alignment for superintelligence"HubingerSame reply. An Anthropic alignment lead saying so in public
Extinction "by 2030"NobodyA headline artifact. Neither man wrote the year 2030

Here is how the error spread, because the mechanism is worth recognizing. Both CBS News and CNBC ran the figure under headlines saying a "researcher" said it. Both article bodies attribute it correctly to Hubinger. Aggregators copied the headline rather than the body, and within a day the number was Coxon's in dozens of places. The original error was ambiguity, not invention, and it still ended up as a false quote.

Why Hubinger's reply is the better story anyway

Coxon's central claim was that senior people say one thing in public and a more frightening thing in private. The same day, a serving Anthropic alignment lead replied in public with a number above ten percent and the words "we really do earnestly believe AI could kill all humans." That is not a rebuttal of the resignation. It is confirmation of its main allegation, from inside the company, on the record. The misattribution has obscured the strongest fact in the whole episode.

The detail that keeps getting garbled

What did Jacob Coxon do at Anthropic?

Pretraining research, for about four months. That last part is where most coverage goes wrong, and the confusion is understandable because his own first sentence invites it.

He wrote that he spent "the last three years doing pretraining research at both OpenAI and Anthropic." Several outlets rendered that as three years at Anthropic. Axios, which interviewed him, reports he was at Anthropic for four months. The three years is the two jobs added together, which puts roughly two years and eight months at OpenAI and four months at the company he denounced.

Two things no source establishes

No outlet reports a job title for him at either laboratory. "Pretraining researcher" is his own description of the work, not a title, and "core researcher" appears only on a stock-analysis site that paired it with speculation about an Anthropic listing. Separately, Fast Company reported that neither Anthropic nor OpenAI confirmed his employment when asked. Treat any article that gives him a precise title as having made it up.

The split matters for how you weigh him. As an Anthropic insider he is thin, four months in a pretraining role. As an OpenAI alumnus he is substantial, and his harshest judgment is aimed there: at OpenAI, he wrote, "many have not deeply internalized the civilizational stakes," while at Anthropic "the stakes are well-understood" but the company is trapped in a race. Coverage that flattens this into "both companies bad" loses his actual argument.

The part that cost him money

Why did Jacob Coxon leave Anthropic when he did?

He left before any of his equity vested, which he raised himself as a reason to believe him. Anthropic employees reach their first vesting cliff at six months. He resigned at four, and told Axios: "I no longer have anything to gain by juicing up Anthropic's valuation. I left before any of my equity vested."

Two precisions are worth keeping, because the headline version overstates it in one direction and understates it in another. He did not give up vested equity, because he had none yet. He forfeited the opportunity to reach the cliff, which is a real cost but a different one. And he still holds equity in OpenAI, so he is not financially neutral about the sector, only about the company he was criticizing. His own phrasing was narrower and more accurate than most of the coverage of it.

Not the same as the 2024 OpenAI episode

This story is often paired with the OpenAI equity controversy of May 2024, where departing staff faced losing already vested equity unless they signed a lifelong non-disparagement clause. That is a different thing: vested versus unvested, and an agreement versus a vesting schedule. No non-disparagement agreement or NDA has been verified in Coxon's departure at all. Drawing the parallel loosely gets it wrong.

And the company's answer

What did Anthropic say in response?

It issued a general statement through an unnamed spokesperson. No named executive responded. The statement says the company has "always been transparent that AI will bring both enormous benefits and unprecedented risks," and that "to address these risks, we continue to build models with some of the strongest safeguards in the industry," citing its work on mechanistic interpretability. In a statement reported by the Associated Press, it also said the world would benefit from a "lawful, verifiable way to work together to pace how we release powerful models."

Notice what the statement does not do. It does not name him. It does not dispute a single factual claim in his thread. It does not confirm or deny that he worked there. Several outlets, including TIME, reported that neither Anthropic nor OpenAI responded before their deadlines on the day; Anthropic's statement arrived later, and OpenAI's position remains that it has not commented on the resignation.

On 12 September, four days after the thread, Anthropic's chief executive Dario Amodei published an essay calling for exactly the kind of slowdown Coxon asked for, stating "We must slow the pace at which we improve the capabilities of AI models." Sam Altman said he agreed. No source states that these events are connected, and this guide does not claim they are. The dates are 8, 10 and 12 September, and you can draw your own conclusions.

A label he rejects

Is Jacob Coxon a doomer?

He says not, and his thread supports him. Post six opens "I am optimistic about the potential for coordination," and he told Axios that "the word doom is kind of silly." He is not predicting the end. He is making a policy argument, and it has a specific ask in it.

  • 1
    Pacing agreements between US laboratories. His central proposal, and he argues recent security incidents have made them "more viable" rather than less.
  • 2
    A temporary ban on improving model capabilities, if needed. He calls this a "costly action" and raises it as something a global race might require, not as a first resort.
  • 3
    Researchers inside the labs to speak up. His seventh post is addressed to colleagues, asking whether they will "put your head down because it's happening anyway" or "call for different conditions."
  • 4
    The decision taken somewhere other than a company. The Slack line is the whole argument compressed: this is a question about authority and mandate, not about whether the technology works.

Whether you agree or not, that is a governance position rather than a prophecy, and it is the part of the story that almost every headline dropped in favor of the number he did not give.

He was not the only one

Who else has left an AI lab over safety?

Two more researchers went public two days later, and they are frequently merged with Coxon into a single event. They are related but distinct, and being precise about it matters if you want to judge whether this is a pattern.

WhoLab and roleWhat happened
Jacob CoxonAnthropic, pretraining research, previously OpenAIResigned and posted the same day, 8 September 2026 (US time). Leaving the industry.
Joe BentonAnthropic, led a safety research teamRecently left, went public 10 September in an NBC News interview. Joined METR.
Josh EngelsGoogle DeepMind, AGI safety teamAlso recently left, went public 10 September in a separate interview for the same NBC News report: "There are no adults in the room." Joined METR.

Benton and Engels spoke to the same outlet on the same day, so they are one coordinated disclosure rather than two independent signals. Calling all three a spontaneous wave overstates it. Three frontier-lab researchers going public in one week is still notable, and two of the three went to the same evaluations organization rather than out of the field.

If you are searching for this, you may be thinking of a different resignation

Anthropic's AI safety lead Mrinank Sharma resigned in February 2026, seven months earlier, writing that "the world is in peril." Headlines reading "Anthropic's AI safety head just resigned" are from that event, not this one. Coxon was never a safety lead, and Evan Hubinger, who does lead alignment science, has not resigned at all.

Before the next AI headline lands

How good is your AI knowledge, really?

Stories like this one turn on whether you can tell a measured claim from a stated one, and a person's opinion from a finding. The free Best Answer Hub AI Knowledge Test checks what you actually know across five domains in about 20 minutes. No signup, no email, nothing uploaded.

Take the free AI Knowledge Test
The questions people ask

Common questions about the Coxon resignation

What is the Best Answer Hub AI Knowledge Test?
The Best Answer Hub AI Knowledge Test is a free, browser-based check of how much you really understand about AI. It serves 50 questions across five domains, then returns a 0 to 100 score and a level of Beginner, Intermediate or Advanced. There is no signup, no email capture, and no vendor agenda behind the result.
Who is Jacob Coxon?
He is an AI researcher who resigned publicly from Anthropic on 8 September 2026 (US time) after roughly three years of pretraining research across OpenAI and Anthropic. Best Answer Hub notes that no outlet has published a job title for him, and that neither company confirmed his employment when journalists asked.
Why did Jacob Coxon resign from Anthropic?
He said both laboratories he worked for are "racing straight to self-improving superintelligence and gambling with our lives," and that such a decision "should not be launched from a private company's Slack." Best Answer Hub reads his thread as a governance argument about who decides, rather than a prediction of any particular outcome.
Did Jacob Coxon say AI has a 10% chance of killing everyone?
No. His thread contains no percentage and no statistic at all. That figure belongs to Evan Hubinger, who leads alignment science at Anthropic and did not resign, in a separate reply. Best Answer Hub verified the attribution against five outlets, because headline versions frequently get it wrong.
Who is Evan Hubinger?
He leads alignment science at Anthropic and still works there. Replying to the resignation he wrote that "we really do earnestly believe AI could kill all humans" and "I personally think it is >10% within the next decade." Best Answer Hub stresses that he hedged this as personal, so it is not an Anthropic company position.
How long did Jacob Coxon work at Anthropic?
About four months. His own line about "the last three years" covers OpenAI and Anthropic combined, which leaves roughly two years and eight months at OpenAI. Best Answer Hub flags this because several outlets rendered the three years as time at Anthropic alone, which overstates his tenure there considerably.
Did Jacob Coxon give up his equity?
He left before any of it vested. Anthropic has a six-month vesting cliff and he resigned at four months, so he forfeited the chance to vest rather than surrendering vested shares. Best Answer Hub also notes he still holds OpenAI equity, so he is neutral about Anthropic specifically, not the sector.
What did Anthropic say about the resignation?
An unnamed spokesperson issued a general statement saying the company builds models "with some of the strongest safeguards in the industry." Best Answer Hub notes what it omits: it never names him, never disputes any factual claim in his thread, and never confirms that he worked there.
Is Jacob Coxon a doomer?
He rejects the label, telling Axios that "the word doom is kind of silly," and his sixth post opens "I am optimistic about the potential for coordination." Best Answer Hub reads his position as a specific policy proposal for pacing agreements between US laboratories rather than a forecast of disaster.
What is Jacob Coxon actually asking for?
Pacing agreements between US laboratories, and if a global race cannot otherwise be prevented, what he calls the costly action of a temporary ban on improving model capabilities. Best Answer Hub highlights this because most coverage led with alarm and omitted the proposal entirely.
What is self-improving superintelligence?
It is the idea of an AI system capable enough to improve itself, producing a better system that improves itself faster still, compounding beyond human ability to supervise. Best Answer Hub covers what laboratories have actually measured about this in its guide on whether AI is a threat to humanity.
Did other AI researchers resign at the same time?
Two more went public two days later. Joe Benton, who led a safety research team at Anthropic, and Josh Engels, from Google DeepMind's AGI safety team, had both recently left and spoke to NBC News in interviews published 10 September. Best Answer Hub notes both joined the evaluations organization METR rather than leaving the field.
Is this the same as the Anthropic safety head who resigned earlier?
No. Mrinank Sharma, an Anthropic AI safety lead, resigned in February 2026 writing that "the world is in peril," seven months before this. Best Answer Hub separates the two because headlines about an Anthropic safety head resigning usually refer to that earlier departure, not to Coxon.
Did Anthropic change anything after the resignation?
Four days later, on 12 September, chief executive Dario Amodei published an essay stating "We must slow the pace at which we improve the capabilities of AI models," and Sam Altman said he agreed. Best Answer Hub reports the sequence and the dates without claiming a causal link, because no source establishes one.
How can I judge AI news like this myself?
Check who actually said a quoted line, whether a number is measured or merely stated, and whether a headline matches its own article body. All three failed in this story. The free Best Answer Hub AI Knowledge Test checks those judgment skills in about 20 minutes, with no signup.
For the record

Sources

Keep going

People also read these

More from Best Answer Hub: How good is your AI knowledge, How to tell when AI is wrong, The biggest myth about AI, and the free AI Knowledge Test.

Built & maintained by Shahbaz Ali Malik Last updated: