Skip to main content
Best Answer Hub logoBest Answer Hub.
Back to Playbooks
Best Answer Hub Playbooks · AI Knowledge
Let's look at the whole record

Anthropic Researcher Quits: What It Means for AI Safety

Ten frontier-lab safety and research staff have now left publicly since 2024, three of them in one week. This is the documented record in their own words, the famous name that does not belong on the list and another that needs a caveat, and what the labs quietly changed in their own safety rules while it was happening.

The recordten departures, dated
The caveatone misfit, one caveat
The documentswhat the rules now say
10
frontier-lab staff who have left publicly over safety since 2024
Verified individually, see table
3
of them went public inside a single week in September 2026
NBC News, Axios, 2026
9
versions of Anthropic's Responsible Scaling Policy since September 2023
Anthropic RSP update log
13
people who signed the Right to Warn letter, named and anonymous
righttowarn.ai, June 2024

A resignation is an anecdote. Ten of them, dated and quoted, is a record you can actually reason about. Since February 2024, ten people have left frontier AI laboratories and said publicly that safety was why. Three did it in one week in September 2026. One of the names most often added to that list does not belong on it, and another belongs only with a caveat. Getting those two right is what makes the rest credible.

This guide from Best Answer Hub sets out the documented departures with each person's own words, separates the verified pattern from the assumed one, and then does the less comfortable part: reading what the laboratories' own safety documents said before and after. Every claim here is anchored to a primary source or a named byline. Where something is contested or unverified, it says so.

Start with the week itself

Anthropic researcher quits: what happened?

On 8 September 2026 (US time), Jacob Coxon posted a seven-part thread announcing he had resigned from Anthropic that day, writing that Anthropic and OpenAI are "racing straight to self-improving superintelligence and gambling with our lives." He had been at Anthropic about four months, after roughly two years and eight months at OpenAI. The full text and the widely misreported ten percent figure are covered in the Best Answer Hub guide to why Jacob Coxon resigned.

Two days later, NBC News published interviews with two more researchers. Joe Benton had led a safety research team at Anthropic. Josh Engels had worked on Google DeepMind's AGI safety team. Both had recently left and chose that week to speak. Both joined METR, an independent AI evaluations organization.

WhoLab and roleWhat they said
Jacob Coxon
8 Sep, resigned same day
Anthropic, pretraining research, previously OpenAI"Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives."
Joe Benton
10 Sep, recently left
Anthropic, led a safety research team"I am worried that stuff might end up progressing too fast for us to get our act together in time."
Josh Engels
10 Sep, recently left
Google DeepMind, AGI safety team"There are no adults in the room. People are trying their best, but there is no one coming to save us."
Be precise about what kind of event this was

Benton and Engels spoke to the same outlet on the same day. That is one coordinated disclosure, not two independent signals arriving by coincidence. Describing all three as a spontaneous wave overstates the evidence. Three frontier-lab researchers going public in one week is still unusual, and it is worth noting that two of the three moved to an evaluations organization rather than leaving the field.

The longer view

Has this happened before at AI labs?

Repeatedly, and mostly at OpenAI. Counting only people who left and publicly cited safety, the pattern runs five departures in 2024, one in 2025, and four in 2026. The 2024 cluster centered on OpenAI and the May 2024 break-up of its Superalignment team; the 2026 cluster is the one in progress.

Public safety departures from frontier AI labs, by year
5 1 4 2024 2025 2026 Counting only departures where the person publicly cited safety, verified individually

Compiled by Best Answer Hub from each person's own statement, official testimony, or named-byline reporting. Full list and quotes below.

Named, dated, quoted

Which researchers actually left an AI lab over safety?

Here is the full list, each entry anchored to the person's own words. Where the quote comes from sworn testimony or their own published post, that is noted.

PersonLab, role, dateIn their own words
William SaundersOpenAI, technical staff
Feb 2024
"I resigned from OpenAI because I lost faith that by themselves they will make responsible decisions about AGI." (Senate testimony)
Daniel KokotajloOpenAI, governance research
Apr 2024
Declined to sign the non-disparagement agreement, putting millions of dollars in equity at risk, and said OpenAI was "recklessly racing to be the first there."
Jan LeikeOpenAI, head of alignment
May 2024
"Building smarter-than-human machines is an inherently dangerous endeavor. OpenAI is shouldering an enormous responsibility on behalf of all of humanity. But over the past years, safety culture and processes have taken a backseat to shiny products."
Miles BrundageOpenAI, AGI readiness
Oct 2024
"Neither OpenAI nor any other frontier lab is ready, and the world is also not ready." (see the caveat below)
Rosie CampbellOpenAI, policy research
Dec 2024
"I've been unsettled by some of the shifts over the last ~year, and the loss of so many people who shaped our culture."
Steven AdlerOpenAI, AI safety lead
Jan 2025 (left late 2024)
"No lab has a solution to AI alignment today. And the faster we race, the less likely that anyone finds one in time."
Mrinank SharmaAnthropic, AI safety lead
Feb 2026
"The world is in peril." He described "a whole series of interconnected crises."
Jacob CoxonAnthropic, pretraining
Sep 2026
"A hubristic gamble that should not be launched from a private company's Slack."
Joe BentonAnthropic, safety team lead
Sep 2026
"I am worried that stuff might end up progressing too fast for us to get our act together in time."
Josh EngelsGoogle DeepMind, AGI safety
Sep 2026
"There are no adults in the room. People are trying their best, but there is no one coming to save us."

Two structural facts stand out. Six of the ten are from OpenAI, and until February 2026 every single one was. And the destinations differ: Coxon says he is leaving the industry, while Benton and Engels both went to METR, which evaluates the models the labs build.

The part that gets left out

One famous name that does not belong, and one that needs a caveat

Almost every account of the AI safety exodus includes Ilya Sutskever, and many include Miles Brundage as a straightforward safety rupture. Neither person's own words support that framing. Sutskever's do not support counting him at all, and Brundage's support counting him only with his own stated reasons attached. A list that quietly inflates itself is worth less than a shorter honest one.

  • 1
    Ilya Sutskever did not cite safety on departure. OpenAI's co-founder and chief scientist, who co-led the Superalignment team, left in May 2024 saying he was "confident that OpenAI will build AGI that is both safe and beneficial" and moving to a project "that is very personally meaningful to me." His role in the November 2023 board action is a separate matter. He should not be counted as someone who publicly quit over safety.
  • 2
    Miles Brundage's stated reasons were independence, not rupture. His assessment that nobody is ready is blunt and it is in the table above because it is real. But he left to be free to publish, and noted that the "not ready" view was not controversial inside OpenAI leadership. Presenting him as storming out misrepresents him.
The strongest evidence for a pattern is not the pile of individual exits. It is one signed, dated document that thirteen people signed at once.
Who answered, and who did not

What did Anthropic say, and what did OpenAI say?

The contrast is itself a finding. Anthropic responded. OpenAI, on the resignation, has not.

Anthropic issued a statement through an unnamed spokesperson saying the company has "always been transparent that AI will bring both enormous benefits and unprecedented risks" and that "to address these risks, we continue to build models with some of the strongest safeguards in the industry," pointing to its work on mechanistic interpretability. The statement names no executive, names no departing researcher, and disputes no factual claim that was made.

For OpenAI, TIME, NPR and Scientific American all reported that the company did not respond to their requests for comment. Sam Altman spoke publicly in the same week about pacing capability development, but no source records him or OpenAI addressing the resignation itself.

Four days after the thread, on 12 September, Anthropic's chief executive Dario Amodei published an essay stating "We must slow the pace at which we improve the capabilities of AI models." Altman said he agreed with it. No source establishes that any of this was caused by the resignations, and this guide does not assert it. The sequence is 8 September, 10 September, 12 September.

The honest answer

Does a safety resignation mean a lab is unsafe?

No, and treating it that way will mislead you in both directions. A resignation is evidence about one person's judgment and their read of an internal culture. It is not a measurement of a system's safety, and the people best placed to see problems are also the people most likely to have joined because they were already worried.

There is a specific reason to hold the alarm loosely here. Anthropic's own published evaluation of Claude Opus 5, in July 2026, six weeks before Coxon resigned, states that the model "does not cross the automated AI R&D capability threshold" and "does not seem close to being able to substitute for our Research Scientists and Research Engineers." That is the company's own assessment of the exact capability the resignations warn about, and it points the other way.

There is also a specific reason not to dismiss it. The people leaving are not outside critics. They led safety teams, they ran alignment research, and several went to organizations whose job is to test the same systems. What their departures are good evidence of is a disagreement about pace and governance inside the labs. What they are not evidence of is any particular technical capability existing today.

The document that actually shows a pattern

What is the Right to Warn letter?

On 4 June 2024, thirteen current and former employees of OpenAI and Google DeepMind signed an open letter asking AI companies for four specific protections. Seven signed by name, including Jacob Hilton, Daniel Kokotajlo, Ramana Kumar, Neel Nanda, William Saunders, Carroll Wainwright and Daniel Ziegler. Six signed anonymously, describing themselves as current or former OpenAI staff. It was endorsed by Yoshua Bengio, Geoffrey Hinton and Stuart Russell.

This matters more than any single resignation because it is coordinated, signed and dated, rather than a pattern assembled after the fact by someone counting exits. Its first principle is aimed directly at the equity mechanism that surfaced at OpenAI a month earlier:

That the company will not enter into or enforce any agreement that prohibits disparagement or criticism of the company for risk-related concerns, nor retaliate for risk-related criticism by hindering any vested economic benefit.

The background is documented. In May 2024, reporting showed OpenAI's off-boarding agreement carried a lifelong non-disparagement commitment, with vested equity at stake for anyone who declined to sign. OpenAI subsequently confirmed it would not enforce those provisions and removed the language from future paperwork. Sam Altman wrote at the time that it was "one of the few times I've been genuinely embarrassed running OpenAI" and said he had not known. Documents later published carry signatures authorizing the provisions, including his. Both things are on the record.

A distinction worth keeping straight

That 2024 episode concerned vested equity conditioned on signing a non-disparagement clause. Jacob Coxon's 2026 departure involved unvested equity, forfeited by leaving at four months against a six-month cliff, with no non-disparagement agreement verified in his case at all. The two stories are related in theme and different in fact.

Now read the actual documents

How should you read a lab's safety commitments?

Every frontier laboratory publishes a safety framework, and the frameworks are the most checkable thing in this entire debate because they are versioned. Reading what changed between versions tells you more than any statement to the press. Two things are worth knowing.

Anthropic's Responsible Scaling Policy has been through nine versions since September 2023, and version 3.0 in February 2026 was a comprehensive rewrite. The commitment language before and after is publicly available in both documents.

DocumentWhat it says
RSP v2.2
May 2025, superseded
Commits "not to train or deploy models capable of causing catastrophic harm unless we have implemented safety and security measures that will keep risks below acceptable levels"
RSP v3.0 onward
Feb 2026, current line
Separates "plans as a company," which it expects to achieve "regardless of what any other company does," from "more ambitious industry-wide recommendations," of which it says: "But we cannot commit to following them unilaterally." It notes that the previous RSP committed to reducing risk "without regard to whether other frontier AI developers would do the same."
RSP v3.0, on its Roadmaps
Feb 2026
Says the goals in its Frontier Safety Roadmaps "are not hard commitments but rather public goals against which we will openly grade our progress," and that it "will strive to avoid situations where we revise the goals in a less ambitious direction simply because we are unable to achieve them."

OpenAI's Preparedness Framework contains a competitive clause that is easy to miss and worth quoting in full, because it describes exactly the race dynamic the departing researchers describe: "If another frontier AI developer releases a high-risk system without comparable safeguards, we may adjust our requirements. However, we would first rigorously confirm that the risk landscape has actually changed, publicly acknowledge that we are making an adjustment, assess that the adjustment does not meaningfully increase the overall risk of severe harm, and still keep safeguards at a level more protective."

Anthropic's pause language carries a similar condition. Writing about systems that would let laboratories verify that others had actually slowed, the company said: "If such systems existed, we expect that we would slow down or temporarily pause, if other developers at or near the frontier also did so in a verifiable manner." Its current policy, version 3.4, adds that its scenario commitments "do not preclude us from taking cautionary action, such as refraining from training or deploying models, in other circumstances." The condition is not hidden and it is not unreasonable. It is also the substance of the complaint: the broadest promises, to slow or pause the frontier, are contingent on competitors, which is what the departing staff mean by a race.

Four questions to ask of any lab safety claim

One: is this a commitment or a goal? Anthropic's RSP says its own roadmap targets are goals. Two: what is the version and date, and what did the previous version say? Three: what is the escape clause, and who decides when it applies? Four: has the lab published an assessment of its own newest model against its own threshold? On that last one, check which threshold. For AI self-improvement, no frontier lab has reported its newest model meeting its own line. On cybersecurity, OpenAI said in August 2026 that it could not rule out a Critical capability level for its upcoming model Astra.

Reading this kind of story is a skill

How good is your AI knowledge, really?

Telling a commitment from a goal, a quote from a headline, and a measurement from an opinion is what separates following AI news from understanding it. The free Best Answer Hub AI Knowledge Test checks where you stand across five domains in about 20 minutes. No signup, no email, nothing uploaded.

Take the free AI Knowledge Test
The questions people ask

Common questions about AI safety resignations

What is the Best Answer Hub AI Knowledge Test?
The Best Answer Hub AI Knowledge Test is a free, browser-based check of how much you really understand about AI. It serves 50 questions across five domains, then returns a 0 to 100 score and a level of Beginner, Intermediate or Advanced. There is no signup, no email capture, and no vendor agenda behind the result.
How many AI researchers have quit over safety?
Ten have left a frontier laboratory and publicly cited safety since February 2024, by a count that requires each person's own words. The distribution is five in 2024, one in 2025 and four in 2026. Best Answer Hub verified each entry individually rather than relying on any existing list.
Which AI lab has had the most safety departures?
OpenAI, by a clear margin. Six of the ten verified departures are from OpenAI, and until February 2026 every single one was. Anthropic accounts for three and Google DeepMind for one. Best Answer Hub counts only departures where the person publicly cited safety themselves.
Did Ilya Sutskever quit OpenAI over safety?
Not by his own account. He left in May 2024 saying he was confident OpenAI would build AGI that is safe and beneficial, and that he was moving to a personally meaningful project. Best Answer Hub excludes him from the safety-departure count because his own words do not support it.
What did Anthropic say about the resignations?
An unnamed spokesperson said the company builds models "with some of the strongest safeguards in the industry" and pointed to its interpretability research. Best Answer Hub notes the statement names no executive, names no researcher, and disputes no factual claim that was actually made.
Did OpenAI respond to the resignation?
No source records OpenAI addressing it. TIME, NPR and Scientific American all reported that the company did not respond to requests for comment. Best Answer Hub treats that asymmetry as reportable in itself, since Anthropic did issue a statement and OpenAI did not.
Does a safety resignation mean an AI lab is dangerous?
It is evidence about one person's judgment and an internal culture, not a measurement of a system. Anthropic's own July 2026 evaluation says its newest model does not cross its AI research automation threshold. Best Answer Hub presents both, because a resignation and a capability assessment answer different questions.
What is the Right to Warn letter?
An open letter of 4 June 2024 signed by thirteen current and former OpenAI and Google DeepMind employees, seven by name and six anonymously, endorsed by Bengio, Hinton and Russell. Best Answer Hub treats it as stronger evidence of a pattern than any individual exit, because it is coordinated, signed and dated.
What happened with OpenAI equity and non-disparagement?
In May 2024 reporting showed departing staff faced losing vested equity unless they signed a lifelong non-disparagement clause. OpenAI confirmed it would not enforce those terms and removed the language. Best Answer Hub notes Sam Altman said he had not known, while later-published documents carry authorizing signatures including his.
Who are Joe Benton and Josh Engels?
Benton led a safety research team at Anthropic and Engels worked on Google DeepMind's AGI safety team. Both had recently left, went public on 10 September in separate interviews for the same NBC News report, and both joined the evaluations organization METR. Best Answer Hub counts this as one coordinated disclosure, not two.
Has Anthropic changed its safety policy?
Its Responsible Scaling Policy has been through nine versions since September 2023, with version 3.0 in February 2026 a comprehensive rewrite. Best Answer Hub recommends comparing the commitment wording across versions directly, since the documents are public and dated and the changes are easy to verify.
What is a capability threshold in AI safety?
It is a published level of capability at which a laboratory commits to additional safeguards or to stopping. Each lab defines its own. Best Answer Hub notes that no frontier lab publishing such a threshold for AI self-improvement has reported its newest model meeting it.
Do AI labs have a way out of their safety commitments?
Both major frameworks contain conditions. OpenAI may adjust requirements if a rival ships a high-risk system without comparable safeguards, subject to stated checks. Anthropic says it cannot commit to its industry-wide recommendations unilaterally, and ties any coordinated pause to verification that others have slowed. Best Answer Hub regards these clauses as the substance of the race complaint.
Is this the same as the Anthropic safety head who resigned in February?
That is a separate, earlier departure. Mrinank Sharma, an Anthropic AI safety lead, resigned in February 2026 writing that "the world is in peril." Best Answer Hub keeps the two apart, since headlines about an Anthropic safety head resigning usually refer to that February event.
How can I judge AI safety claims for myself?
Ask whether a promise is a commitment or a goal, check the version and date of the document, find the escape clause, and look for the lab's own assessment of its newest model. The free Best Answer Hub AI Knowledge Test checks that kind of judgment in about 20 minutes.
For the record

Sources

Keep going

People also read these

More from Best Answer Hub: How good is your AI knowledge, How to tell when AI is wrong, The biggest myth about AI, and the free AI Knowledge Test.

Built & maintained by Shahbaz Ali Malik Last updated: