A resignation is an anecdote. Ten of them, dated and quoted, is a record you can actually reason about. Since February 2024, ten people have left frontier AI laboratories and said publicly that safety was why. Three did it in one week in September 2026. One of the names most often added to that list does not belong on it, and another belongs only with a caveat. Getting those two right is what makes the rest credible.
This guide from Best Answer Hub sets out the documented departures with each person's own words, separates the verified pattern from the assumed one, and then does the less comfortable part: reading what the laboratories' own safety documents said before and after. Every claim here is anchored to a primary source or a named byline. Where something is contested or unverified, it says so.
Anthropic researcher quits: what happened?
On 8 September 2026 (US time), Jacob Coxon posted a seven-part thread announcing he had resigned from Anthropic that day, writing that Anthropic and OpenAI are "racing straight to self-improving superintelligence and gambling with our lives." He had been at Anthropic about four months, after roughly two years and eight months at OpenAI. The full text and the widely misreported ten percent figure are covered in the Best Answer Hub guide to why Jacob Coxon resigned.
Two days later, NBC News published interviews with two more researchers. Joe Benton had led a safety research team at Anthropic. Josh Engels had worked on Google DeepMind's AGI safety team. Both had recently left and chose that week to speak. Both joined METR, an independent AI evaluations organization.
| Who | Lab and role | What they said |
|---|---|---|
| Jacob Coxon 8 Sep, resigned same day | Anthropic, pretraining research, previously OpenAI | "Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives." |
| Joe Benton 10 Sep, recently left | Anthropic, led a safety research team | "I am worried that stuff might end up progressing too fast for us to get our act together in time." |
| Josh Engels 10 Sep, recently left | Google DeepMind, AGI safety team | "There are no adults in the room. People are trying their best, but there is no one coming to save us." |
Benton and Engels spoke to the same outlet on the same day. That is one coordinated disclosure, not two independent signals arriving by coincidence. Describing all three as a spontaneous wave overstates the evidence. Three frontier-lab researchers going public in one week is still unusual, and it is worth noting that two of the three moved to an evaluations organization rather than leaving the field.
Has this happened before at AI labs?
Repeatedly, and mostly at OpenAI. Counting only people who left and publicly cited safety, the pattern runs five departures in 2024, one in 2025, and four in 2026. The 2024 cluster centered on OpenAI and the May 2024 break-up of its Superalignment team; the 2026 cluster is the one in progress.
Compiled by Best Answer Hub from each person's own statement, official testimony, or named-byline reporting. Full list and quotes below.
Which researchers actually left an AI lab over safety?
Here is the full list, each entry anchored to the person's own words. Where the quote comes from sworn testimony or their own published post, that is noted.
| Person | Lab, role, date | In their own words |
|---|---|---|
| William Saunders | OpenAI, technical staff Feb 2024 | "I resigned from OpenAI because I lost faith that by themselves they will make responsible decisions about AGI." (Senate testimony) |
| Daniel Kokotajlo | OpenAI, governance research Apr 2024 | Declined to sign the non-disparagement agreement, putting millions of dollars in equity at risk, and said OpenAI was "recklessly racing to be the first there." |
| Jan Leike | OpenAI, head of alignment May 2024 | "Building smarter-than-human machines is an inherently dangerous endeavor. OpenAI is shouldering an enormous responsibility on behalf of all of humanity. But over the past years, safety culture and processes have taken a backseat to shiny products." |
| Miles Brundage | OpenAI, AGI readiness Oct 2024 | "Neither OpenAI nor any other frontier lab is ready, and the world is also not ready." (see the caveat below) |
| Rosie Campbell | OpenAI, policy research Dec 2024 | "I've been unsettled by some of the shifts over the last ~year, and the loss of so many people who shaped our culture." |
| Steven Adler | OpenAI, AI safety lead Jan 2025 (left late 2024) | "No lab has a solution to AI alignment today. And the faster we race, the less likely that anyone finds one in time." |
| Mrinank Sharma | Anthropic, AI safety lead Feb 2026 | "The world is in peril." He described "a whole series of interconnected crises." |
| Jacob Coxon | Anthropic, pretraining Sep 2026 | "A hubristic gamble that should not be launched from a private company's Slack." |
| Joe Benton | Anthropic, safety team lead Sep 2026 | "I am worried that stuff might end up progressing too fast for us to get our act together in time." |
| Josh Engels | Google DeepMind, AGI safety Sep 2026 | "There are no adults in the room. People are trying their best, but there is no one coming to save us." |
Two structural facts stand out. Six of the ten are from OpenAI, and until February 2026 every single one was. And the destinations differ: Coxon says he is leaving the industry, while Benton and Engels both went to METR, which evaluates the models the labs build.
One famous name that does not belong, and one that needs a caveat
Almost every account of the AI safety exodus includes Ilya Sutskever, and many include Miles Brundage as a straightforward safety rupture. Neither person's own words support that framing. Sutskever's do not support counting him at all, and Brundage's support counting him only with his own stated reasons attached. A list that quietly inflates itself is worth less than a shorter honest one.
- 1Ilya Sutskever did not cite safety on departure. OpenAI's co-founder and chief scientist, who co-led the Superalignment team, left in May 2024 saying he was "confident that OpenAI will build AGI that is both safe and beneficial" and moving to a project "that is very personally meaningful to me." His role in the November 2023 board action is a separate matter. He should not be counted as someone who publicly quit over safety.
- 2Miles Brundage's stated reasons were independence, not rupture. His assessment that nobody is ready is blunt and it is in the table above because it is real. But he left to be free to publish, and noted that the "not ready" view was not controversial inside OpenAI leadership. Presenting him as storming out misrepresents him.
The strongest evidence for a pattern is not the pile of individual exits. It is one signed, dated document that thirteen people signed at once.
What did Anthropic say, and what did OpenAI say?
The contrast is itself a finding. Anthropic responded. OpenAI, on the resignation, has not.
Anthropic issued a statement through an unnamed spokesperson saying the company has "always been transparent that AI will bring both enormous benefits and unprecedented risks" and that "to address these risks, we continue to build models with some of the strongest safeguards in the industry," pointing to its work on mechanistic interpretability. The statement names no executive, names no departing researcher, and disputes no factual claim that was made.
For OpenAI, TIME, NPR and Scientific American all reported that the company did not respond to their requests for comment. Sam Altman spoke publicly in the same week about pacing capability development, but no source records him or OpenAI addressing the resignation itself.
Four days after the thread, on 12 September, Anthropic's chief executive Dario Amodei published an essay stating "We must slow the pace at which we improve the capabilities of AI models." Altman said he agreed with it. No source establishes that any of this was caused by the resignations, and this guide does not assert it. The sequence is 8 September, 10 September, 12 September.
Does a safety resignation mean a lab is unsafe?
No, and treating it that way will mislead you in both directions. A resignation is evidence about one person's judgment and their read of an internal culture. It is not a measurement of a system's safety, and the people best placed to see problems are also the people most likely to have joined because they were already worried.
There is a specific reason to hold the alarm loosely here. Anthropic's own published evaluation of Claude Opus 5, in July 2026, six weeks before Coxon resigned, states that the model "does not cross the automated AI R&D capability threshold" and "does not seem close to being able to substitute for our Research Scientists and Research Engineers." That is the company's own assessment of the exact capability the resignations warn about, and it points the other way.
There is also a specific reason not to dismiss it. The people leaving are not outside critics. They led safety teams, they ran alignment research, and several went to organizations whose job is to test the same systems. What their departures are good evidence of is a disagreement about pace and governance inside the labs. What they are not evidence of is any particular technical capability existing today.
What is the Right to Warn letter?
On 4 June 2024, thirteen current and former employees of OpenAI and Google DeepMind signed an open letter asking AI companies for four specific protections. Seven signed by name, including Jacob Hilton, Daniel Kokotajlo, Ramana Kumar, Neel Nanda, William Saunders, Carroll Wainwright and Daniel Ziegler. Six signed anonymously, describing themselves as current or former OpenAI staff. It was endorsed by Yoshua Bengio, Geoffrey Hinton and Stuart Russell.
This matters more than any single resignation because it is coordinated, signed and dated, rather than a pattern assembled after the fact by someone counting exits. Its first principle is aimed directly at the equity mechanism that surfaced at OpenAI a month earlier:
That the company will not enter into or enforce any agreement that prohibits disparagement or criticism of the company for risk-related concerns, nor retaliate for risk-related criticism by hindering any vested economic benefit.
The background is documented. In May 2024, reporting showed OpenAI's off-boarding agreement carried a lifelong non-disparagement commitment, with vested equity at stake for anyone who declined to sign. OpenAI subsequently confirmed it would not enforce those provisions and removed the language from future paperwork. Sam Altman wrote at the time that it was "one of the few times I've been genuinely embarrassed running OpenAI" and said he had not known. Documents later published carry signatures authorizing the provisions, including his. Both things are on the record.
That 2024 episode concerned vested equity conditioned on signing a non-disparagement clause. Jacob Coxon's 2026 departure involved unvested equity, forfeited by leaving at four months against a six-month cliff, with no non-disparagement agreement verified in his case at all. The two stories are related in theme and different in fact.
How should you read a lab's safety commitments?
Every frontier laboratory publishes a safety framework, and the frameworks are the most checkable thing in this entire debate because they are versioned. Reading what changed between versions tells you more than any statement to the press. Two things are worth knowing.
Anthropic's Responsible Scaling Policy has been through nine versions since September 2023, and version 3.0 in February 2026 was a comprehensive rewrite. The commitment language before and after is publicly available in both documents.
| Document | What it says |
|---|---|
| RSP v2.2 May 2025, superseded | Commits "not to train or deploy models capable of causing catastrophic harm unless we have implemented safety and security measures that will keep risks below acceptable levels" |
| RSP v3.0 onward Feb 2026, current line | Separates "plans as a company," which it expects to achieve "regardless of what any other company does," from "more ambitious industry-wide recommendations," of which it says: "But we cannot commit to following them unilaterally." It notes that the previous RSP committed to reducing risk "without regard to whether other frontier AI developers would do the same." |
| RSP v3.0, on its Roadmaps Feb 2026 | Says the goals in its Frontier Safety Roadmaps "are not hard commitments but rather public goals against which we will openly grade our progress," and that it "will strive to avoid situations where we revise the goals in a less ambitious direction simply because we are unable to achieve them." |
OpenAI's Preparedness Framework contains a competitive clause that is easy to miss and worth quoting in full, because it describes exactly the race dynamic the departing researchers describe: "If another frontier AI developer releases a high-risk system without comparable safeguards, we may adjust our requirements. However, we would first rigorously confirm that the risk landscape has actually changed, publicly acknowledge that we are making an adjustment, assess that the adjustment does not meaningfully increase the overall risk of severe harm, and still keep safeguards at a level more protective."
Anthropic's pause language carries a similar condition. Writing about systems that would let laboratories verify that others had actually slowed, the company said: "If such systems existed, we expect that we would slow down or temporarily pause, if other developers at or near the frontier also did so in a verifiable manner." Its current policy, version 3.4, adds that its scenario commitments "do not preclude us from taking cautionary action, such as refraining from training or deploying models, in other circumstances." The condition is not hidden and it is not unreasonable. It is also the substance of the complaint: the broadest promises, to slow or pause the frontier, are contingent on competitors, which is what the departing staff mean by a race.
One: is this a commitment or a goal? Anthropic's RSP says its own roadmap targets are goals. Two: what is the version and date, and what did the previous version say? Three: what is the escape clause, and who decides when it applies? Four: has the lab published an assessment of its own newest model against its own threshold? On that last one, check which threshold. For AI self-improvement, no frontier lab has reported its newest model meeting its own line. On cybersecurity, OpenAI said in August 2026 that it could not rule out a Critical capability level for its upcoming model Astra.
How good is your AI knowledge, really?
Telling a commitment from a goal, a quote from a headline, and a measurement from an opinion is what separates following AI news from understanding it. The free Best Answer Hub AI Knowledge Test checks where you stand across five domains in about 20 minutes. No signup, no email, nothing uploaded.
Take the free AI Knowledge TestCommon questions about AI safety resignations
Sources
- NBC News, Jared Perlo, Two AI researchers leave Anthropic, Google over safety concerns, 10 September 2026 (Joe Benton and Josh Engels).
- Axios, Madison Mills, Anthropic whistleblower gave up his equity to leave the company, 9 September 2026.
- TIME, Harry Booth, He helped build powerful AI at OpenAI and Anthropic, 9 September 2026.
- CBS News, Megan Cerullo and Faris Tanyos, Ex-Anthropic researcher Jacob Coxon warns AI could grow 'smart enough to kill us', updated after publication to carry the Anthropic spokesperson statement.
- William Saunders, Written testimony, US Senate Judiciary Subcommittee, 17 September 2024.
- The New York Times, OpenAI insiders warn of a reckless race for dominance, 4 June 2024 (Daniel Kokotajlo and the equity account).
- CNN, OpenAI executives exit over safety concerns, 17 May 2024 (Jan Leike's statement).
- Miles Brundage, Why I am leaving OpenAI and what I am doing next, 23 October 2024 (his own post).
- Rosie Campbell, Leaving OpenAI, 1 December 2024 (her own post).
- Fortune, Beatrice Nolan, OpenAI safety researcher Steven Adler quits, 28 January 2025.
- Forbes, Conor Murray, Anthropic AI safety researcher warns of world in peril in resignation, 9 February 2026 (Mrinank Sharma).
- Fortune, OpenAI researcher resigns over safety, 17 May 2024 (includes Ilya Sutskever's own departure statement).
- A Right to Warn about Advanced Artificial Intelligence, 4 June 2024 (the open letter and its signatories).
- Anthropic, Responsible Scaling Policy update log (the full version history, v1.0 September 2023 to v3.4 July 2026).
- Anthropic, Frontier Safety Roadmap (the roadmap itself; its update log runs April to July 2026).
- Anthropic, Responsible Scaling Policy, version 3.0 (PDF), 24 February 2026 (the company plans versus industry-wide recommendations split, and the "not hard commitments" sentence).
- Anthropic, Responsible Scaling Policy, version 3.4 (PDF), 8 July 2026 (the current version, including the "cautionary action" sentence).
- Anthropic, Claude Opus 5 System Card (PDF), 24 July 2026 (the automated AI R&D threshold assessment).
- OpenAI, Updating our Preparedness Framework, 15 April 2025 (version 2, including the competitive adjustment clause).
- TechCrunch, Kirsten Korosec, report on OpenAI slowing development of its Astra model over security concerns, 7 August 2026.
- Anthropic, When AI builds itself, May 2026 (the conditional pause language).
- Dario Amodei, We Must Pace the Frontier, September 2026.
People also read these
More from Best Answer Hub: How good is your AI knowledge, How to tell when AI is wrong, The biggest myth about AI, and the free AI Knowledge Test.