Case TS-5D47EE6610 Sept 2026safetyCompound claim

AI

“Anthropic researcher Jacob Coxon resigned, stating that both OpenAI and Anthropic are acting irresponsibly by racing toward self-improving superintelligence and 'gambling with our lives,' and that people building AI genuinely believe it could kill everyone by the end of the decade." Secondary claim in the caption: "He also said Anthropic…”

Plain restatementA researcher named Jacob Coxon resigned from Anthropic and posted publicly that neither OpenAI nor Anthropic is acting responsibly, that they are racing to self-improving superintelligence and gambling with our lives, and that people building AI earnestly believe it could kill everyone by the end of the decade.

Mostly accurateConfidence High
What this verdict means →

Distortion codes this site does not recognise yet: misattribution, rumor_as_fact, capability_extrapolation. Not collectible until the field guide has an entry.

This one is essentially real. Jacob Coxon did resign from Anthropic on September 8, 2026, and posted a thread on X saying neither Anthropic nor OpenAI is acting responsibly, that they are "racing straight to self-improving superintelligence and gambling with our lives," and that people building AI earnestly believe it could kill us all by the end of the decade. Those are his exact words, quoted identically by Bloomberg, TIME, TechCrunch and others, and he gave interviews to the Wall Street Journal and TIME. The strongest corroboration is that a current Anthropic alignment lead, Evan Hubinger, publicly replied that Coxon was correct and put his own odds of AI killing all humans in the next decade above 10 percent. There is one real error in the post: the caption says Coxon stated Anthropic has no plan for aligning superintelligence and is not clearly on track, but that sentence was written by Hubinger, not Coxon, and the caption credits only Coxon's account. The post also leaves out that Hubinger said in the same breath that he considers the risk from today's models low, with his concern being future self-improving systems. Anthropic declined to comment and has not confirmed or denied anything, and none of this establishes that the underlying forecast about extinction is correct, only that these people said it.

The drift / as claimed vs as evidenced

[drifted from the evidence:] Anthropic researcher Jacob Coxon resigned, [drifted from the evidence:] stating that [drifted from the evidence:] both OpenAI [drifted from the evidence:] and Anthropic [drifted from the evidence:] are acting [drifted from the evidence:] irresponsibly by racing [drifted from the evidence:] toward self-improving superintelligence and 'gambling with our lives,' and that people building AI [drifted from the evidence:] genuinely believe it could kill everyone by the end of the decade." [drifted from the evidence:] Secondary claim in the caption: "He also said Anthropic does not yet have a clear solution for aligning superintelligent AI and is not clearly on track to solve the problem.


[added by the neutral restatement:] A researcher [added by the neutral restatement:] named Jacob Coxon resigned [added by the neutral restatement:] from Anthropic and posted publicly that [added by the neutral restatement:] neither OpenAI [added by the neutral restatement:] nor Anthropic [added by the neutral restatement:] is acting [added by the neutral restatement:] responsibly, that they are racing [added by the neutral restatement:] to self-improving superintelligence and gambling with our lives, and that people building AI [added by the neutral restatement:] earnestly believe it could kill everyone by the end of the decade.

Red-tinted words in the claim drifted from the evidence. Green-tinted words are what a neutral restatement needs.

The trace / claim to source

Where it appeared
misattribution
rumor_as_fact
⌿ Omitted qualifier
A load-bearing condition from the source quietly disappears from the claim.
capability_extrapolation
Source
Coxon's own X thread (@hilbertspaess), 2026-09-08/09
Secondary sourcenamed-outlet journalism, direct interview with the subject
TIME interview with Coxon, 2026-09-09
Secondary sourcenamed-outlet journalism
Bloomberg, "Anthropic Worker Quits Over AI Firms 'Gambling With Our Lives'", 2026-09-09
Secondary sourcenamed-outlet journalism, first interview
Wall Street Journal exclusive interview (accessed via multiple outlets citing and quoting it, not retrieved directly)
Secondary sourcenamed-outlet tech journalism
TechCrunch, 2026-09-09
Secondary sourcenamed-outlet journalism
Deadline, 2026-09-09, reproducing the embedded X thread
Secondary sourcenamed-outlet journalism
South China Morning Post, 2026-09-09
Secondary sourcemixed-quality aggregation of the same thread
Quartz, Newsweek, Fast Company, CoinDesk, IBTimes, MercoPress, Daily Sabah
Primary sourceAnthropic Alignment Science lead, account labelled "opinions my own"
Evan Hubinger (@EvanHub) X post of 2026-09-09, quote-replying to Coxon
● Primary source found
What is true
  • Jacob Coxon publicly announced his resignation from Anthropic in a thread on X under the handle @hilbertspaess on 2026-09-08, and the thread went viral.
  • The quoted phrase "gambling with our lives" is verbatim, not a paraphrase. The exact wording is "They are racing straight to self-improving superintelligence and gambling with our lives."
  • The characterisation of both companies is his, not the post's invention. He wrote that "Neither company is acting responsibly."
  • The extinction-belief line is his, near-verbatim. He said the people racing to build this technology "earnestly believe it could kill us all by the end of the decade."
  • He worked on pretraining at both labs. SCMP reports he spent the past three years pretraining AI models, first at OpenAI and then, this year, at Anthropic.
  • He is leaving the industry, not just the company. SCMP reports he decided to leave the industry, accusing both US companies of "gambling with our lives."
  • The substance was endorsed by a serving Anthropic alignment lead, which is a stronger corroboration than any of the press coverage.
What is misleading
  • Misattribution: the Instagram caption says "He also said Anthropic does not yet have a clear solution for aligning superintelligent AI and is not clearly on track to solve the problem," and cites only "Jacob Coxon / X" as the source. That sentence is Evan Hubinger's, a current Anthropic alignment science lead replying to Coxon, not Coxon's. Assigning it to a departing employee makes it read as an ex-insider's parting accusation. It is in fact a serving alignment lead's on-record statement about his own employer, which is a materially different and arguably more significant thing. The caption's single source line erases the second speaker entirely.
  • Omitted qualifier: the caption and headline framing carry none of the scoping Hubinger attached in the same breath. He was explicit that his concern is not about deployed models. He cited Anthropic's latest risk report saying present-model risk is "low," with his worry attaching to superintelligence arising from recursive self-improvement. The post presents only the alarming half.
  • Capability extrapolation, in the underlying claim rather than the post's handling of it: statements such as systems that "can hack anything" and "revolutionize any field overnight" are one researcher's forecast, not a demonstrated capability. The post presents them as reported news rather than as prediction. Note this is a distortion inherited from the source thread, and the post did quote it accurately.
What is uncertain
  • I did not retrieve Coxon's X thread directly. Its wording is established by verbatim quotation and embedded reproduction across Bloomberg, TIME, TechCrunch, Deadline and SCMP, which agree word for word, but the artifact itself was not loaded.
  • I did not retrieve the WSJ article directly. Its contents are known here only through outlets quoting it, so it is recorded as reported-by, not verified.
  • Anthropic has not confirmed Coxon's employment, title, tenure or departure. It declined to comment. That is neither confirmation nor denial. Fast Company hedged accordingly, describing him as someone who claims to have worked at both companies.
  • Specific biographical details carried by single outlets, including his age of 27, a Cambridge degree, OpenAI tenure from 2023 to July 2026, and core contributor status on a named model, rest on one report each and are not independently corroborated.
  • The truth of the forecast itself, that AI could kill everyone by 2030, is not something this investigation assesses. The claim under test is that he said it, and he did.
  • View counts cited on the post (67.6M) are a moving figure. Outlets recorded 76M, 79M and 90M at different hours of the same day, so the number is a snapshot, not a fixed fact.
Evidence summary

Multiple independently reporting, editorially accountable outlets report the same event within the same 24 hour window, quoting the same words. TechCrunch reports that Jacob Coxon, a researcher who said in a social media post Tuesday evening that he spent the last three years working on pretraining research at both OpenAI and Anthropic, accused the firms of failing to act responsibly, and said the people racing to build this technology "earnestly believe it could kill us all by the end of the decade." TechCrunch quotes the thread directly: "They are racing straight to self-improving superintelligence and gambling with our lives." Bloomberg reports that an artificial intelligence researcher has resigned from Anthropic PBC and called on other staffers to rethink their work, citing his concern that the company and its top competitor OpenAI are acting irresponsibly. TIME reports that for three years Coxon helped train increasingly powerful AI systems at OpenAI and Anthropic, that on Sept. 8 he walked away, and quotes the post: "I resigned from Anthropic today. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives." The most important corroboration is not journalistic. A currently serving Anthropic alignment lead publicly endorsed the substance. Evan Hubinger wrote: "Jacob is correct here, we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to." Hubinger went on to cite Anthropic's own latest risk report saying the risk from present models is "low," while saying he is worried about superintelligence arising from recursive self-improvement. On the secondary caption claim, the "no plan for alignment" sentence belongs to Hubinger, not Coxon. The Anthropic alignment lead is the one who wrote "I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to." Quartz frames it the same way: a colleague at Anthropic responded by saying the company has no plan to address alignment risks from superintelligence.

Complete reasoning
As of 2026-09-10, every element of the primary claim text checks out against the subject's own public posting, two on-the-record interviews, and reporting by multiple independent named outlets, with the quoted phrases matching verbatim rather than paraphrased. "Accurate" was considered and rejected only because the accompanying caption misattributes Evan Hubinger's "no plan to solve alignment for superintelligence" statement to Coxon, and cites Coxon's X account as the sole source for it. That is a genuine misattribution, but it sits in the caption rather than in the operative proposition, and the operative proposition survives intact, so the contradiction test does not close the accurate family. "Source exists but framing is misleading" was considered and rejected because the post does not exaggerate or recontextualise what Coxon said; the headline quote is his own words and the summary is faithful. Confidence is High because a serving Anthropic alignment lead publicly wrote "Jacob is correct here" using "we," which corroborates insider status far more strongly than press repetition, and because no party has denied any element.
Use this case

The reply is formatted for pasting into the thread where the claim is circulating.

Compact share page: ai.trueseeker.com/s/5d47ee66ce9b/E20VfAburutukBPP3B0BYrN2D4z

Ask this case

Answers come only from the case file above; nothing is added.

Did Jacob Coxon really resign from Anthropic and say this?

Yes. Coxon posted a thread on X on September 8, 2026 announcing his resignation and saying neither Anthropic nor OpenAI is acting responsibly. The phrase 'gambling with our lives' and the line about believing AI could kill everyone by the end of the decade are his exact words, quoted the same way by Bloomberg, TIME, TechCrunch and others.

Did Coxon say Anthropic has no plan for aligning superintelligent AI?

No. That statement was written by Evan Hubinger, a current Anthropic alignment lead who replied to Coxon's post, not by Coxon himself. The caption in question credits it only to Coxon, which misattributes it.

Does anyone at Anthropic agree with Coxon's warning?

Yes. Evan Hubinger publicly replied that Coxon was correct and said he personally puts the odds of AI killing all humans within the next decade above 10 percent. This is treated as the strongest corroboration since it comes from someone still working at the company.

Is the risk from AI happening right now, according to these statements?

No. Hubinger specifically said the risk from Anthropic's current models is low, based on the company's own risk report. His stated concern is about future superintelligent systems arising from recursive self-improvement, not present-day AI.

Has Anthropic confirmed any of this?

No. Anthropic declined to comment and has not confirmed or denied Coxon's employment, title, tenure, or resignation. The investigation could not independently verify these details beyond media reporting.

Similar cases on record