The AI Extinction Warning Isn’t Hype, It’s a Confession

The AI Extinction Warning Isn’t Hype, It’s a Confession

On Wednesday, a pretraining researcher named Jacob Coxon resigned from Anthropic and told the internet that the two most powerful AI labs on earth are “gambling with our lives.” Resignations like this have become a genre at frontier labs. This one broke the pattern, because the people still inside didn’t deny a word of it. They confirmed it, in public, with numbers.

I’ve spent three days going through the thread, the replies, and the forum fallout, and my takeaway is simple. The argument over whether AI could end humanity is basically settled inside the labs. The live argument is who gets to keep building it anyway, and that one is being settled by game theory rather than by the people who know the most.

The reply that mattered more than the resignation

Coxon spent about three years at OpenAI before joining Anthropic this summer. His exit thread passed 100 million views within a day, which tells you how starved the public is for confirmation from the inside.

I went into the replies expecting the usual corporate non-denial. Instead Evan Hubinger, Anthropic’s alignment science lead, replied directly: “Jacob is correct here, we really do earnestly believe AI could kill all humans.” He put his personal odds above 10% within the next decade, then added that the field doesn’t yet have a plan to solve alignment for superintelligence and isn’t clearly on track to find one. The safety lead of a frontier lab said, out loud, that they don’t know how to make the thing safe. And this is the lab that built its entire reputation on being the careful one.

Read Also:  Anthropic's Alignment Lead Admitted There's No Plan for Superintelligence

He wasn’t alone. Samuel Marks, who leads scalable oversight at Anthropic, confirmed that developers believe extinction-level outcomes could arrive “in the next few years.” Paul Christiano, who co-authored foundational safety work with Dario Amodei, said most people could die and announced he was joining OpenAI’s nonprofit safety team. Geoffrey Hinton called 10% a not unreasonable estimate. Others go higher still, with one researcher citing roughly 50% odds tied to decisions made this decade. That’s not alarmism. That’s a consensus forming in real time.

The AI Extinction Warning Isn't Hype, It's a Confession

Here’s the detail almost everyone missed. Roughly 92% of the people who saw Coxon’s thread only saw the headline. I dug into the viewership breakdown myself, and the section explaining the race dynamics, the part that makes the warning coherent instead of hysterical, was read by a rounding error. The internet consumed “researcher says AI could kill everyone” as content and scrolled on. Half my feed that same week was arguing about an Ocarina of Time remake. We metabolize everything the same way now, including our own obituary.

If they believe it, why keep building?

Every thread I read kept circling that question, and the honest answer is uglier than a conspiracy. It’s game theory. Each lab believes that if it slows down, a rival, usually framed as China, won’t, so the individually rational move adds up to a collectively insane outcome. Coxon’s own framing was that neither company acts responsibly because neither can afford to go second.

The internal language leaks out in fragments. Colleagues describe the next couple of years as “crunch time for humanity,” even “endgame.” Anthropic’s own August risk report only modestly raised its misalignment concerns for current models, which tells you the fear isn’t today’s chatbots. It’s recursive self-improvement, systems improving themselves faster than anyone can check the work. One comment I came across put the technical problem more plainly than most papers manage: we struggle to align these models when we build them, so how does that go when they build themselves?

Read Also:  TikTok's $400 Million Settlement Isn't a Fine, It's an Exit Fee

And here’s the fact that should retire the “he’s doing it for clout” take. Coxon left before his equity vested, walking away from real money to say any of this. You can argue he’s wrong about the odds. It’s much harder to argue he’s performing. Days after he quit, Anthropic published a report on bioweapon misuse threats, and other safety-motivated researchers have walked out of Anthropic and Google DeepMind under similar clouds this year.

Publicly, not much changes unless pressure converts into policy. More than 1,300 researchers signed an open letter in August warning that runaway AI could outpace human control, and the Pentagon’s posture so far reads as competitive anxiety rather than alarm. Trust in AI agents was already fragile, we’ve covered the trust problem before, and a week like this hardens it considerably.

The sentence I can’t shake isn’t “AI could kill all humans.” It’s “we do not yet have a plan.” You can live with a warning. A warning followed by business as usual is a different thing, because it turns every shipped model into an informed decision. My honest reading of this week is that the people with the best information on earth put extinction odds at one in ten and clocked in the next morning anyway. At some point the rest of us have to stop asking whether they believe it and start asking why a handful of private companies get to make that call for everyone else.

With ten years in the Industry, I write to provide our readers with the best material and great experience.

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *