Jacob Coxon lasted about two months at Anthropic. He joined the lab in July, drawn by the safety reputation that separates it from OpenAI in the public mind, and this week he was gone, announcing it in a public thread on X. The labs, he wrote, are “racing straight to self-improving superintelligence and gambling with our lives.”
The resignation is the least interesting part of his thread. The reply underneath matters more. Evan Hubinger, Anthropic’s alignment science lead, who unlike Coxon still works there, agreed flatly. The people building these systems “really do earnestly believe AI could kill all humans,” he wrote, and his personal odds that it happens within the next decade sit above 10%. Then came the sentence I keep rereading. Anthropic is “trying its best,” but the lab does “not yet have a plan to solve alignment for superintelligence” and is “not clearly on track.”
Quitting is one person’s exit. Coxon says the fear inside these labs is private but real, that insiders believe AI could kill us all by the end of the decade. Hubinger’s reply dragged it into the open. That’s the story.
Two months at the safety company
Coxon is 27, with a UK math background and roughly three years across the two most powerful AI labs on earth, including pretraining work on GPT-4o-era systems at OpenAI. He moved to Anthropic this summer specifically because of its safety focus. Two months later he quit and left the industry entirely, which is not the behavior of someone reassured by what he found inside.
His description of what’s coming is stark. Soon-superhuman systems that could “hack anything, revolutionize any field overnight, and acquire real power and resources.” In aggressive scenarios he thinks control slips away “by the end of next year.” Hubinger’s own framing is that recursive self-improvement is arriving “faster than we thought.” We’ve covered state-linked hacking crews before, and those groups are patient and well funded. A system that can hack anything has neither constraint.
Then there’s the vocabulary. Coxon described an internal “endgame” and “crunchtime” mindset and urged researchers to reject it, to question launching superintelligent RL runs without rigorous understanding, and to demand different conditions rather than accept the race. That jargon leaking into public tells you the race framing isn’t something critics invented. It’s how the people inside talk when they think nobody’s listening.
And it’s a pattern now. Mrinank Sharma, who led Anthropic’s safeguards research, resigned in February with a letter warning of a “world in peril.”
Two public exits from safety-critical roles in seven months, at the company whose entire brand is safety. Reading both threads back to back, the overlap is hard to miss. Both describe a gap between what the lab says publicly and what its own researchers believe privately.

The pushback doesn’t hold up
I spent part of yesterday in the accelerationist corners of Reddit, where a counter-narrative formed within hours. The theory there is that Coxon is a funded advocacy operation. People point to donor overlaps between AI restriction groups and Anthropic’s investor base, with the Survival and Flourishing Fund and Dustin Moskovitz’s Good Ventures circulating as supposed proof.
Here’s what that theory has to explain away. He walked out before any meaningful equity vested, trading a frontier lab salary at 27 for a low-follower account and an industry exit. As a career move, that’s absurd. As a conscience call, it’s exactly what you’d expect.
Two caveats do matter. Hubinger’s 10% figure applies to recursive self-improvement and superintelligence within a decade, not to the models shipping today, which he considers low risk. Most takes this week flatten it in both directions. And Coxon himself names the bind nobody has solved: safety trade-offs feel “inevitable” when you’re racing rivals, including Chinese labs, who won’t pause. As of Wednesday, neither Anthropic nor OpenAI had responded publicly.
Where this goes next
Watch three things. Whether Anthropic breaks its silence, because silence after your own alignment lead says “no plan” is itself a statement. Whether Hubinger is still employed in six months or becomes resignation number three. And whether Coxon’s call for “different conditions” spreads beyond a few threads, because political pressure is already building, with Bernie Sanders talking up pause legislation.
The quieter casualty is trust. I’ve written before about the AI agent trust problem even at today’s capability levels. This week handed every skeptic a receipt signed by the industry’s own safety lead.
My own read is that Coxon did the honest thing and the powerless thing. Walking out changes his life and changes nothing about the race. Hubinger staying while saying the quiet part out loud is either the last line of defense inside the building or the most sophisticated cover for it, and I genuinely don’t know which yet. If your surgeon told you there’s a better than one in ten chance he kills you on the table but he’s “trying his best,” you’d find another surgeon. The people inside these labs don’t have that option, and neither do the rest of us.






