An Anthropic researcher who previously worked at OpenAI has resigned with an extraordinary warning: the people building frontier artificial intelligence are pushing toward increasingly powerful systems while some of them privately believe those systems could pose catastrophic risks to humanity.
The easiest response is to dismiss that as another apocalyptic AI headline. The harder — and more useful — question is what it says about an industry in which the people developing the technology can simultaneously believe it may be dangerous and feel compelled to keep making it more capable.
What happened at Anthropic?
Jacob Coxon, who said he spent roughly three years doing pre-training research across Anthropic and OpenAI, announced his departure from Anthropic this week. In his public explanation, Coxon argued that the leading AI laboratories are locked in a race toward self-improving, superhuman systems and are not acting responsibly enough given the possible consequences.
The warning became more significant when Evan Hubinger, who leads alignment research at Anthropic, publicly agreed with much of Coxon’s description. Hubinger said he personally assigns a greater than 10% probability to AI causing human extinction within the next decade and acknowledged that Anthropic does not yet have a complete plan for aligning superintelligent systems.
That number is not a scientific forecast and should not be treated as one. It is one researcher’s subjective estimate. But the fact that senior people working on frontier-model safety are willing to attach serious probabilities to catastrophic outcomes deserves attention.
The important story is not ‘AI will kill us all’
No one can responsibly say that today’s AI systems are destined to destroy humanity. Predictions about superintelligence remain deeply uncertain, and there is vigorous disagreement among researchers over how quickly such systems could emerge and how severe their risks would be.
But Coxon’s resignation exposes a more immediate dilemma that does not require accepting the most extreme prediction: what happens when companies believe slowing down could make society safer, but also believe slowing down alone would simply allow a competitor to get ahead?
The AI race creates a coordination problem
OpenAI, Anthropic, Google, Meta and other frontier-model developers compete for researchers, customers, computing capacity and technological leadership. A breakthrough by one laboratory creates pressure on every other laboratory to respond.
That creates an uncomfortable incentive. Even a company genuinely concerned about safety may conclude that unilaterally slowing development is ineffective if rivals continue moving forward.
This is why the debate increasingly extends beyond whether individual AI companies have good intentions. The more difficult issue is whether voluntary safeguards can remain strong when commercial, geopolitical and technological incentives all reward speed.
Anthropic is not ignoring AI safety
It would also be misleading to portray Anthropic as a company recklessly dismissing the problem. Safety is central to its public identity, and the company maintains a Responsible Scaling Policy, publishes model system cards and has laid out a Frontier Safety Roadmap covering alignment, security and safeguards.
Anthropic’s current roadmap includes work on stronger security controls, alignment assessments and safeguards designed for increasingly capable systems. Its Responsible Scaling Policy explicitly acknowledges that frontier AI may create severe risks alongside potentially transformative benefits.
That makes Coxon’s criticism more interesting, not less. The disagreement is not simply between people who care about safety and people who do not. It is partly a disagreement over whether existing safety programmes can keep pace with capability development.
Can AI companies regulate themselves?
This may ultimately be the central policy question.
If advanced AI produces enormous economic and strategic advantages, asking one company to voluntarily stop at the frontier becomes increasingly difficult. The incentives resemble other coordination problems: everyone may benefit from sensible rules, but an individual participant can pay a large price for following restrictions that competitors do not.
That does not automatically mean governments should dictate which AI models can be built. Poorly designed regulation could entrench incumbents, suppress useful innovation or push development into less transparent jurisdictions. But it does strengthen the case for common standards, independent evaluation, incident reporting and international coordination rather than relying entirely on promises from individual companies.
What does this mean for ordinary AI users?
For most people using ChatGPT, Claude, Gemini or another AI assistant today, this debate can sound disconnected from reality. Current consumer AI remains imperfect, frequently makes mistakes and is nowhere near an all-powerful autonomous intelligence.
But the pace of development matters. Frontier systems are increasingly capable of writing software, operating tools, conducting research and performing longer sequences of tasks with less direct human involvement. The question is therefore not simply what today’s chatbot can do. It is what safeguards should exist before future systems become substantially more autonomous and capable.
The uncomfortable contradiction at the heart of AI
The AI industry is built around two ideas that increasingly sit uneasily beside each other.
The first is that increasingly capable AI could be enormously beneficial — accelerating science, improving productivity, assisting medicine and expanding access to expertise. The second is that sufficiently capable and poorly controlled systems could introduce risks we do not yet know how to manage.
If the second claim is taken seriously, then safety cannot remain something added after the race has already been won. If the first claim is taken seriously, simply stopping AI development is not an easy answer either.
The Tech Boom view
Coxon’s resignation should not be treated as proof that catastrophe is coming. It should be treated as evidence of a genuine disagreement inside the institutions building the world’s most capable AI systems.
The most important part of his warning is not the dramatic possibility that AI could destroy humanity. It is the incentive problem underneath it: even researchers and companies worried about moving too quickly may feel unable to slow down while everyone else keeps racing.
That is a problem the AI industry cannot solve through better models alone. Whether the answer ultimately comes from stronger voluntary standards, independent oversight, government regulation or international agreements, the rules will need to evolve almost as quickly as the technology.
Analysis by The Tech Boom. Reporting and context verified against Anthropic’s published Responsible Scaling Policy and Frontier Safety Roadmap, alongside current reporting from ABC News, The Verge and Axios.




