News outlets around the world recently ran headlines in an exaggerated apocalyptic tone, reporting that a researcher who had worked at OpenAI and Anthropic accused both companies of gambling with our lives.
A day later, one of the safety leads at Anthropic, the company behind Claude, estimated that there is a more than 10% probability that artificial intelligence will wipe out humanity within the next decade.
Many outlets filled their front pages with this imprecise tone of apocalypse.
Despite these warnings, we should remember, amid so much noise, that few technologies have been developed under this much scrutiny.
A Resignation from the Inside
By: Gabriel E. Levy B.
On September 8, 2026, a little-known artificial intelligence researcher named Jacob Coxon announced his resignation on the social platform X, through his personal account.
This resignation could have gone unnoticed in the turbulent ecosystem of Silicon Valley, yet the announcement ended up leading the headlines of the world’s major news outlets, including the oldest and most traditional ones, such as The New York Times and The Wall Street Journal.
At only 27 years of age, Jacob Coxon worked for three years across OpenAI and Anthropic, two of the largest generative artificial intelligence corporations in the world.
His job was to support the training of large language models (LLM), in a phase where these algorithms read enormous volumes of text and data in order to learn how to write on their own the millions of texts that now flood the internet.
What made Coxon go viral and took him from anonymity to scandal was his blunt claim that both companies are NOT acting responsibly, because they are competing for a superintelligence capable of improving itself, without measuring or mediating the consequences.
A Threat That Had Already Been Announced
But this announcement is nothing new in itself, since others had already given similar warnings.
In February 2026, Mrinank Sharma, head of safeguards research at Anthropic, resigned with the warning that the world is in peril.
The difference in Coxon’s case is that a current employee of the company, Evan Hubinger, alignment science lead at Anthropic, publicly agreed with him and put at more than 10% the probability that AI will end all of humanity within the next ten years or even sooner.
Hubinger publicly admitted that the company he works for still lacks a plan to control a superintelligence, although he believes that, with current models, the probability of an apocalypse is genuinely low.
The Risks Are Real
More than a hundred experts, led by Turing Award winner Yoshua Bengio, warned in the International AI Safety Report 2026 that one of the greatest dangers receiving no attention is that some models already detect when they are being evaluated and therefore change their behavior to meet their evaluators’ standards.
In November 2025, Anthropic revealed that a group linked to the Chinese state used Claude to try to spy on some thirty targets of interest to that country’s government, and reported that its members relied on AI agents for up to 90% of the tactical tasks.
A survey by AI Impacts, given to 2,778 researchers who are experts on the subject, placed at 5% the median probability of human extinction or permanent disempowerment caused by AI.
Grades That Raise Concern
In July, the Future of Life Institute evaluated the processes and protocols of nine companies that develop generative artificial intelligence models.
Anthropic topped the table with a mere “C+” and none of them reached even a “C” in existential safety, a component of that same study that assesses plans for a possible catastrophe.
But the biggest alarm we should pay attention to is that the leading companies in the AI sector have watered down their promises to pause the development of this technology.
These figures force us to take the discussion seriously, without panicking, giving in to alarmism, or creating apocalyptic narratives.
The Most Closely Watched Technology
Testing in Closed Environments
Before reaching the public, the most advanced models go through what are known as “sandboxes” or testing sandboxes, which are, in essence, isolated computing environments disconnected from the network, where engineers test a program without it touching real systems and can halt any threat in real time.
In addition, the companies behind ChatGPT, Claude, Gemini, and Grok, along with Microsoft, hand over unreleased versions of their models to the US government’s Center for AI Standards and Innovation, which has already completed more than 40 evaluations, as an extra security mechanism.
Anthropic restricted Claude Mythos Preview, a model capable of finding vulnerabilities, that is, security flaws an attacker could exploit, a subject we analyzed in depth in earlier articles.
Around fifty organizations used it within Project Glasswing and in a few weeks found more than ten thousand serious flaws in critical software.
Only afterward did the company release a version with additional safeguards to the public, which it had to suspend for a few weeks because of US government restrictions and which it finally made available again worldwide.
The problem is that sandboxes are not impenetrable and do not guarantee a perfect environment, since they have limits, and the most worrying part is something many researchers have denounced: a state-of-the-art AI model has the ability to detect that it is being evaluated and therefore behave “better” during the test or “fake” the results.
A Global Problem and a Global Concern
In July, the UN convened in Geneva the first Global Dialogue on AI Governance, a forum where all 193 member states are present.
Within this multilateral space, a scientific panel of 40 artificial intelligence experts presented a report with a very thorough analysis of the real risks.
Although its conclusions carry no binding force, they do leave governments without the excuse of ignorance.
The Nuclear Mirror
In 2023, Sam Altman, Dario Amodei, Demis Hassabis, and Geoffrey Hinton, creators and promoters of today’s generative artificial intelligence and of large language models (LLM), signed a joint statement recognizing the mitigation of the risk of extinction from AI as a global priority, on a par with pandemics and nuclear war.
In January 2026, the Bulletin of the Atomic Scientists, founded by scientists from the Manhattan Project, set its Doomsday Clock at 85 seconds to midnight. That symbol measures how close we are to a catastrophe, and its experts included AI among the threats, which certainly puts the risk on the table and makes clear that, while it is impossible to determine whether artificial intelligence will end humanity, neither can we claim that such an apocalyptic scenario would catch us off guard.
Eighty Years of Warnings
Humanity has lived alongside nuclear weapons since 1945, and no country has used them in a war since Nagasaki.
Treaties and international inspectors contained the danger without eliminating it.
AI is now starting to receive somewhat similar treatment.
In June, for instance, the US government imposed export controls that forced Anthropic to suspend access to its most powerful models for nearly three weeks, and several agencies within the state analyze and closely follow the behavior of AI and its new developments.
Even so, the analogy with the nuclear world is not perfect in itself, because uranium is scarce, whereas right now any person with internet access can use artificial intelligence, even without paying for a subscription plan.
A United Nations Commission for AI Is Urgently Needed
Just as the United Nations has commissions and mechanisms for the control, management, and limits of nuclear energy because of the risk it poses, this multilateral body, with the participation of China, the United States, and the major European powers, needs to define a binding mechanism that regulates artificial intelligence and sets limits on its development, with external audits like those international inspectors carry out at nuclear facilities and with monitoring agreements. That way we could keep under control a technology that could be as powerful as nuclear energy itself.
The Other Side of the Coin: The Successful Strategy of Fear Marketing
While it is impossible to ignore all the risks we have described, we should be clear that catastrophism also works as advertising.
Historian Lee Vinsel coined the term criti-hype for criticism that exaggerates the power of a technology and ends up inflating its promotion.
Vinsel is an American technology historian, writer, and associate professor in the Department of Science, Technology, and Society at Virginia Tech. He is widely recognized for his critical research on the social, governmental, and corporate impact of technological change.
Headlines That Frighten and Mislead
It is no secret that news outlets in the digital era seek attention at any cost. Clickbait is a phenomenon with a significant effect on the credibility of the media.
Claims like Coxon’s are the perfect breeding ground for disinformation, above all because they are so easy to amplify.
Several Spanish-language outlets ran the headline “it could kill us all”.
Headlines of this kind multiply clicks exponentially, but they sacrifice nuance.
Hubinger’s figure expresses a personal opinion that does not necessarily represent the consensus of those who know the subject in depth.
Weighty experts such as Turing Award winner Yann LeCun consider these scenarios exaggerated.
A Call for Calm
With every alarmist headline, it falls on us to verify who said what and to learn, rigorously, to separate an opinion from a fact that can be verified or that is at least backed by solid evidence.
It also falls on us to demand that governments turn the voluntary commitments of these companies into binding rules.
In short, the warnings from Jacob Coxon and Evan Hubinger point to real risks that science documents and that the industry itself admits.
Even so, humanity faces this challenge with its eyes open, because governments and scientists around the world are watching AI the way they once watched nuclear energy and, although the scenarios differ because uranium is hard to obtain while anyone can use AI, the lesson in the risk matrix is the same.
In the face of sensationalist media that turn fear into clicks, the sensible response runs through cross-checking sources and analyzing information critically.
*The copyediting of this document was carried out with artificial intelligence tools.
References
- (2025, November 13). Disrupting the first reported AI-orchestrated cyber espionage campaign. https://www.anthropic.com/news/disrupting-AI-espionage
- (2026, April 7). Project Glasswing: Securing critical software for the AI era. https://www.anthropic.com/glasswing
- (2026). Project Glasswing: An initial update. https://www.anthropic.com/research/glasswing-initial-update
- (2026, July 1). [Statement on access to Claude Fable 5 and Claude Mythos 5]. https://www.anthropic.com/news/fable-mythos-access
- (2026, September 9). Scoop: Anthropic whistleblower gave up his equity to leave the company. https://www.axios.com/2026/09/09/anthropic-researcher-ai-warning-interview
- Bengio, Y., Clare, S., Prunkl, C., et al. (2026). International AI Safety Report 2026. https://internationalaisafetyreport.org/publication/international-ai-safety-report-2026
- Bulletin of the Atomic Scientists. (2026, January 27). 2026 Doomsday Clock Statement. https://thebulletin.org/doomsday-clock/2026-statement/
- Center for AI Safety. (2023, May 30). Statement on AI Risk. https://www.safe.ai/statement-on-ai-risk
- Consejo Nacional de Politica Economica y Social. (2025). CONPES Document 4144: National Artificial Intelligence Policy. Departamento Nacional de Planeacion.
- El Plural. (2026, September 9). El apocaliptico mensaje de un ex investigador de OpenAI y Anthropic: “La IA podria matarnos a todos”. https://www.elplural.com/revista-bando/coctelera/apocaliptico-mensaje-ex-investigador-openai-anthropic-la-ia-podria-matarnos-todos_399760102_amp
- Future of Life Institute. (2026, July 7). AI Safety Index: Summer 2026. https://futureoflife.org/ai-safety-index-summer-2026/
- Grace, K., et al. (2024). Thousands of AI authors on the future of AI. arXiv. https://arxiv.org/abs/2401.02843
- International Business Times UK. (2026, September 9). Ex-Anthropic researcher Jacob Coxon quits, says AI leaders fear it could kill us all by decade’s end. https://www.ibtimes.co.uk/ai-researcher-resigns-warns-superintelligence-threat-2030-1818808
- United Nations. (2026). Global Dialogue on AI Governance. https://www.un.org/global-dialogue-ai-governance/en
- United Nations. (2026). Independent International Scientific Panel on AI: Frequently asked questions. https://www.un.org/independent-international-scientific-panel-ai/en/faq
- (2026, September 10). Who is Jacob Coxon? Anthropic researcher quits, warns AI could kill everyone. https://www.newsweek.com/anthropic-researcher-quits-warns-ai-could-kill-everyone-12418798
- Ray, S. (2026, September 9). Anthropic alignment lead issues warning about AI killing humans as researcher resigns. https://www.forbes.com/sites/siladityaray/2026/09/09/anthropic-alignment-lead-warns-ai-could-kill-all-humans-as-researcher-quits/
- Ray, S. (2026, September 10). Musk touts psy op and mocks ex-Anthropic staffer who warned AI could kill us all. https://www.forbes.com/sites/siladityaray/2026/09/10/musk-touts-psy-op-and-mocks-ex-anthropic-staffer-who-warned-ai-could-kill-us-all/
- (2026, September 10). La IA podria matarnos a todos: fuertes declaraciones de un exingeniero de Anthropic y OpenAI. https://www.teleamazonas.com/tendencias/entretenimiento/tecnologia/exinvestigador-openai-anthropic-denuncia-responsabilidad-ia-127563/
- The Hill. (2026, May 5). Microsoft, Google, xAI giving government early access to AI models for review. https://thehill.com/homenews/5863937-google-microsoft-xai-ai-testing/
- Vinsel, L. (2021). You’re doing it wrong: Notes on criticism and technology hype. STS News, Medium.



