Anthropic Researcher Warns AI Could Pose Existential Risk

Date:

A senior artificial intelligence safety researcher at Anthropic has warned that rapidly advancing AI systems could eventually pose an existential threat to humanity, putting his own estimate of the risk at more than 10% over the next decade.

Evan Hubinger, who works on AI alignment at Anthropic, said currently available models presented a relatively low risk but expressed concern about what could happen as systems become more capable and potentially able to improve aspects of their own development.

In a post on X, Hubinger said he believed the possibility of AI causing human extinction was serious enough to warrant attention, while acknowledging the considerable uncertainty surrounding predictions about future systems.

He did not describe a specific scenario in which AI could cause humanity’s extinction. His comments instead centre on a broader concern within AI safety research: whether increasingly autonomous and capable systems can reliably remain under human control.

The warning comes as researchers, technology executives and governments intensify debate over how quickly advanced AI should be developed and what safeguards should be required.

Concerns over superintelligence

Hubinger’s comments followed the departure of AI researcher Jacob Coxon from Anthropic. Coxon, who previously worked at OpenAI, publicly criticised both companies’ approaches to the development of increasingly powerful AI.

He argued that future systems could acquire capabilities far beyond those available today, including advanced cybersecurity skills and the ability to accelerate scientific and technological research.

Hubinger responded by saying that concerns about catastrophic AI risks within Anthropic were genuine. He also acknowledged that researchers had not yet solved what is known as the alignment problem for a future superintelligence.

AI alignment refers to efforts to ensure artificial intelligence behaves according to human intentions and values, particularly as systems become more autonomous and capable.

This becomes increasingly difficult if an AI system develops abilities that exceed human expertise in important areas. Researchers want to understand how such systems can remain predictable and controllable even when operating on tasks that humans may struggle to supervise directly.

Anthropic has previously acknowledged uncertainty around some of these risks. In an August 2026 safety assessment, the company described the likelihood of certain forms of serious model misalignment as low, while also saying it had become less confident about some assessments as AI capabilities continued to develop.

The report also examined whether highly capable AI could eventually perform automated research and development at a level capable of accelerating further technological progress. Anthropic said it was seeing early indications of possible acceleration, although that does not mean today’s systems have reached such a level.

Debate moves towards regulation

The increasingly strong warnings coming from within AI laboratories are adding pressure on governments to determine how advanced systems should be governed.

Darren Jones, a former senior UK government official, has called for international cooperation around the development of superintelligent AI, arguing that countries may need a multinational framework similar to agreements used to manage other technologies carrying significant global risks.

Computer scientist Dame Wendy Hall has also questioned the extraordinary nature of some of the public warnings coming from researchers employed by companies developing the technology. The debate creates an unusual situation in which some of the organisations racing to build the world’s most capable AI systems are simultaneously warning that future versions of those systems could become dangerous.

Questions have also emerged over how independently the most advanced models can be evaluated. The Financial Times reported that Anthropic did not provide its latest model to the UK’s AI Security Institute for assessment. The UK government has said it continues to work with AI companies on model safety.

The issue goes beyond Anthropic.

Researchers and executives across the industry have increasingly discussed the possibility that future AI systems could become capable of performing complex tasks with limited human supervision. OpenAI chief scientist Jakub Pachocki has recently called for extreme caution as capabilities advance, while other prominent figures have advocated stronger mechanisms for controlling the pace at which frontier AI is developed.

More than 1,000 employees working across the AI industry have also backed an open letter calling for international efforts to develop technical and governance mechanisms capable of slowing frontier AI development if necessary.

For policymakers, however, determining the appropriate response remains difficult. Predictions about artificial general intelligence, superintelligence and extinction-level risks are highly uncertain, and researchers disagree substantially about both the likelihood and timing of such outcomes.

There is also an important distinction between today’s AI systems and hypothetical future systems. Current models can make significant mistakes, generate false information and behave unpredictably, but they do not possess the capabilities assumed in many scenarios involving superintelligence.

The concern raised by Hubinger and other safety researchers is about where the technology could be heading rather than what existing AI can currently do.

That distinction matters. A prediction that there is a greater than 10% chance of an existential catastrophe is not evidence that such an outcome will occur, nor is there scientific consensus around a precise probability.

What makes the debate significant is the source of the warnings. Some are coming from researchers working directly on the most advanced AI systems and studying how those systems behave.

As competition between AI companies accelerates, the challenge for governments and the industry will be deciding how seriously to treat uncertain but potentially enormous risks without losing sight of the technology’s considerable potential benefits.

The central question is increasingly becoming one of control: as artificial intelligence becomes more capable and autonomous, can researchers ensure that humans remain firmly responsible for what it does?

LEAVE A REPLY

Please enter your comment!
Please enter your name here

Related articles

How AI Is Transforming Agriculture in Rwanda

Artificial intelligence is changing many industries around the world, and agriculture is no exception. In Rwanda, AI and...

CPU vs GPU vs NPU: What’s the Difference?

As artificial intelligence becomes more common in smartphones, computers and data centres, you may increasingly hear three terms:...

Inside an AI Data Centre: How AI Computing Works

Artificial intelligence may feel like something that simply lives on your phone or computer. But behind many AI...

How Artificial Intelligence Works: A Simple Guide

Artificial intelligence (AI) is becoming part of everyday life. From chatbots and smartphones to banking, healthcare and agriculture,...