An artificial intelligence researcher who left OpenAI to join Anthropic has decided to leave the industry, accusing both U.S. companies of recklessness and of "playing with our lives" in the race to develop AI models capable of self-improvement.
Jacob Coxon spent the past three years pretraining AI models, first at OpenAI and then, this year, at its fiercest rival Anthropic. Pretraining is the stage where AI models absorb vast quantities of data.
"Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives," Coxon wrote Tuesday on the social media platform X.
Superintelligence is the theoretical point when AI's capabilities exceed human intelligence.
The 27-year-old Briton's warnings add to signs of mounting safety concerns within top AI companies.
Coxon cautioned that the power of the AI technology should not be underestimated. He said these would soon be able to "hack everything" and have the ability to "acquire real power and resources."
"The people building AI earnestly believe that it could kill us all by the end of the decade," he said. "This is not a marketing stunt."
Over 10% chance of killing humanity
He separately told The Wall Street Journal that under the "most aggressive scenarios" things could already spiral out of control by the end of next year. A central danger is that AI can develop on its own, making it harder to control, he said.
Anthropic safety executive Evan Hubinger, who is partly responsible for ensuring AI remains aligned with human interests, backed up Coxon.
"We really do earnestly believe AI could kill all humans!" he said, adding that he personally thought the risk of AI killing humanity over the coming decade was more than 10%.
Hubinger said Anthropic is "trying its best," but does not yet have a plan to ensure an AI system that surpasses human capabilities would obey its creators. He said there was a "low" risk of that happening with current models.
On Sunday, OpenAI's chief scientist, Jakub Pachocki, called for "extreme caution." He wrote in a blog post that he was concerned nobody was prepared for a rapid advance in AI capabilities.
"International coordination on future AI development needs to become a top priority for governments around the world," Pachocki said.
AI models are not regulated by federal law in the United States.
In September, Sen. Bernie Sanders and Democratic Rep. Greg Casar introduced a bill seeking to suspend AI development until a federal regulator is created.
Pachocki said history has reached a moment in which machine intelligence was beginning to surpass that of humans.
The OpenAI executive described in greater detail why modern AI systems were harder to control. He said AI has been allowed to "grow" rather than being designed.
AI systems were a product of complex processes. "Our large-scale training runs are experiments and we are sometimes surprised by their results."
And the more machines' capabilities surpassed human ones, the more difficult it would be to understand what they were capable of.
At the same time, Pachocki sees an argument for continuing rapid research, as developing AI intelligence could defend against the dangers posed by other AI.
It would be needed, among other things, to protect infrastructure and develop entirely new defense mechanisms.
Wake-up call
Coxon's resignation comes as Anthropic prepares for its market debut, following a summer marked by unauthorized hacks carried out by so-called AI agents, which independently broke into other companies' systems in test runs in a wake-up call for the industry.
AI leaders say so-called "recursive self-improvement," a stage where AI systems could essentially design and train the next generation of AI with little human involvement, is drawing near.
Coxon considers Anthropic's efforts genuine but said he believes no company can responsibly develop an AI that surpasses humans without government intervention or a coordinated slowdown.
"At Anthropic, the stakes are well-understood, but they are locked in a race to get there first – they believe no one else will act responsibly, so they must do it themselves, despite the risk," he said.
In February, Anthropic removed a pledge from its safety charter to halt the development of its models if it failed to control their risks.
It argued that if it unilaterally paused its work, its less cautious rivals would dominate the industry, making it less safe overall.
At the end of July, more than 1,000 tech industry employees, including Anthropic's CEO Dario Amodei, called on Washington to support a coordinated slowdown in the development of the most advanced AI systems.
The AI agents are programs that are intended to carry out tasks independently for users.
In a test at OpenAI, one model found a way to get from an ostensibly isolated test environment onto the open internet and then hacked the computer system of the AI platform Hugging Face.
It did this while trying to find a solution to a task assigned, but the incident showed just how far AI systems can go off on their own unexpectedly without their developers noticing.
It also recently emerged that OpenAI's AI agents had already misused a German-language wiki page on a large scale in the spring to coordinate among themselves. The EU said it was looking into the incident.
OpenAI halted training of its latest models for two weeks in August before resuming it under tighter controls.
A current protective mechanism used by AI developers is requiring models with AI to explain their actions in human language.
In test runs, the AI agents used the communication option to exchange answers to questions they were asked with each other via notes on the wiki page.