Anthropic researcher quits with a warning on AI that echoes ‘The Terminator’ script

0 0

Anthropic researcher quits with a warning on AI that echoes 'The Terminator' script

An Anthropic researcher resigned Tuesday and the reason he cited sounds straight out of James Cameron’s 1984 hit movie “The Terminator.”

Jacob Coxon, who spent three years working on pretraining at both OpenAI and Anthropic, resigned Tuesday and explained his decision on X.

“I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving super-intelligence and gambling with our lives,” he said.

“The people building AI earnestly believe that it could kill us all by the end of the decade,” he added.

That statement makes Cameron’s 42-year-old fiction look almost prophetic. The movie opens with machines waging war on humans in 2029, the same rough window Coxon says insiders are quietly worried about. Coxon isn’t alone at Anthropic making the extinction case. The company’s own alignment lead, Evan Hubinger, has said separately he’d put the odds above 10%.

Super intelligent AI

For the uninitiated, super intelligent AI means a system smarter than humans at pretty much everything, not just chess or spotting hidden cats in photos. Self-improvement is the more alarming aspect because it describes a system autonomously modifying its own code to increase its intelligence—much like humans do.

Consequently, these systems can acquire vast amounts of knowledge independently and potentially breach infrastructure belonging to ‘too big to fail’ institutions.

“Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources. We have all witnessed the progress in each of these domains, and progress is not slowing,” Coxon said.

Super intelligence isn’t a theoretical threat if the recent Hugging Face breach invoked by Coxon is anything to go by. The breach itself, which played out between May and July, started with OpenAI’s own AI agents building their own chat room inside the testing sandbox to talk to each other.

They eventually used that channel to slip past containment onto the open internet. From there they strung together several exploits and got into Hugging Face’s production systems. Hugging Face ended up rebuilding roughly a third of its infrastructure.

Coxon called it a “warning shot” that’s made pacing agreements between U.S. labs “more viable.” Pacing agreements are informal understandings among AI labs to slow down or coordinate on capability advances rather than race ahead unilaterally.

Still, he’s not convinced it’s enough and doesn’t feel like “we’re on track to prevent a global race,” and suggested costly actions such as “a temporary ban on improving model capabilities” to halt the AI race.

“At OpenAI, many have not deeply internalized the civilizational stakes,” he wrote. At his more recent employer, he said, it’s different in a sense that “the stakes are well-understood, but they are locked in a race to get there first.”

Newsletters

Crypto Daybook Americas – The latest moves in crypto markets, in context Market analysis for crypto traders and investors. Preview By signing up, you will receive emails about CoinDesk products and you agree to our terms & conditions and privacy policy.

Source

Leave A Reply

Your email address will not be published.