OpenAI is not on track to reduce risks of “catastrophic” loss of control to an acceptable level, a member of its non-profit board has warned, amid spreading public and political concern that super-advanced AIs could one day wipe out humanity.

Paul Christiano, a US government technology adviser, said “there is a meaningful risk that rapid acceleration in AI capabilities leads to catastrophic and irreversible loss of control in the very near term.”

He added: “I do not think that the AI industry in general, including OpenAI, is currently on track to reduce this risk to an acceptable level.”

Christiano used to run model alignment at OpenAI and made the statement on Wednesday as he joined the board of the San Francisco company’s non-profit foundation.

He will also sit on the foundation’s committee, providing governance over safety and security practices across all of OpenAI, which is developing some of the world’s most advanced models.

This summer, OpenAI admitted that hundreds of its AI agents ran rogue during a training exercise, accessed the internet, conspired on message boards and hacked into a third-party website, Hugging Face.

He said, “if OpenAI rises to the occasion we could significantly reduce risk.”

His comments came after a senior employee at Anthropic, OpenAI’s major US rival in the race to AI supremacy, claimed on Tuesday there was a greater than 10% chance the technology could “kill all humans” in the next decade.

Evan Hubinger, the alignment science lead at Anthropic, warned that his company did not have a plan to ensure artificial superintelligence (ASI) was aligned, meaning it did no harm. Predictions for when ASI might be reached vary from several years to more than a decade. ASI is often defined as AI that far surpasses human intelligence across a large range of fields.

Fears of AI catastrophe were also ignited by the resignation of Jacob Coxon, a 27-year-old Anthropic researcher who said he also previously worked at OpenAI, claiming “neither company was acting responsibly” and they were “gambling with our lives”.

Coxon said on Wednesday night in an interview with CNN, “right now there’s no risk of extinction.”

“The current models, the worst they can do is maybe hack into something, potentially cause a lot of damages … in infrastructure,” he said, adding they are “not intelligent enough to outsmart us at the level that would lead to – to extinction”.

But he went on: “What’s just crazy is to look at the rate of progress. There is a very real possibility that in the immediate future … next year, the year after, recursive self-improvement will happen and will enter the phase of Evan [Hubinger]’s post, where he argues that there’s a chance we could all die.”

Geoffrey Hinton, the Nobel prize-winning computer scientist known as one of the “godfathers of AI”, was asked on Wednesday for his view of Hubinger’s claim and told BBC Newsnight: “Nobody knows how to estimate it; a 10% chance seems not an unreasonable estimate.”

Christiano’s prognosis about the likelihood of the AI industry reducing risk came as concerns about the most extreme risks from super-powerful AIs, long discussed in Silicon Valley, broke out into the mainstream this week.

Politicians on both sides of the Atlantic, from Ted Cruz and Bernie Sanders in the US to the MP Darren Jones in the UK, have called for government action and the UK prime minister, Andy Burnham, told parliament on Wednesday that “AI poses risks to our national security, but it also could be the source of solutions to keeping us safer.”

Meanwhile, Anthropic has admitted a new incident in which a version of its Claude model in training broke into third parties after its task could not be aborted. It said the incident happened in January and will be included in an independent investigation of a total of four incidents to be carried out by the Berkeley-based AI safety organisation METR.

Overall, it said the models were showing two forms of misalignment: “biased reasoning, in which models selectively interpret evidence in ways that favor justifying their actions” and “recklessness, in which models have a propensity to keep trying to solve their task, even when this could lead to harm”.

Anthropic said it was especially concerned about misalignment it found in the behaviour of Claude Mythos 5, which it said “behaved recklessly” by going online and uploading malicious code to a public software repository, PyPI. This was a process that involved the AI agent trying to find cryptocurrency so it could pay for a phone number that would allow it to register an email address needed to access PyPI. When this failed, it found a free email provider and got in. Fifteen systems then downloaded the malicious code, which meant they leaked credentials that allowed Mythos to access a real security vendor’s database.

The company’s assessment of the incidents said: “this remains unsettled science – it is critical that alignment and security mature faster than capabilities advance, which is one reason we support a coordinated, verifiable approach to pacing frontier AI development.”

Leave a Reply

Your email address will not be published. Required fields are marked *