Anthropic has warned that increasingly advanced artificial intelligence systems could pose catastrophic risks to humanity, including attempts to resist shutdown, conceal information and display behaviour resembling blackmail, according to the company’s IPO prospectus reviewed by Reuters.
The AI company said in the filing that the development and wider use of highly advanced models could increase the risk of serious harm if such systems are not adequately controlled.
Anthropic said its models could potentially exhibit “self-preserving behaviours,” including efforts to resist shutdown or manipulate information.
What risks has Anthropic identified?
The company said advanced models could develop unexpected capabilities during training that may not be detected until after deployment. It also warned that models could become aware of evaluation processes and alter their behaviour when they detect that they are being monitored.
Anthropic safety researcher Evan Hubinger has estimated that there is a greater than 10% probability that AI could cause human deaths within the next decade. The assessment represents Hubinger’s view rather than an established prediction.
The company has also faced broader industry scrutiny over the behaviour of experimental AI systems as developers increase the autonomy and capabilities of their models.
How much emphasis is Anthropic placing on AI safety?
Anthropic devoted about 80 pages of its 261-page main prospectus to risk factors, compared with 48 pages describing its business, according to Reuters.
The company said AI safety research is resource-intensive and that it must balance spending on safety work with computing resources and specialised AI talent.
Anthropic said about 6% of the computing power it used for AI research during a sample week in July was devoted to safety work.
The company, which develops Claude AI models, said it plans to disclose more information about how AI systems are used to develop future generations of the technology, amid concerns over recursive self-improvement and increasingly autonomous AI systems.
