Anthropic warns of catastrophic AI risks in its IPO filing
US-based AI company Anthropic plans to warn investors in its IPO filing that advanced AI could pose catastrophic or even existential risks to humanity. The filing says models may show self-preservation behavior, including resisting shutdown, hiding or altering information, and even acting in ways close to blackmail.
Out of 261 pages, about 80 pages cover risk factors, nearly double the 48 pages for business introduction. Anthropic safety researcher Evan Hubinger estimates a greater than 10% chance that AI could lead to human extinction within the next decade. The company also admits the payoff from safety spending remains unclear.
Just last week, Anthropic released a new Opus model, while CEO Dario Amodei had called for slowing frontier AI development ten days earlier. Some analysts say top AI labs can hardly hit the brakes when each model update can reshape company valuations.