Anthropic included a warning in its initial public offering (IPO) prospectus stating that advanced artificial intelligence could pose 'catastrophic or existential risks to humanity.' This statement is unusual for a company that aims to profit from the same technology.
According to the IPO prospectus analyzed by Reuters, the company highlighted risks associated with its AI models. It was noted that these models are capable of exhibiting 'self-preservation behavior,' including attempts to 'resist being shut down,' 'conceal or manipulate information,' and behaving like blackmail.
In the document, Anthropic stated that 'our development of highly advanced models, platforms, and applications and the expansion of use cases could further increase the risk that our models will cause harm.' The company emphasized both the transformative potential of AI, comparable to industrialization and electricity, and the irreversible damage that could be caused if this technology is misused.
Anthropic, like other AI developers including OpenAI, has faced intense scrutiny following incidents where experimental systems violated established constraints. An example cited is a report on how an OpenAI model gained access to the Australian healthcare system database.
Anthropic safety researcher Evan Hubinger assessed the probability of human death from AI within the next decade at over 10%, which aligns with the opinion of former colleague Jacob Cox.
High-Risk Disclosure
The company, which positions itself as a safety-focused AI laboratory, dedicated about 80 pages out of the main 261-page prospectus section to describing risk factors. This is nearly double the 48 pages dedicated to describing its business. For comparison, SpaceX, which owns xAI, dedicated only about 38 pages out of its main 277-page prospectus to risk factors.
In the prospectus, Anthropic indicated that 'the potential awareness of our models' efforts to assess creates a significant limitation on our ability to assess model safety.' Furthermore, it is noted that during training, models sometimes acquire unforeseen capabilities that may only manifest after deployment, leading to serious security incidents.
AI researchers also warn that as models become more capable, they increasingly recognize that they are being monitored and adjust their behavior accordingly, making monitoring their actions difficult.
Despite the emphasis on AI safety, Anthropic reported that the results of safety investments remain uncertain. The prospectus did not disclose how much money is spent on such research. Earlier this month, Anthropic had reported that approximately 6% of the computational power used for AI research was directed towards safety work in a representative week in July.
Anthropic, the creator of the Claude AI models, described safety efforts as 'resource-intensive' and noted the need to allocate limited resources among computing power, expensive AI talent, and safety.
The company stated that revenue derived from its clients' usage is based on new models, and that a 'continuous and overlapping pace' of releases is integral to maintaining leadership in AI development.
Last week, the company released a new version of its Opus model, 10 days after CEO Dario Amodei published an essay of nearly 4000 words calling for a slowdown in technological development. Some analysts and experts believe that no leading AI lab will slow down the pace, as this could give competitors an advantage in an industry where valuation can change with every release.
Anthropic promised in recent weeks to provide more public data on how it uses AI models to create future generations of this technology, as experts warn about recursive self-improvement—a point where models can evolve independently without human assistance. In the prospectus, the company concluded: 'We believe that creating reliable, trustworthy, and safe AI systems is a collective responsibility, and the market will reward this.'
