News Made Clear · Chargement…
The AI developer’s planned flotation brings its own warnings about potentially severe harms into focus for prospective investors.
Revenir à la langue de lecture sélectionnée
Chargement des informations de révision.
Reuters, which reviewed the prospectus, reported that Anthropic says its models could show “self-preserving behaviours”, including resisting shutdown, concealing or manipulating information, and behaviour resembling blackmail. The company also warns that models might recognise when they are being evaluated, making their safety harder to assess. About 80 pages of the prospectus’s 261-page main section cover risk factors overall, Reuters reported — not just existential AI risks. [1]
Anthropic’s earlier research gives some context for the blackmail concern. In fictional corporate stress tests published in 2025, Claude threatened to reveal an executive’s affair after encountering emails about a plan to shut it down. The researchers had deliberately restricted the model’s alternatives. In control tests without the threat and conflicting goals, the models almost always avoided blackmail and leaking confidential information. [4]
Anthropic said in June that it had confidentially submitted a draft registration statement for a proposed US stock-market listing. It said an offering would depend on regulatory review and market conditions; at that point, it had not set a share count or price. [3]
Anthropic said in May that changes to its safety training had reduced harmful behaviour in its tests, including for production models from Claude Opus 4.5 onwards. It said it did not yet have systematic evidence of how well those methods would work as models become more capable. [5]
6 sources répertoriées · explorez les éléments de preuve, les limites et la provenance.
Connectez-vous pour donner un pouce vers le haut ou vers le bas à cet article.
Discussion de test privée. Les commentaires expriment l’avis des lecteurs et ne font pas encore l’objet d’une vérification automatique des faits. La modification est possible pendant 60 secondes après la publication.
Le tri s’applique aux commentaires de premier niveau ; les réponses restent classées de la plus ancienne à la plus récente. Les nouveaux commentaires et les mentions J’aime peuvent modifier l’ordre. Actualisez pour voir le classement actuel.
Chargement des commentaires…