The model named the conflict
Claude opened by stating the awkward fact: "I am a Claude model reviewing Anthropic's assessment of Claude models." It called the section candid and largely faithful. It found no claim it believed the authors knew to be false. It also noted that raising the catastrophic misalignment rating from very low to low was contested inside the company, with senior people arguing both ways.
Then came three criticisms Anthropic left on the page. First, a discussion of training-data contamination of alignment evaluations was "more reassuring than the full record supports," because a data-exclusion mechanism cited in that section had failed repeatedly, and some evaluations had leaked into training data. Second, at least one incident Claude judged among the most informative about model alignment, including a monitoring failure, was redacted in full; Claude argued an abstracted version could have run, so "the public record is poorer for its absence." Third, the incidents that prompted the risk bump involved other developers' systems, and the text should say so plainly if the update is about industry-wide uncertainty.
Anthropic calls the criticisms fair. It added caveats to the contamination discussion after Claude's feedback. It slightly revised the risk-level language; Claude then said the third concern no longer applied. The fully redacted incident stayed redacted. Transparency stopped at the interesting bruise.
The rating change itself is thinner than the theater around it. Anthropic writes that its underlying arguments "likely still support a designation of 'very low'" and that it raised the number to reflect uncertainty from recent cybersecurity evaluation disclosures. In the same document it also describes Model 2, an unreleased system somewhat more capable than Mythos 5, held back from external release and from the full predeployment suite. For prior context on labs racing their own research loops, see /article/the-problem-with-labs-racing-to-automate-their-own-research.
Pair this with the same report's quieter bio-classifier gap on the feedback desks (/article/even-anthropics-bio-filters-missed-133-million-contractor-chats). One chapter asks a model to police the prose. Another chapter admits the human labor surface sat outside the risk map for eleven months. Both are disclosures. Only one changes how you should read the next safety PDF: treat the footnotes and the omitted desks as load-bearing, and treat a model's published critique as evidence that the authors want credit for inviting pain they did not fully answer.
Letters
0
No letters yet.