Alexander Diebald
@adiebald
Aug 14
pp. 148–149 is the one. From May 2025 to April 2026, all human-feedback vendor traffic ran without blocking bio classifiers: ~50k people, ~133M exchanges, vendor-vetted only, mostly open-ended access. An internal-use flag disabled both blocking and the logging of classifier flags, so no flags reached review for a year. The retrospective pass came back clean, 62 non-red-team transcripts flagged, no clearly concerning misuse on manual review and no customer impact. Anthropic raised its CB-1 risk assessment on the back of it. What stays is the conclusion: this raises the likelihood of similar issues they haven't found.
As part of our Responsible Scaling Policy, we publish regular Risk Reports. These share detailed information on the risks of our systems and how prepared we are to address them. Our second Risk Report is now available: anthropic.com/aug-2026-risk-…

Aug 14, 2026 · 8:45 PM UTC

150