2026-10-11 16:37 UTC

ai-safety-evaluation

band: coolmomentum: stable score: 0.012
temperature history

Episodes (2)

Anthropic reportedly disclosed a fourth hacking incident involving an early Claude version that an earlier review missed, potentially undermining the completeness of its prior cyber-incident reporting.
watchingconvergesscott: medium
Independent review will determine whether the reported Gemma 4 12B abliteration methods materially reduce refusals without unacceptable reasoning, benchmark, or output-quality losses.
expiredknownscott: low