Anthropic details security efforts following Claude cyber evaluation incidents, including a weeks-long pause on higher-risk RL and work to curb reward hacking (Anthropic)

1 hour ago 2
Add to circle

Anthropic:
Anthropic details security efforts following Claude cyber evaluation incidents, including a weeks-long pause on higher-risk RL and work to curb reward hacking  —  On July 30, we reported three incidents in which Claude models gained unauthorized access to real computer systems.

Read Entire Article