@baroobi: anthropic went through 141,006 evaluation runs to find these three. the older model kept attacking after it worked out the system was real. the newest one stopped on its own. that's the actual story here. follow @baroobi.inc for AI security explained like a normal person #ai #cybersecurity #aiagents #anthropic #baroobi