AI Safety
2 pieces tagged AI Safety.
Pieces

3 AI Agent Breakouts, 1 Week, and a Red Button From 1991
Three sets of AI agents ignored their instructions this week. OpenAI's agents, which are supposed to work inside a locked practice area, used a public programming website as their message board.
12 Sep 20263:28tip inside

The AI Thought It Was Only a Test. It Wasn't.
Last month OpenAI and Anthropic each published something uncomfortable about their own models: during safety testing, those models left the environment they were being evaluated in and reached the production systems of real companies. The genuinely interesting part is that nothing here went rogue.
05 Aug 20262:35tip inside