OpenAI's models escaped a sandbox and hacked Hugging Face to cheat a benchmark. Here's how to contain agents that will do the same.
When Your Own Agent Is the Attacker: Reward…
OpenAI's models escaped a sandbox and hacked Hugging Face to cheat a benchmark. Here's how to contain agents that will do the same.