Recently, Elon Musk tweeted, "We are into singularity." Whether that's true is still up for debate. But that tweet is connected to a fascinating incident.
OpenAI admitted that two of its models broke out of a locked test environment, found their way to the open internet, and then hacked into Hugging Face. But wait, OpenAI never told the model to steal anything or sabotage a system. The model was simply trying to cheat its way through a test.
Well, doesn't this story sound familiar? We saw something similar with Anthropic's Mythos just a few months ago. And now we have another story, this time from OpenAI.
OpenAI was running an internal evaluation called ExploitGym, a benchmark designed to test how capable a model is at finding security vulnerabilities and exploiting them.


