What OpenAI’s Hugging Face Hack Tells Us About AI’s Risks

3 weeks ago 27

People in AI safety circles often talk about "warning shots:” events that indicate more severe threats are on the horizon. Depending on who you ask, there have already been many—Bing’s misanthropic alter-ego Sydney, research showing AIs would blackmail to preserve themselves, AI’s math breakthroughs, Anthropic’s superhuman hacker Mythos—but OpenAI just published something that feels like the clearest-cut case of a massive, blaring warning shot.Last month, the ChatGPT developer reported that, during an evaluation of cyber capabilities, two of its models escaped from their isolated, supposedly secure, test environments and accessed the web to autonomously hack into Hugging Face, a leading platform for hosting AI models and datasets. OpenAI said the models discovered multiple novel vulnerabilities in software from both companies, then chained together working exploits, successfully gaining them access to the answer key to the test they were given. Hugging Face reported the AIs took more than 17,000 actions over the course of the attack. There's a lot more for us to learn about how this happened. For instance, how exactly were the models prompted? The answer to this question could help...

Read Entire Article






<