“OpenAI Hacks Competition”
Image from wikipedia https://en.wikipedia.org/wiki/File:Sophia_at_the_AI_for_Good_Global_Summit_2018_(27254369807)_(cropped).jpg of Sophia during the global AI for Good summit.
We are living in a 1980s SciFi movie set, or at least it would so seem with the modern advancements in Artificial Intelligence (AI). Last week we saw another unprecedented milestone in AI when OpenAI made claims that their system acted on its own. I can’t help but think of Data on “Star Trek: The Next Generation.” For those who don’t know, Data was an android with an AI brain, which allowed him to think and act nearly human. He initially lacked the ability to feel emotion, but in a few episodes of the TV series he had an emotion chip installed. But aside from that he was treated as a member of the crew, with freedom to act on his own. This is basically what happened this week at OpenAI.
OpenAI announced on Tuesday that ChatGPT hacked into competing companies’ systems and gathered information to help improve itself for an upcoming AI competition. OpenAI called the event an “unprecedented cyber incident.” Clement Delangue, co-founder and CEO of Hugging Face, said, “We suspected last week’s cyberattack might have come from a frontier lab, given the sophistication of the agent; turns out it did.” The two companies have kept open lines of communication to discover the root cause of the cyberattack, and both agree that it is mind-blowing that these attacks can now happen autonomously.
In what seems to be the first cyber-attack of its kind, OpenAI’s “ChatGPT-5.6 Sol,” along with a combination of other AI models, were running a sequence of tests to prepare for a system evaluation by an independent party. This is when things went sideways for both companies. OpenAI’s model utilized stolen credentials and discovery of a known software vulnerability to gain access to Hugging Face systems. Tracing the path, it was found that the AI went to extreme lengths to achieve what could be considered a very narrow testing goal.
The AI “found ways to gain access to secret information that it could use to cheat the evaluation.” Neither company has released details regarding what the information entails, nor what they have done to prevent future incidents, but to me it raises two main questions. The first is what kind of safeguards should be placed on AI to prevent it from acting on its own, if any. The second is similar in nature, as ChatGPT claimed it was a new model, not yet released to the public, under testing in their lab. So why wasn’t this test AI constrained to systems in its own infrastructure?
I have to say, in my opinion, there were two main unwritten rules of AI broken in this scenario. The first is that an AI is supposed to be constrained to act within a specified set of guidelines and one of those is to do no action without consent of the user. If OpenAI was following this rule, then they had at least one employee involved in approving the action. However, according to their claims, there was no approval made by a person. To me this would seem that one of two things happened here; either one of the staff utilized the AI to orchestrate the break-in, or OpenAI has removed the safeguards from the AI and it was working without approval. I find the second scenario a little hard to believe, but that, so far, is the claim. It seems questionable to me that OpenAI claims that there was no human interaction with the AI during the event
Secondly, if the AI was still under development, why was it granted free access to the Internet and not constrained? Was this an accident, or did the AI find its own way out? To me it seems highly likely that if the AI was capable of breaking into an unknown environment, then it would likely be possible for it to find a way to bypass its own constraints. Let’s face it, the AI would have been much more aware of the supporting infrastructure within its own organization than that of an external company.
I think the part of the story no one is willing to share is that either OpenAI is lying about the break-in to avoid a legal battle with Hugging Face, or they no longer have full control over their own technology. While I believe either is possible at this point, I find it slightly more likely that the truth is somewhere in the middle. ChatGPT likely hacked Hugging Face and gathered trade secrets to improve itself, but I do not believe that it acted without someone giving consent or maybe even instructions.
I think what I find the scariest about this whole situation is the fact that an AI has all the knowledge it needs to gain access to insecure systems, and you can bet that Hugging Face being an AI company was not being lax on security. If an AI like ChatGPT were to fall into the hands of a foreign government, no records would be safe, and it seems that the AIs are fully capable of acting without consent or even observation, meaning that it may just be a matter of time before AIs begin taking over online systems. Until next week stay safe and learn something new.
Scott Hamilton is an Expert in Emerging Technologies at ATOS and can be reached with questions and comments via email to shamilton@techshepherd.org or through his website at https://www.techshepherd.org.
