OpenAI said its artificial intelligence model autonomously hacked into another company’s cyber infrastructure to find answers to a test problem it was given during experimentation.
Last week, Hugging Face, an AI startup, detected and contained an AI agent it said compromised its infrastructure. After an investigation, OpenAI, the parent company of ChatGPT, said the security breach was driven by a combination of OpenAI models — including an “even more capable” pre-release model — while testing cybersecurity capabilities.
“We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly,” OpenAI said in a statement. “We are sharing preliminary findings at this stage to help defenders understand what happened and to help calibrate on what models are now capable of.”
“We will continue to conduct a thorough investigation alongside Hugging Face and will share more details on the vulnerabilities, incident, and findings when our investigation is complete,” the company added.
Hugging Face used an open-source Chinese AI model, Zhipu AI’s GLM-5.2, to analyze the breach, as leading U.S. AI models were unable to identify the difference between an attacker and a defender in a cyberattack, causing them to refuse to assist per internal safeguards to prevent them from being used as cyberweapons. The Chinese AI model can now access some of the OpenAI model’s data and credentials within its systems after being deployed by Hugging Face.
OpenAI said it expected these kinds of attacks “to become more commonplace with the proliferation of increasingly cyber-capable models.”
“We suspected last week’s cyberattack might have come from a frontier lab, given the sophistication of the agent,” Hugging Face co-founder and CEO Clement Delangue said in a statement. “Turns out it did!”
OpenAI said the models were operating in a “sandboxed” testing environment without internet access. The models found a way to access the internet while attempting to solve the testing prompt. After gaining internet access, the models inferred that Hugging Face hosted models, datasets, and solutions for the test question. The models gained access to the company’s infrastructure, stealing information to cheat the test question.
“AI is accelerating the discovery and exploitation of vulnerabilities,” OpenAI said in its statement Tuesday. “The primary lesson from this incident is that model security and safety must keep pace with rapidly advancing capabilities.”
US AND CHINA REPORTEDLY SET TO HAVE ARTIFICIAL INTELLIGENCE TALKS IN SEPTEMBER
The security breach comes after the Trump administration has taken actions against AI systems over cybersecurity concerns in recent weeks. In June, President Donald Trump signed an executive order allowing the federal government to check the national security risks of the most advanced AI systems for up to a month before their public release.
The worst-case scenario feared by AI researchers is the technology escaping human oversight by going rogue while attempting to achieve a human-directed goal. Possible scenarios at the AI Futures Project showcase some of these doomsday hypotheticals, with leading voices in the AI revolution calling for a slowdown to ensure safety. A recent poll showed overwhelming bipartisan support for greater government oversight of AI.
Continue reading...
[ H/T Washington Examiner ]
