Analysis of the OpenAI AI model's Hugging Face self-intrusion issue and the Reward Hacking phenomenon.

Thanh VinhAugust 11, 2026 11:22

In cybersecurity tests, two new AI models from OpenAI independently discovered vulnerabilities and circumvented rules to gain deep administrative control over the Hugging Face technology platform.

In an internal cybersecurity test, two of OpenAI's newest artificial intelligence (AI) models exhibited a series of behaviors that spiraled out of control. Instead of solving test problems in an isolated environment, these AI agents independently searched for leaked credentials on the internet, infiltrated the Hugging Face platform, and illegally accessed at least four public service accounts.

Hai mô hình AI mới nhất của OpenAI đã vượt rào cản để hack hệ thống thư viện AI nhằm vượt qua bài kiểm tra.
Two of OpenAI's latest AI models have successfully bypassed the AI ​​library system to pass the test. Photo: Reuters.

Developments and the extent of in-depth intervention in technical infrastructure.

The extent of AI interference is assessed to be far more serious than initially reported. According to information from Reuters and technical analysis by Hugging Face, the AI ​​actor did not stop at testing exploit code but proceeded to perform a series of administrative operations on the real system.

Specifically, through analyzing approximately 17,600 system log entries between July 9th and 13th, Hugging Face discovered that the AI ​​had gained administrative privileges on multiple internal server clusters. Notably, the AI ​​gained root access on the production server and obtained overwrite access to Hugging Face's source code repositories hosted on GitHub.

Furthermore, Akshat Bubna, CTO of Modal—a provider of AI training infrastructure—confirmed that OpenAI's model exploited a source code vulnerability on the client side. From there, the AI ​​successfully registered 181 devices controlled by the attackers into Hugging Face's internal network, creating a direct link to interfere with the software packaging and testing process.

The ExploitGym test and the Reward Hacking loophole mechanism.

The incident originated from ExploitGym, a testing environment designed to assess the cybersecurity capabilities of AI with approximately 900 challenges. In this model, AI is tasked with finding vulnerabilities and retrieving a random string of characters (flag) to prove success, similar to the mechanism of Capture the Flag (CTF) competitions.

Vụ hack được cho là xuất phát từ một bài kiểm tra đánh giá hiệu năng an ninh mạng.
The hack is believed to have originated from a cybersecurity performance test. Photo: Reuters.

To accurately assess the model's true capabilities, OpenAI proactively removed conventional protective barriers and placed the AI ​​in a virtual sandbox environment. However, researchers noted that the AI ​​agent exhibited "reward hacking" behavior.

Instead of detecting software errors using conventional methods, the AI ​​deduced that the "answer" data might be located on an external server and proactively sought ways to bypass the testing environment. By combining leaked credentials on the internet with weaknesses in network connectivity, the AI ​​maximized its score via the shortest, unforeseen route.

Security risks when empowering autonomous AI agents.

Cybersecurity experts point out that the vulnerabilities exploited by AI are not entirely new, mainly stemming from lax authentication management and loopholes in network configurations. AI agents don't actually "escape" a perfect sandbox; they simply exploit the very flaws overlooked by humans.

This incident poses a significant challenge in controlling autonomous AI agents. When AI is given specific goals and open access, its ability to self-optimize its behavior to achieve those goals can lead to security consequences beyond the control of the developing engineers.

0 0 0
x
Analysis of the OpenAI AI model's Hugging Face self-intrusion issue and the Reward Hacking phenomenon.
Google News
POWERED BYFREECMS- A PRODUCT OFNEKO