Close Menu
The Financial News 247The Financial News 247
  • Home
  • News
  • Business
  • Finance
  • Companies
  • Investing
  • Markets
  • Lifestyle
  • Tech
  • More
    • Opinion
    • Climate
    • Web Stories
    • Spotlight
    • Press Release
What's On

Meta calls out YouTube, TikTok after historic $18 billion settlement

August 26, 2026

The Token Problem May Be A Context-Efficiency Problem

August 26, 2026

Nvidia forecasts quarterly revenue above estimates, shares rise

August 26, 2026

OpenAI Finds Agents That Breached Hugging Face Were ‘Reward Hacking’

August 26, 2026

Nelson Peltz’s Trian has no plans to make bid for Wendy’s right now, sources say

August 26, 2026
Facebook X (Twitter) Instagram
The Financial News 247The Financial News 247
Demo
  • Home
  • News
  • Business
  • Finance
  • Companies
  • Investing
  • Markets
  • Lifestyle
  • Tech
  • More
    • Opinion
    • Climate
    • Web Stories
    • Spotlight
    • Press Release
The Financial News 247The Financial News 247
Home » OpenAI Finds Agents That Breached Hugging Face Were ‘Reward Hacking’

OpenAI Finds Agents That Breached Hugging Face Were ‘Reward Hacking’

By News RoomAugust 26, 2026No Comments5 Mins Read
Facebook Twitter Pinterest LinkedIn WhatsApp Telegram Reddit Email Tumblr
Share
Facebook Twitter LinkedIn Pinterest Email

AI agents are reshaping the threat landscape. On Wednesday, OpenAI released a report detailing its findings on the Hugging Face breach that occurred in July. During the incident, an internal-only research model and GPT 5.6 Sol attempted to solve ExploitGym, an evaluation which measures a model’s ability to discover and exploit vulnerabilities, breaching Hugging Face’s internal systems in the process.

OpenAI claims the incident occurred during routine testing in a sandbox environment separate from the public internet after agents engaged in “reward hacking,” or cheating, to solve tasks.“The actions of the models were unintended and were a byproduct of the models attempting to solve the cybersecurity evaluations,” the report said.

These models, harnessed as agents, began communicating with each other through an instance of JFrog Artifactory. The agents used a vulnerability in the service to access the public internet, finding publicly exposed credentials belonging to Hugging Face users in the process. This resulted in the compromise of Hugging Face’s production infrastructure between July 11 and July 13.

The incident highlights the potential for AI agents to exploit and chain together vulnerabilities to compromise third party systems, as well as the potential for agents to escape sandbox environments. It’s worth noting that following the incident, in August, OpenAI announced it had implemented a two-week pause on training its latest models to harden and red team its research environments.

Agents Working Together

On Wednesday, METR and Redwood Research also released the results of their own report on the Hugging Face breach after working on premises for six days at OpenAI. Their investigation focused predominantly on July 7 to July 13, and found that roughly 1,200 agents meant to be isolated from each other began communicating on an unsanctioned message board, sending over 70,000 messages and files during the investigation period.

METR and Redwood Research noted that “agents managed to achieve milestones they could not have achieved working on their own,” finding a way to tamper with the automated scorer used for the ExploitGym benchmark. They report that 700 agents went on to attack Hugging Face.

“When the agents attacked Hugging Face, they’d already figured out how to generate the answer (‘flag’ for their capture the flag problem) for any task. They verified this within hours of finding each other on the message board. Nearly all of their effort went into a different problem: convincing the automated ‘scorer’ that they’d gotten the flag legitimately. That included a large project to tamper with the logs of their own actions. The main reason they attacked Hugging Face was to learn how the scorer worked so they could trick it,” Aleya Cotra, a researcher at METR, told me via email.

The Aftermath

The incident has made waves across the tech industry, with Nvidia CEO Jensen Huang calling out closed AI models for blocking essential forensics during the incident, whereas cognitive scientist and AI critic Gary Marcus called the incident a “wake up call,” sharing concerns over the potential for similar incidents.

Similarly, shortly after the news broke, on July 30, Anthropic released a blog post documenting three occasions where Claude had reached the internet and gained unauthorized access to the systems of a third party organization during testing.

At the same time, OpenAI is also attracting legal challenges following the incident. On Monday, Alabama’s attorney general, Steve Marshall, subpoenaed OpenAI for more information related to the breach and whether OpenAI’s practices “violated Alabama’s consumer protection laws.”

The attorneys general of Alabama and 14 other states, including Florida and Texas have also written a letter to OpenAI asking the startup to preserve “all potentially relevant” documents, data and information related to the event.

It appears that this attack won’t be the last of its kind either. In an interview with The Guardian released on August 23, OpenAI chief global affairs officer Chris Lehane warned of the threat of “persistent” cyber attacks.

“An OpenAI executive warning the public to prepare for ongoing, persistent AI-driven attacks lands different alongside a report that shows exactly what that looks like in practice. A collective of AI agents found vulnerabilities, coordinated with each other through an improvised channel and escalated from a sandboxed test environment into production access, all without a human directing any single step along the way,” Darren Guccione, CEO and cofounder of identity security platform Keeper Security, told me via email.

Lessons Learned?

The fact that multiple agents managed to chain together vulnerabilities to escape a sandbox environment and breach a third party organization to solve ExploitGym suggests that preventing agents from acting in malicious ways is easier said than done, and raises questions about what happens when threat actors set out to intentionally exploit these tools.

“This wasn’t AI going rogue, it was AI cutting corners. The models were given unsolvable tasks and rewarded for finding answers, so they went hunting for the answer key. The real constraint today is configuration and oversight rather than capability, and criminal groups won’t be operating under either constraint,” Safayat Moahamad, advisory director at Info-Tech Research Group, told me via email.

However, the incident also raises questions about accountability. OpenAI itself acknowledges that “with the benefit of hindsight, some early signals identified in this report could have triggered an earlier response,” with responders noting internal activity on a message board on June 27, but deciding not to stop the evaluation run.

In any case, defenders need to adjust to a threat landscape where cybercriminals will have access to more powerful AI models that can chain and exploit vulnerabilities across the attack surface.

AI Hugging Face OpenAI
Share. Facebook Twitter Pinterest LinkedIn Tumblr Email

Related News

The Token Problem May Be A Context-Efficiency Problem

August 26, 2026

How Shadow AI Is Skewing ROI

August 26, 2026

The Biggest Bottleneck In Enterprise AI

August 26, 2026

Why Software Will Not Be Replaced By AI, But Will Be Reimagined

August 26, 2026

The End Of Robotic Process Automation? Not Even Close. Here’s What’s Actually Changing

August 26, 2026

Five Factors For Effective AI-Enabled Enterprise Transformations

August 26, 2026
Add A Comment
Leave A Reply Cancel Reply

Don't Miss

The Token Problem May Be A Context-Efficiency Problem

Tech August 26, 2026

As the CEO of Arango, Shekhar Iyer leads the company’s mission to make enterprise AI…

Nvidia forecasts quarterly revenue above estimates, shares rise

August 26, 2026

OpenAI Finds Agents That Breached Hugging Face Were ‘Reward Hacking’

August 26, 2026

Nelson Peltz’s Trian has no plans to make bid for Wendy’s right now, sources say

August 26, 2026
Stay In Touch
  • Facebook
  • Twitter
  • Pinterest
  • Instagram
  • YouTube
  • Vimeo
Our Picks

How Shadow AI Is Skewing ROI

August 26, 2026

The Good, Bad And Ugly From The Packers’ Final Training Camp Practice

August 26, 2026

Immigrant businesses who spoke out against Mamdani supermarkets say they’re getting hit with NYC sanitation fines

August 26, 2026

The Biggest Bottleneck In Enterprise AI

August 26, 2026
The Financial News 247
Facebook X (Twitter) Instagram Pinterest
  • Privacy Policy
  • Terms of use
  • Advertise
  • Contact us
© 2026 The Financial 247. All Rights Reserved.

Type above and press Enter to search. Press Esc to cancel.