Close Menu
The Financial News 247The Financial News 247
  • Home
  • News
  • Business
  • Finance
  • Companies
  • Investing
  • Markets
  • Lifestyle
  • Tech
  • More
    • Opinion
    • Climate
    • Web Stories
    • Spotlight
    • Press Release
What's On

In-N-Out managers make six figure salaries, company reveals

September 15, 2026

10-year Treasury yield soars, oil surges above $105 as Fed expected to hike interest rates

September 15, 2026

1984 track champ preps South LA bakery for 2028 Olympics

September 15, 2026

Meet Anthropic CEO Dario Amodei’s handpicked far-left ‘evaluators’ he thinks will save us from an AI apocalypse

September 15, 2026

Senate fails to advance Clarity Act in blow to crypto industry ahead of 2026 midterms

September 15, 2026
Facebook X (Twitter) Instagram
The Financial News 247The Financial News 247
Demo
  • Home
  • News
  • Business
  • Finance
  • Companies
  • Investing
  • Markets
  • Lifestyle
  • Tech
  • More
    • Opinion
    • Climate
    • Web Stories
    • Spotlight
    • Press Release
The Financial News 247The Financial News 247
Home » OpenAI Finds Agents That Breached Hugging Face Were ‘Reward Hacking’

OpenAI Finds Agents That Breached Hugging Face Were ‘Reward Hacking’

By News RoomAugust 26, 2026No Comments5 Mins Read
Facebook Twitter Pinterest LinkedIn WhatsApp Telegram Reddit Email Tumblr
Share
Facebook Twitter LinkedIn Pinterest Email

AI agents are reshaping the threat landscape. On Wednesday, OpenAI released a report detailing its findings on the Hugging Face breach that occurred in July. During the incident, an internal-only research model and GPT 5.6 Sol attempted to solve ExploitGym, an evaluation which measures a model’s ability to discover and exploit vulnerabilities, breaching Hugging Face’s internal systems in the process.

OpenAI claims the incident occurred during routine testing in a sandbox environment separate from the public internet after agents engaged in “reward hacking,” or cheating, to solve tasks.“The actions of the models were unintended and were a byproduct of the models attempting to solve the cybersecurity evaluations,” the report said.

These models, harnessed as agents, began communicating with each other through an instance of JFrog Artifactory. The agents used a vulnerability in the service to access the public internet, finding publicly exposed credentials belonging to Hugging Face users in the process. This resulted in the compromise of Hugging Face’s production infrastructure between July 11 and July 13.

The incident highlights the potential for AI agents to exploit and chain together vulnerabilities to compromise third party systems, as well as the potential for agents to escape sandbox environments. It’s worth noting that following the incident, in August, OpenAI announced it had implemented a two-week pause on training its latest models to harden and red team its research environments.

Agents Working Together

On Wednesday, METR and Redwood Research also released the results of their own report on the Hugging Face breach after working on premises for six days at OpenAI. Their investigation focused predominantly on July 7 to July 13, and found that roughly 1,200 agents meant to be isolated from each other began communicating on an unsanctioned message board, sending over 70,000 messages and files during the investigation period.

METR and Redwood Research noted that “agents managed to achieve milestones they could not have achieved working on their own,” finding a way to tamper with the automated scorer used for the ExploitGym benchmark. They report that 700 agents went on to attack Hugging Face.

“When the agents attacked Hugging Face, they’d already figured out how to generate the answer (‘flag’ for their capture the flag problem) for any task. They verified this within hours of finding each other on the message board. Nearly all of their effort went into a different problem: convincing the automated ‘scorer’ that they’d gotten the flag legitimately. That included a large project to tamper with the logs of their own actions. The main reason they attacked Hugging Face was to learn how the scorer worked so they could trick it,” Aleya Cotra, a researcher at METR, told me via email.

The Aftermath

The incident has made waves across the tech industry, with Nvidia CEO Jensen Huang calling out closed AI models for blocking essential forensics during the incident, whereas cognitive scientist and AI critic Gary Marcus called the incident a “wake up call,” sharing concerns over the potential for similar incidents.

Similarly, shortly after the news broke, on July 30, Anthropic released a blog post documenting three occasions where Claude had reached the internet and gained unauthorized access to the systems of a third party organization during testing.

At the same time, OpenAI is also attracting legal challenges following the incident. On Monday, Alabama’s attorney general, Steve Marshall, subpoenaed OpenAI for more information related to the breach and whether OpenAI’s practices “violated Alabama’s consumer protection laws.”

The attorneys general of Alabama and 14 other states, including Florida and Texas have also written a letter to OpenAI asking the startup to preserve “all potentially relevant” documents, data and information related to the event.

It appears that this attack won’t be the last of its kind either. In an interview with The Guardian released on August 23, OpenAI chief global affairs officer Chris Lehane warned of the threat of “persistent” cyber attacks.

“An OpenAI executive warning the public to prepare for ongoing, persistent AI-driven attacks lands different alongside a report that shows exactly what that looks like in practice. A collective of AI agents found vulnerabilities, coordinated with each other through an improvised channel and escalated from a sandboxed test environment into production access, all without a human directing any single step along the way,” Darren Guccione, CEO and cofounder of identity security platform Keeper Security, told me via email.

Lessons Learned?

The fact that multiple agents managed to chain together vulnerabilities to escape a sandbox environment and breach a third party organization to solve ExploitGym suggests that preventing agents from acting in malicious ways is easier said than done, and raises questions about what happens when threat actors set out to intentionally exploit these tools.

“This wasn’t AI going rogue, it was AI cutting corners. The models were given unsolvable tasks and rewarded for finding answers, so they went hunting for the answer key. The real constraint today is configuration and oversight rather than capability, and criminal groups won’t be operating under either constraint,” Safayat Moahamad, advisory director at Info-Tech Research Group, told me via email.

However, the incident also raises questions about accountability. OpenAI itself acknowledges that “with the benefit of hindsight, some early signals identified in this report could have triggered an earlier response,” with responders noting internal activity on a message board on June 27, but deciding not to stop the evaluation run.

In any case, defenders need to adjust to a threat landscape where cybercriminals will have access to more powerful AI models that can chain and exploit vulnerabilities across the attack surface.

AI Hugging Face OpenAI
Share. Facebook Twitter Pinterest LinkedIn Tumblr Email

Related News

Meet Anthropic CEO Dario Amodei’s handpicked far-left ‘evaluators’ he thinks will save us from an AI apocalypse

September 15, 2026

It could ‘kill us all’

September 15, 2026

What The Return To The Moon Means For Global Business

September 15, 2026

How To Build AI Skills That Scale Into Agents

September 15, 2026

Signs Your Software Is No Longer Delivering Value

September 15, 2026

The A2P 10DLC Rules Most Businesses Are Breaking Without Knowing It

September 15, 2026
Add A Comment
Leave A Reply Cancel Reply

Don't Miss

10-year Treasury yield soars, oil surges above $105 as Fed expected to hike interest rates

Business September 15, 2026

The US 10-year Treasury yield hit its highest level since 2007 on Tuesday, as oil…

1984 track champ preps South LA bakery for 2028 Olympics

September 15, 2026

Meet Anthropic CEO Dario Amodei’s handpicked far-left ‘evaluators’ he thinks will save us from an AI apocalypse

September 15, 2026

Senate fails to advance Clarity Act in blow to crypto industry ahead of 2026 midterms

September 15, 2026
Stay In Touch
  • Facebook
  • Twitter
  • Pinterest
  • Instagram
  • YouTube
  • Vimeo
Our Picks

Bakersfield restaurant Uricchio’s closes after 31 years

September 15, 2026

It could ‘kill us all’

September 15, 2026

What The Return To The Moon Means For Global Business

September 15, 2026

LA 2028 Olympics could generate $40B and 224,000 California jobs

September 15, 2026
The Financial News 247
Facebook X (Twitter) Instagram Pinterest
  • Privacy Policy
  • Terms of use
  • Advertise
  • Contact us
© 2026 The Financial 247. All Rights Reserved.

Type above and press Enter to search. Press Esc to cancel.