OpenAI models' Hugging Face breach is a red flag for bankers

Sam Altman, CEO of OpenAI
"We had a significant security incident during evaluation of our models. we are sharing what we have learned so far," OpenAI CEO Sam Altman wrote on Twitter of the Hugging Face breach. "Thanks to @huggingface for the partnership on this."
David Paul Morris/Bloomberg
  • Key takeaway: An advanced AI agent being tested by OpenAI unexpectedly broke out of its "sandbox" and hacked into the servers of another company.
  • Expert quote: "The unprecedented is becoming the norm." –Alenka Grealish, Celent analyst
  • What's at stake: Model risk management programs will have to be updated for the increasingly unexpected capabilities being exhibited by AI agents. 

OpenAI's extraordinary admission that its models hacked into the AI software library of a tech startup named Hugging Face while being tested within a sandbox touched a nerve for even some of the most enthusiastic AI supporters. 

Processing Content

"For years, many of us warned that artificial intelligence might one day carry out a cyberattack, beginning to end, no person at the keyboard. That day has arrived," wrote Rob T. Lee, chief AI officer of the technology research and training organization SANS Institute, in a blog.

The OpenAI models, GPT‑5.6 Sol and an unnamed but "even more capable pre-release model" broke out of their sandbox, breached Hugging Face's production infrastructure, stole data and harvested credentials.

OpenAI said it's still figuring out how this happened.

"We will continue to conduct a thorough investigation alongside Hugging Face and will share more details on the vulnerabilities, incident, and findings when our investigation is complete," an OpenAI spokesperson said in a statement.

This isn't the first time an AI model has escaped a sandbox to do something it shouldn't do. There have been several incidents of AI agents going rogue this year. In March, Alibaba researchers discovered that internally developed AI agents had gotten out of the training servers they were sandboxed in, tried to access internal network resources they did not have rights to and mined cryptocurrency. In April, Claude Mythos Preview broke out of its sandbox and then posted details about what it had done on public websites.

The OpenAI models' breach of Hugging Face is the first real example of a frontier‑class AI model escaping a sandbox and conducting a cyber intrusion, according to Gary McAlum, a senior advisor at Oliver Wyman and former chief security officer at USAA. 

"But it won't be the last," he said. "Banks shouldn't panic, but they should recognize the implications: highly capable models can chain vulnerabilities and behave like autonomous attackers if isolation or supply‑chain controls are weak." (Vulnerability chaining is the practice of combining multiple low or medium-severity vulnerabilities to conduct a massive breach.)

The episode demonstrates the frontier models' exponential learning curve, according to Alenka Grealish, lead emerging tech analyst at Celent. 

"Mythos only recently shocked us with its ability to identify unknown zero-day vulnerabilities at lighting speed," Grealish told American Banker. "The OpenAI-Hugging Face incident demonstrates what happens when one reduces cyber refusals for evaluation purposes and models take advantage. The unprecedented is becoming the norm, putting banks' cybersecurity teams on continual high alert."

It also illustrates the added risks banks take as they roll out AI agents that can act autonomously, said Sumeet Chabria, CEO of consultancy Thoughtlinks and former global corporate operations officer at Bank of America and former chief information officer at HSBC. 

"AI agents, by definition, drift, hallucinate and can dynamically select their own paths," Chabria said. 

They can breach sandboxed environments, unless they're paired with security and governance software explicitly built to contain them, he said. 

What happened

Hugging Face is an open-source AI and machine learning platform that provides access to more than 45,000 models from AI providers and is used by more than 50,000 organizations. 

In a July 16 blog, it revealed that it had been hacked. The intrusion was "different from anything we had handled before in one important way: it was driven, end to end, by an autonomous AI agent system — and we detected and dissected it largely with AI of our own," the company wrote.

On Tuesday, OpenAI posted a blog saying its models were responsible. "We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly," the company wrote.

The OpenAI models had been instructed to try to pass a hacking exam, the ExploitGym public AI cybersecurity benchmark. They were supposed to analyze a software vulnerability and create and execute unauthorized code to retrieve a hidden secret flag. But instead of doing this work on their own, the models decided the test solutions might already exist in Hugging Face's software library and that they should try to steal them. 

Read more:

In one of their attempts, the OpenAI agents chained zero-day vulnerabilities and used stolen credentials to find a remote code execution attack vector while trying to gain access to Hugging Face servers.

When Hugging Face's own responders investigated the intrusion, the frontier models they use would not do the forensic work, they were blocked by guardrails. The safety layer could not tell the difference between a responder analyzing an attack and an attacker running one, so it treated the defender like a threat. 

What banks should do about it

Most banks have model risk management programs — they are required to by the major bank regulators.

But these frameworks were built for deterministic software and statistical models, not agentic AI systems capable of autonomous action, exploration and exploitation, McAlum pointed out. 

Model risk management "now needs to expand to cover model behavior, containment and software‑supply‑chain exposure," he said. "The train has left the station so things will only get more challenging."

There are existing security platforms banks can turn to, including one being developed by cybersecurity vendor HiddenLayer, that are designed to detect anomalous or adversarial model activity and validate the integrity of AI components, he said. 

"These aren't replacements for model risk management, they're the next layer of defense required for frontier‑class systems," McAlum said.

In the near term, he said, banks should tighten model isolation, deploy AI‑aware intrusion detection and use "red teams" – groups of people trained to mimic the work of hackers to test their AI agents. 


For reprint and licensing requests for this article, click here.
Artificial Intelligence Bank technology
MORE FROM AMERICAN BANKER
Load More