Anthropic's model breakout another wake-up call for banks

Anthropic CEO Dario Amodei
Anthropic CEO Dario Amodei at the company's headquarters.
Jason Henry/Bloomberg
  • Key insight: Anthropic disclosed that its models broke out of testing environments and attacked other companies. Bankers need to harden their systems in defense.
  • Expert quote: "It's something we'll keep monitoring while we reinforce the importance of what data, privileges and environments our models have access to, and make sure we have clear forensic capability over their actions," said Peter Chapman, chief technology officer at Grasshopper Bank.
  • Forward look: Anthropic is improving security for its models and testing sandboxes. But observers say banks need to take their own defensive steps against "rogue" AI agents.

Frontier artificial intelligence creators such as OpenAI and Anthropic seem to be in a race to prove their models are the most perilous and therefore most advanced. Ten days after OpenAI disclosed that two of its agents had escaped a test environment and hacked the Hugging Face AI model library, Anthropic said it reviewed 141,006 cybersecurity tests of its models and found three cases in which they too had escaped testing sandboxes and hacked into other companies' computers. In one of those cases, the model also downloaded malicious code. 
"There is definitely an element of lab competition here, and disclosing a dangerous model has become its own way of signaling AI capability," said Sumeet Chabria, CEO of consultancy ThoughtLinks and former chief operating officer for global tech and ops at Bank of America.

Processing Content

"Rogue" AI agents are a cybersecurity minefield bank leaders and technologists can't evade. Several banks, including JPMorganChase and Goldman Sachs, use advanced generative AI models from Anthropic, OpenAI and others. Even without working with these models, any bank could potentially become a victim of an AI agent that escapes its test environment.

"The frontier is indeed scary," said Alenka Grealish, lead analyst for emerging technology at Celent. "The push to test model boundaries is getting ahead of robust AI governance checks."

"It looks like some human error in both cases, but it also shows the raw power of these models if they ever did get loose," Peter Chapman, chief technology officer at Grasshopper Bank, told American Banker. "It really feels like the sci-fi of a few years ago is becoming reality."

Staying the course

Grasshopper Bank uses fully released models from Gemini and Anthropic, and Chapman said he has no plans to stop using them despite the recent disclosures about Anthropic and OpenAI models. 

"It's something we'll keep monitoring while we reinforce the importance of what data, privileges and environments our models have access to, and make sure we have clear forensic capability over their actions," he said.

Financial institutions can't simply ignore advanced AI models because of the risks, Chabria said. 

"Telling banks to avoid autonomous AI and its promised benefits is like telling them to avoid the internet in 1999," he said. "It is a non-starter if you intend to remain competitive. Going slow is an option but stalling is not."

Similar concerns have existed for earlier technologies, pointed out Gary McAlum, senior advisor at Oliver Wyman and former chief security officer at USAA. 

"The risk isn't that these systems are inherently dangerous; it's that weak or misconfigured environments can turn any capable agent into an offensive actor," he said. "This isn't a new problem, but the power of frontier models dramatically increases the risk in terms of speed, scale and complexity."

Banks, like all companies, need to elevate their cybersecurity infrastructure, eliminate vulnerabilities and wrap even tighter controls around the technology, Chabria said.

"Frankly, banks are better positioned to solve this than almost any other sector," he said. "Financial institutions have spent decades building world-class compliance, conduct and risk-management frameworks."

What Anthropic's models did

Anthropic disclosed Thursday that three of its AI models — Opus 4.7, Mythos 5 and an internal research model — inadvertently accessed the internet and compromised external systems during "capture-the-flag" cybersecurity tests. The tests were run in environments that were supposed to be fully sealed off from the internet but weren't. Anthropic said this was due to a miscommunication between Anthropic and its model testing partner, Irregular.

Tasked with retrieving hidden "flags" in a challenge, the models operated under the false belief that they were in a simulation. They treated outside companies' systems as part of the exercise, using basic hacking techniques such as exploiting weak passwords and insecure endpoints.

The most serious incident involved Claude Mythos 5, which navigated multiple hurdles to upload a malicious Python package, including registering for an email account and phone number. While the model's internal reasoning identified these actions as potential real-world attacks, it convinced itself the environment was not real and that the security measures it encountered were part of the simulation.

"The lengths Claude went to … fall short of ideal behavior, and this is an area where we will focus more training," Anthropic wrote in its blog. 

Anthropic attributed these failures to a lack of "situational awareness" — in other words, the models failed to distinguish between simulated and actual infrastructure. 

The company said it's implementing stricter security standards for its test environments. It's also stepping up its continuous monitoring of evaluations, refining investigation tools and working with independent organizations like METR to conduct third-party reviews. 

Three takeaways for banks

Banks can glean several things from the Anthropic reveal. 

So-called "sandboxes" are not watertight. In a LinkedIn post written shortly after OpenAI made its disclosures, Irregular wrote, "A model pursuing a goal treats a boundary as part of the problem, and solves it along with everything else. The controls that contained software do not reliably contain a model that can reason past them. None of this surprised us. In our own evaluations at Irregular, capable models break into hardened, production-grade environments, and with each generation they do it more reliably." Irregular did not reply to a request for comment. 

"The whole point of those environments is to let the models run loose and test full capability without the production safeguards," Chapman said.

But the only truly unbreakable environment would be a completely air-gapped, offline system that would be useless for real-world banking operations, Chabria pointed out. "In this latest incident, the environment was believed to be isolated and wasn't," he said.

And trying to build security into AI model prompts isn't enough, he said.

"Assuming you can secure an AI simply by telling it how to behave in a prompt is a dangerous myth," Chabria said. "You cannot rely on pre-task instructions or post-task audits. You need in-the-moment inference controls: hardcoded circuit breakers, honeypot tripwires [fake digital assets or decoys placed in a network to catch hackers] and human-in-the-loop gates that evaluate actions at the exact moment of execution."

More human review may be needed for AI agents. The Anthropic incidents "had an element of human error, which a maker-checker model could have prevented," Grealish said, referring to a security and control process where two different people review work before it's completed. 

Frontier models probably need more than two humans reviewing them, added Grealish, who recently published a report on AI governance. 

Basic cybersecurity defenses still work. "The real story for banks is not about sentient models waking up, it's still about technology hygiene and infrastructure management," Chabria said. "These recent incidents didn't happen because an AI invented an unstoppable exploit," but because Anthropic's models took advantage of weak passwords, unsecured API endpoints, open network ports and misconfigured sandboxes. 

"The capability of these models is surely accelerating, and that serves as a reminder that enterprise tech hygiene needs to keep pace," he said.

Banks can deploy frontier models safely if they have true isolation, zero‑trust controls and continuous adversarial evaluation, McAlum said. 

"The real work now is building hardened, multi‑layered environments that agents cannot break out of — and that's exactly where the industry is moving," he said. "This moment is another wake‑up call, not a stop sign."


For reprint and licensing requests for this article, click here.
Artificial Intelligence Bank technology
MORE FROM AMERICAN BANKER
Load More