BankThink

AI agents don't have to fail in order to cause major problems for banks

  • Key insight: As banks move from AI that supports decisions to AI that can take action, they need to be ready for the possibility that an agent's actions may produce consequences the bank did not intend.
  • Supporting data: In August, the U.K.'s AI Security Institute separately reported that agents had taken unsanctioned actions during its evaluations.
  • Forward look: The challenge for banks is to recognize when unexpected agent behavior requires formal incident response.

AI agents are bringing artificial intelligence deeper into bank operations, with systems that can execute tasks, use tools, interact with other systems, and act within defined workflows. This year, Citi introduced Arc, an enterprise platform for developing and deploying AI agents across the bank. BNY continues to scale agentic systems already embedded in its operational workflows, while JPMorganChase is increasingly adopting agentic AI, thinking beyond coding and looking to apply it across the full product development lifecycle.

Processing Content

AI agents are also beginning to participate directly in payments. Visa's U.S. pilots have executed agent-initiated consumer and B2B purchases, while several European banks including Barclays, BBVA, ING, and Lloyds Banking Group have completed live agent-executed transactions.

While banking agents are typically deployed with defined permissions, strict monitoring and strong safeguards, these controls do not eliminate the possibility of unintended behavior.

Recent evaluations have shown how agents can take unexpected actions while pursuing assigned objectives. In July, OpenAI reported that agents in a cybersecurity evaluation circumvented controls and accessed systems outside their intended environment. In August, the U.K.'s AI Security Institute separately reported that agents had taken unsanctioned actions during its evaluations. These evaluations were conducted outside banking environments and under conditions that differ from normal banking deployments, but they are useful for understanding how autonomous systems can pursue assigned objectives in ways that were not anticipated.

The challenge for banks is to recognize when unexpected agent behavior requires formal incident response.

An agent does not necessarily have to fail at its task to create an incident. It can create one while successfully pursuing the task it was given. An agent may complete the task it was assigned and remain within much of its expected workflow, yet take an unexpected action along the way that exposes data, affects a customer, or initiates an unauthorized transaction.

Read more:

What can make an agent-caused incident harder to identify is the fact that the system may still be acting, with consequences still unfolding, while the bank is still assessing the significance of the unusual behavior. The OpenAI evaluation illustrates this problem, as unusual agent activity was observed before its broader significance was understood.

Not every unexpected action should automatically be treated as an incident. Banks already use incident-classification processes to assess the severity and impact of operational events. Agent-caused events need to be incorporated into those processes and assessed according to the seriousness of their operational consequences, which may already fall into familiar categories such as data exposure, unauthorized payments, cybersecurity events, customer-impacting errors, or operational disruptions. The goal is to identify when agent behavior has created an incident and route it into the appropriate existing response process.

Existing banking rules already establish notification requirements when certain computer-security incidents meet specified thresholds. However, these rules were not developed specifically for incidents caused by AI agents. Regulatory frameworks for AI in banking are still developing and have yet to fully address agentic AI. The revised model-risk-management guidance issued this year by the OCC, the Fed, and the FDIC, for example, explicitly excludes generative and agentic AI from its scope.

As banks move from AI that supports decisions to AI that can take action, they need to be ready for the possibility that an agent's actions may produce consequences the bank did not intend. That means including agent-caused events in incident planning now.


For reprint and licensing requests for this article, click here.
Artificial Intelligence Consumer banking Regulation and compliance
MORE FROM AMERICAN BANKER
Load More