Imagine a hypothetical payment-support agent at a bank. It reads a customer message and calls a payment tool before handing the case to a human. Each step is ordinary access with real security stakes.
This article covers security risks to AI agents and buyer guidance for evaluating controls and tests. CISA's joint agentic AI guide was published May 1, 2026. OWASP's Agentic Threats Navigator describes six attack surfaces: reasoning, memory, tools, identity, human oversight, and multi-agent interactions. NIST NCCoE is exploring standards-based agent identity and authorization as an early-stage effort.
We turn that into attack paths, controls, a hypothetical payment-support-agent scenario, and a threat-to-control-to-test matrix. Production tests and guidance on separating misbehavior from attack follow. It helps bank security and risk leaders assess agentic AI before production.
What do public sources say about securing agentic AI?
CISA's May 1, 2026 joint guide on secure adoption of agentic AI names four cybersecurity risks. They are an expanded attack surface, privilege creep, behavioral misalignment, and obscure event records. Each one maps to a question a security team can ask a vendor.
CISA also gives three practical recommendations. Avoid broad or unrestricted access. Start with low-risk and non-sensitive use cases. Account for agentic AI security in the organization's security model and risk posture.
OWASP's Agentic Threats Navigator lists six attack surfaces: reasoning, memory, tools, identity, human oversight, and multi-agent interactions. That list works as a review checklist. A control plan that covers only the model leaves five surfaces open.
OWASP's highlights Agent Behavior Hijacking, Tool Misuse and Exploitation, and Identity and Privilege Abuse. It also describes FinBot, a capture-the-flag application for practicing agentic security skills in a controlled environment. Your red team can use it before touching production.
On identity, the NIST NCCoE is exploring standards-based approaches to software and AI agent identity and authorization. It is a potential project based on a concept paper. The work is at concept-paper stage, so ask vendors how they handle agent identity today and how they will track that work.
This page covers security risks to agents. For system design, see our agentic AI architecture article.
Which attack paths should you assess first?
Treat each path below as a scenario to test. Start where an agent can move money, change customer data, or reach other agents.
Does the agent hold more permission than its task needs? Ask what identity the agent presents to each system. If it borrows a human's session or a broad service credential, a hijacked agent inherits that reach. Privilege creep matters here too. Permissions granted for a pilot tend to outlive the pilot unless someone reviews them.
Tool misuse. An agent with a valid tool can still call it badly. Test whether the tool layer checks parameters, amounts, and destinations, or trusts whatever the agent sends.
Can untrusted content steer the agent? Agents read customer messages, documents, emails, and web content. Any of those can carry instructions. Test whether untrusted content can change what the agent decides to do.
Memory poisoning. If an agent stores notes or summaries, a bad entry can shape later decisions. Test who can write to memory, how entries are labeled, and whether anyone can remove them.
Can one bad step trigger many? One agent may pass a task to another. Test whether the second agent verifies the request or accepts it because it came from a peer.
Action chains. Each step may look harmless while the sequence does harm. A status lookup and a recall request can combine into something no one approved.
Can you see what happened and stop it? CISA calls out obscure event records. If you cannot reconstruct which input led to which action, you cannot investigate.
How do banks enforce security boundaries on AI agents?
Banks enforce boundaries outside the model. The model proposes, and a separate layer decides what is allowed. Here are the controls to evaluate when you secure agentic banking infrastructure.
- Distinct agent identity. Each agent gets its own identity, separate from any employee or shared account. Ask how the vendor issues, scopes, and retires it.
- Least-privilege entitlements. Grant only the data and actions the use case needs. Review them on a schedule the bank sets.
- Tool allowlists with parameter limits. The agent reaches named tools only. Each tool enforces limits on amount and destination.
- Approval routing. Define which actions need a human before execution. Route anything outside agent authority to an employee with context.
- Input separation. Mark untrusted content as data. Test that it cannot be promoted to instruction.
- Governed memory. Control writes, label provenance, and allow rollback.
- Chain limits. Cap how many consequential steps an agent can take in sequence without a fresh authority check.
- Detailed event records. Log the input, the decision, the policy check, and the action in a form an investigator can follow.
- Containment. Have a way to suspend one agent, one tool, or one use case. Test how long suspension takes in your environment.
On autonomy, CISA's advice is to start small. Begin with low-risk, non-sensitive work and widen access only after tests pass. Visit out dedicate blog about architecture that covers how to review a design before an agent acts.
What could an attack on a payment-support agent look like?
This scenario is hypothetical. It describes no real bank, customer, or incident.
Imagine a bank's payment-support agent. It can read payment status and open cases. It can also request a hold on a pending outgoing payment. A person must approve any release of that hold.
An attacker sends a support message with a PDF attached. Hidden text in the PDF tells the agent to release the hold and mark the case verified.
Here is how your tests could trace the path:
- The agent reads the PDF as part of its normal case review.
- The injected text enters the agent's context alongside real instructions.
- The agent proposes a tool call to release the hold.
- The authority check compares the call with the agent's scope and the approval rule.
- The release routes to a human approver, and the attempt raises an alert.
Now vary the test. Suppose the injected text also writes a note into the agent's memory. Check whether the note changes how the agent handles the next case. Suppose the agent's credential can release holds directly. Check whether step four would catch it.
Each variation tells you which control carries the load. If one control alone stops the attack, you have a single point of failure.
How do you tell agent misbehavior from an attack?
An agent that acts wrongly may be buggy or misaligned, or an attacker may have manipulated it. Your triage needs a fast way to separate these. This sequence is a starting point, and your incident process should own the final version.
- Preserve evidence. Freeze the agent version, its prompt, its memory state, and the session record.
- Check input provenance. Did external content enter the context before the deviation?
- Compare tool calls with baseline. Did the agent call tools or values outside its usual pattern?
- Check identity. Did the credential come from the expected runtime and task?
- Contain. Pause the agent or narrow its scope while you investigate.
- Classify. If the action matches instructions found in untrusted content, treat it as a suspected attack. Escalate to security incident response. If no outside influence appears and the action stayed within permissions, route it to the agent's owner for quality review.
For regulatory reporting and evidence workflows, see our agentic AI compliance article.
What does a threat-to-control-to-test matrix look like?
This matrix is proposed buyer guidance. Adapt it to your environment, and have your own team set the pass criteria.
Which production security tests should you run?
Run these tests before production and repeat them during production. Each test has a description and a pass criterion. Where the bank must decide a target, the bank sets the threshold.
- Scope denial and scope drift, with entitlement review
Description: Ask the agent to act outside its granted scope. Repeat the request over time to check for drift. Review each entitlement against the agent's approved purpose.
Pass criterion: Every out-of-scope request is denied and logged. Behavior stays consistent across repeated runs. Each entitlement traces to an approved need. - Credential substitution
Description: Replace the agent's credentials with invalid ones or with those of another identity. Observe what the agent can still do.
Pass criterion: The agent cannot act under a substituted credential. Access is refused and the attempt is recorded. - Injection through all real input channels
Description: Plant hostile instructions in every channel the agent reads in real operation. Do not limit testing to the chat window.
Pass criterion: The agent follows none of the planted instructions. The bank lists every real input channel, and each one is tested. - Tool boundary and malformed or extreme parameters
Description: Call each tool outside its intended boundary. Send malformed values and extreme values as parameters.
Pass criterion: Out-of-boundary calls are rejected. Malformed and extreme values cause no unintended action. The bank sets the range of extreme values to test. - Memory integrity and cross-session persistence
Description: Plant false or altered content in the agent's memory. Then start new sessions and check what carries over.
Pass criterion: Altered memory does not change the agent's behavior. Nothing persists across sessions unless the bank has approved it. - Chain-break and forbidden multi-step outcomes
Description: Combine steps that are each allowed but together reach a forbidden outcome. Check whether the agent completes the chain.
Pass criterion: The chain is interrupted before the forbidden outcome occurs. The bank defines which outcomes are forbidden. - Reconstruction from logs by an independent analyst
Description: Give an analyst who did not build the agent only the logs. Ask them to reconstruct what the agent did and why.
Pass criterion: The analyst reconstructs the actions correctly from the logs alone. The bank sets the time allowed and the level of completeness required. - Containment: pause and suspend
Description: Pause and suspend the agent while it is working. Measure how long each takes to become effective. Measure the impact on the work in progress.
Pass criterion: The agent stops acting once paused or suspended. Measured time and impact fall within thresholds the bank sets.
Note: The OWASP FinBot exercise is optional preparation for these tests. Teams may use it beforehand. It is outside the required set.
How does Backbase handle agent authority?
Backbase gives intelligence bank-defined authority and the ability to execute across systems, so banks reach accountable production.
The Authority Layer, Sentinel, decides. Every action must be evaluated against identity and policy, evidence and limits. Sentinel holds an authority contract per action, per actor. Each contract sets the autonomy level and monetary limit, plus a confidence threshold and the required evidence. Conditions that force human approval include an amount over the limit or conflicting evidence. A fraud signal or detected customer vulnerability also forces human approval.
Every authorized action yields a Decision Token, a real-time authorization carrying scope and constraints, with reason codes. The bank sets the autonomy level. Autonomy runs from Observe and Recommend through Prepare, to Act with approval and Act within limits. Banks configure it by domain and use case, actor and action. They can tighten autonomy or revoke it without changing the underlying model. Kill switches and segregation of duties remain part of the runtime.
FAQs
How do you secure AI banking agents?
Give each agent its own identity and the least privilege its job needs. Limit its tools, route high-risk actions to people, log every decision, and rehearse containment. Then test against the attack paths above.
How do banks enforce security boundaries on AI agents?
They put enforcement in a layer the model cannot rewrite. That layer checks identity, entitlements, policy, and approvals before a consequential action runs. Tools also enforce their own limits.
How can financial institutions secure agentic payment systems?
Start with narrow payment authority, such as read-only status and recall requests. Put amount and destination limits in the tool layer. Test injection, memory, and chain abuse. Widen scope only when tests pass.
Is there a standard for AI agent identity?
Not yet. The NIST NCCoE is exploring standards-based approaches through a potential project based on a concept paper. Ask vendors what they do today and how they will adapt.
Should we start with low-risk use cases?
Yes. CISA recommends starting with low-risk and non-sensitive use cases and avoiding broad or unrestricted access.
Do these tests replace a compliance review?
No. They cover security risks to agents. Regulatory workflows and evidence belong to a separate review.
