2026-2027 Pub. 16 Issue 1

from AI within three years, and roughly one-third devote more than 20% of their digital budget to it while systematically tracking KPIs.6 For an institution considering AI implementation, these three practices can be most useful in driving meaningful impact down the line: 1. Set a baseline metric, a target and a named owner for each AI deployment from day one. 2. Redesign the underlying workflow end-to-end rather than automating a single step. For every process, label each task explicitly as AI-led, human-in-the-loop or human-only, with documented escalation rules. 3. Pair that with a multiyear value thesis, for instance, “+3 points of NPS from AI in loan servicing,” and stage gates that define what evidence is required to move from pilot to scale. What ‘AI’ Actually Means Today Part of the confusion on this topic is linguistic. “AI” is more than just large language models; large language models are only one of roughly five families of machine-learning models, each suited for different jobs. Tree-based models work like auto-generated if-this-then-that rules and excel at fraud detection and delinquency-risk prediction. Linear and logistic models capture straightforward relationships and support risk scoring and segmentation. Neural networks recognize complex patterns and underpin default prediction and fraud detection. Unsupervised models find hidden structure in data, grouping borrowers by behavior or financial stress without being told the right answer. Large language models, the technology behind tools such as ChatGPT, predict the most likely next word in a sequence and, on their own, have no goals or understanding. The more consequential shift for financial institutions is the arrival of AI agents. An agent is best understood as a workflow orchestrator that can think, decide and act, not merely respond. Agents break goals into steps, call systems, and enterprise data and adapt as they go. In servicing and collections, this translates into concrete roles: QA agents who review interactions against custom scorecards, loan-servicing agents who handle payment plans and hardship options, delinquency-management agents who tailor outreach, and back-office agents who process documents and complete routine tasks. The New Risk Categories Because agents take action across multiple steps within a process rather than simply generate text, they introduce risks that traditional model governance was not designed to catch. The most counterintuitive is compounding error. A 98% per-step success rate sounds excellent until five steps are chained together, at which point overall accuracy falls to roughly 90%. That means one in ten workflows would fail. In a regulated collections process, that is not an acceptable margin. Three further failure modes deserve board-level attention. Agents can hallucinate or misrepresent policy, inventing non-compliant financial advice, referencing competitors the institution does not endorse, or making insensitive assumptions about a borrower’s circumstances. They can miss edge cases, failing to recognize manipulative language or legal traps; a borrower who says “I have no money to pay you” may be invoking a script that makes further contact a potential FDCPA or RFDCPA violation. And they can be manipulated directly, through prompt-injection attempts that try to hijack the agent into saying something off-brand or non-compliant. Mitigations for these cases should be non-negotiable. Hard-coded guardrails and specially trained oversight models that check other models’ work can catch hijacking attempts and block poor responses before they reach a borrower. Rigorous evaluation in simulated environments also quantifies error rates across categories such as jailbreaking, hallucinations, content moderation, competitor references and policy-sensitive requests, such as bankruptcy or cease-and-desist notices. Finally, human-in-the-loop review remains essential for high-uncertainty cases. These measures build confidence in risk management by grounding error rates in concrete data and expressing them as a number that can be insured against. What Regulators Expect Supervisory expectations from the CFPB, FDIC, OCC and FFIEC are converging around four themes, and institutions should map any AI deployment against all of them: • Explainability: The ability to explain, in plain language, how AI is used in a product or process; specific and accurate reasons for credit denials and other adverse actions, with no “black box” excuses; and documented model purpose, assumptions and limitations. • Borrower Impact: Demonstrated testing for disparate impact and unfair bias across protected classes; ongoing monitoring of approval rates, pricing and servicing outcomes; and clear disclosures, opt-outs where applicable and easy access to a human review or complaint path. • Model Governance: Treating AI as “models” under SR 11-7-style guidance, with an inventory, named owners and use-case documentation; independent validation before go-live and regular back-testing for drift and performance; and board and senior-management oversight. • Data Privacy: Lawful data use under FCRA, GLBA and UDAAP, with limits on surveillance and non-traditional data; strong vendor controls, including rights to audit, explain, remediate and shut down AI systems; and cybersecurity, access control and logging around training data, prompts and outputs. Where the Opportunities Are When deployed with clarity and discipline, AI can drive material gains across an institution. Use cases most beneficial for banks 17 Colorado Banker

RkJQdWJsaXNoZXIy MTg3NDExNQ==