Treasury's 230 Controls: The Financial-Services AI RMF Examiners Will Quietly Use to Scope Your Next Exam

Treasury turned NIST's AI RMF into 230 control objectives across seven domains. Voluntary on paper, it will anchor exam scoping for regulated financial firms, here is how a mid-market team should map to it before an examiner does it for them.

On February 19, 2026, the US Treasury published a Financial Services AI RMF, a sector-specific translation of NIST's voluntary AI Risk Management Framework into 230 discrete control objectives across seven domains. If you run a regulated mid-market financial firm, a community bank, a lending fintech, a payments company, a servicer inside a private-equity portfolio, read the word "voluntary" carefully and then set it aside. Voluntary frameworks do not stay voluntary in financial services. They become the structure an examiner uses to organize a conversation, the checklist a state regulator borrows for scope, and the diligence grid a PE owner staples to the next audit. The CFPB, the OCC, the state banking departments, and the prudential examiners do not need a new rule to ask you 230 questions. They need a credible, government-published taxonomy. Treasury just handed them one. This briefing is not about fair lending or disparate impact, that is a separate problem with its own controls. This is about the framework itself as an examination benchmark: what the seven domains are, why your current control inventory almost certainly does not line up against them, and what a mid-market team should do in the next quarter to map what it already has before someone external does the mapping for it. What Treasury actually published NIST's AI RMF is a general-purpose framework organized around four functions, Govern, Map, Measure, Manage. It is deliberately abstract, sector-neutral, and outcome-oriented. That abstraction is its strength for a research lab and its weakness for a bank examiner, who needs to know whether a specific model has a specific control with specific evidence. Treasury closed that gap. It took the four NIST functions and decomposed them into 230 control objectives grouped under seven domains covering the AI lifecycle as a financial institution actually experiences it: governance and accountability; data management and quality; model development and validation; deployment and integration; monitoring and performance; third-party and vendor risk; and resilience, security, and incident response. The names vary slightly depending on how you read the document, but the shape is consistent, it walks an AI system from the moment a business unit proposes it to the moment it fails in production and someone has to explain why. The number matters more than it looks. Two hundred and thirty objectives is granular enough that "we have an AI policy" is not a defensible answer to any single domain. It is the difference between a regulator asking "do you govern AI?" and a regulator asking "for this credit model, show me the validation evidence, the data lineage, the vendor attestation, the drift monitoring threshold, and the incident runbook." The first question you can talk your way through. The second you cannot. Why your current control model misses it Most mid-market financial firms already have a control framework. It is almost never built around AI. It is built around the things examiners have asked about for twenty years: model risk under SR 11-7, third-party risk, BSA/AML, information security mapped to FFIEC or a flavor of NIST CSF. Those frameworks are real and useful, and they cover perhaps half of what Treasury's seven domains demand. The other half is the part that did not exist when those frameworks were written. The gaps cluster in predictable places. SR 11-7 governs models, but it assumes a model you built or licensed and can document, not a foundation model accessed through an API whose weights you will never see, whose training data you cannot inspect, and whose behavior changes when the vendor ships a new version. When GPT-5.5 Instant became the ChatGPT default on May 5, 2026, it was exposed as a floating "chat-latest" alias, meaning a firm pointing production at that endpoint had its model silently swapped underneath it. Treasury's model-development and monitoring domains expect you to detect and govern exactly that. Your SR 11-7 inventory probably does not even list it as a model. Third-party risk is the second gap. Your vendor framework was built for a SaaS provider with a SOC 2. It was not built for the supply chain underneath an AI feature, where a poisoned package can reach you transitively, the LiteLLM PyPI backdoor in March 2026 propagated into CrewAI, DSPy, and GraphRAG, libraries a fintech engineer might pull in without a procurement review. Treasury's vendor-risk and resilience domains expect provenance, attestation, and an incident path for the AI dependency tree. Most mid-market vendor inventories stop at the named vendor. The third gap is security as it applies to AI specifically. The vulnerabilities now landing are not in the model, they are in the orchestration layer wrapped around it. Microsoft's Semantic Kernel, the framework many teams use to wire models into applications, carried remote-code-execution flaws (CVE-2026-26030 and CVE-2026-25592) disclosed on May 7, 2026. Prompt injection has become a primary attack class against any system that lets a model read untrusted input and then take an action. Treasury's resilience-and-security domain treats all of this as in scope. A traditional infosec program treats an AI feature as "a vendor product we already trust" and never tests the layer where these failures actually live. What the examiner will ask Assume the examiner has the 230 objectives open and is not going to walk all of them. They will sample. They pick a handful of AI systems that are actually in production, a fraud model, a chatbot, a document-processing pipeline, and they trace each one across the seven domains. For each system they want a coherent story: who owns it (governance), where its data comes from and whether that data is fit for use (data management), how it was validated before launch (model development), how it was integrated and what it is allowed to touch (deployment), how you know it is still working (monitoring), who is on the hook in the vendor chain (third-party), and what happens when it breaks or is attacked (resilience). The failure mode is not having no answers. It is having answers that do not reconcile. The model your CRO describes in the governance interview is not the model the engineer describes in the deployment interview, and neither matches what the vendor contract actually permits. When the seven domains do not tell a single consistent story about one system, the examiner stops sampling and starts writing. The mapping artifact to build now The work we do with regulated mid-market clients almost always starts with one artifact, and it is the same one Treasury's structure now rewards. Build a control-mapping register, a single sheet, one row per production AI system, with columns that mirror the seven domains. For each system and each domain, three fields: the control objective in plain language, the existing control you already run that satisfies it (with a pointer to the evidence), and a gap flag where you have nothing. The register is not impressive when it is full of green. It is valuable precisely when it shows red, because a documented, dated gap with a remediation owner is a control posture, and a blank is a finding. The columns: System | Domain | Treasury Objective | Existing Control | Evidence Location | Gap (Y/N) | Owner | Target Date. Eight columns, every production AI system, refreshed quarterly. That is the difference between handing an examiner a narrative they can audit and handing them a vacuum they will fill with their own conclusions. This is also where the broader picture matters: as we documented in The State of Mid-Market AI Compliance 2026, the firms that struggle in exams are rarely the ones with the most AI, they are the ones who cannot produce a single sheet showing what AI they run and which controls cover it. What we recommend First, inventory before you map. You cannot map 230 objectives against systems you have not listed. Run a two-week pass to enumerate every production and pilot AI system, including the ones embedded in SaaS tools your staff turned on without procurement, model-backed features in vendor platforms, anything pointed at a floating model alias. Second, pin your models. Replace every floating alias with a stable, versioned endpoint, for Claude, the stable id "claude-opus-4-8" rather than a moving pointer, so a vendor's silent upgrade is a decision you make, not an event you discover mid-exam. The GPT-5.5 default swap is the warning: production traffic can be redirected onto a moving target without anyone inside the firm choosing it. Third, map against the seven domains, not against your comfort zone. The temptation is to map the four domains you already cover well and wave at the other three. The examiner will sample the three you skipped. Fourth, date your gaps and assign owners. A gap with a name and a target date is a managed risk. A gap with neither is a finding waiting to be written. A short diagnostic against the seven domains will tell you which of the two you are holding. The framework is voluntary. The exam is not. Build the register before someone builds it for you.