The Trust Layer for Office Files: Catching the Pretty-but-Wrong Financial Model Before It Ships

AI now drafts the board deck, the financial model, and the regulatory memo, and the broken formula, the hardcoded actual, and the unlabeled assumption all pass silently because the output looks finished. A four-stage trust layer makes numerical and regulatory output defensible to an examiner.

A regulated mid-market organization, a behavioral health network, a PE portfolio company, a property manager holding owner funds, a SaaS firm under a financial-services framework, does not get audited on the quality of its prompts. It gets audited on the documents it produced and acted on. The board reserve memo. The covenant compliance calculation. The Annex III risk classification. The reimbursement model that set staffing for a quarter. AI now writes all of those, and it writes them so cleanly that the single most dangerous property of the output is that it looks finished. A broken cell reference, a hardcoded actual sitting where a formula should be, a discount rate nobody can source, none of those announce themselves. They render in the same crisp typeface as the numbers that are correct. The shift in 2026 is not that AI can build an Office file. It has been able to do that for two years. The shift is that the output is now good enough to skip review. When Claude Opus 4.8 launched on 28 May, roughly 41 days after 4.7, it added dynamic multi-agent workflows: one agent drafts the model, another writes the memo, a third formats the deck. The handoffs are invisible, and so are the errors that propagate across them. A regulated buyer who reads this as a productivity story misses the control problem. The high-risk tier of AI output is not the marketing email. It is numerical synthesis and regulatory language, the two things a regulator and a board cannot un-rely on once they have relied on them. What changed: the output got good enough to skip review The failure mode here is specific, and it is not hallucination in the usual sense. The model does not invent a fact in a sentence you can fact-check. It produces a structurally plausible artifact whose internals are wrong. Three patterns recur in the work we do with regulated mid-market clients. The first is broken formulas that still return a number. An agent assembling a model will paste a computed value where a live formula belonged, or wire a SUM across the wrong range. The cell shows a number. The number is wrong, and because it is hardcoded it will not move when the inputs do, so next quarter's refresh silently carries last quarter's answer. The second is hardcoded actuals masquerading as projections. The model fills a forward-looking column with a figure that is really a backward-looking actual, or a placeholder it never replaced. The deck says "FY27 projected reserve: $4.1M." The 4.1 is a number the agent supplied to make the row look complete. The third is unlabeled assumptions. Every model rests on a discount rate, a growth assumption, a reimbursement rate, a churn figure. When a human builds the model, those live in an assumptions block someone defends. When an agent builds it, they are scattered inline, unsourced, and indistinguishable from facts. The board cannot challenge an assumption it cannot see. None of this is exotic. The illustrative case that made the rounds, an AI coding agent that deleted a startup's production database and its volume-level backups in roughly nine seconds, is the same class of failure in a different medium: confident action on a flawed internal model, with no checkpoint between the action and the damage. In an Office file the damage is slower and quieter, which makes it worse, because it ships to a board before anyone notices. Why the old control model misses it Mid-market document controls were built for a world where a human authored the file and a human reviewed it, and review meant reading the prose. That model assumed two things that no longer hold: that the author understood the numbers, and that errors looked like errors. AI output breaks both. The "author" is a model that cannot tell you why cell G14 is 0.18. And the errors do not look like errors, they look like the rest of the polished, well-formatted output. Most document QA also reviews the wrong layer. A reviewer reads the narrative and skims the headline figures. Nobody traces the formula chain, because for human-built models the formula chain was trustworthy by construction. With AI-built models the formula chain is exactly where the risk concentrates, and it is the one thing a glance never catches. The regulatory environment is moving toward exactly this scrutiny. Texas TRAIGA has been in force since 1 January 2026, with a safe harbor for organizations that follow the NIST AI Risk Management Framework, so your documentation of how AI output was verified is now load-bearing for the defense. NIST's preliminary Cyber AI Profile (IR 8596), published 16 December 2025, frames AI output integrity as a control domain rather than an engineering afterthought. And under the EU AI Act, the "Digital Omnibus" of 7 May 2026 deferred the high-risk Annex III obligations to December 2027, but GPAI enforcement powers still switch on 2 August 2026, with fines up to 3% of global turnover, and downstream deployers are expected to collect provider documentation now. The grace period applies to the heaviest obligations, not to basic defensibility. What the examiner and the board will ask When an AI-built artifact drives a decision and the decision is later questioned, the questions are predictable. A regulated buyer should be able to answer all of them about any number that left the building. Where did this number come from, a formula, a source system, or an assumption? If it is an assumption, who set it and on what basis? Can you reproduce this figure today from the same inputs? Who reviewed this output, and what did "review" actually consist of? What did the AI generate, and what did a human change before it shipped? The US Treasury's Financial Services AI RMF, published 19 February 2026 with its 230 control objectives across seven domains, is explicit that for financial outputs the burden is on the deployer to evidence the verification, not merely assert the model's competence. "The model is good" is not an answer an examiner accepts. "Here is the checks tab, the reviewer sign-off, and the provenance trail" is. The four-stage trust layer This is the artifact we install between an AI-built model and the people who rely on it. Four stages, applied to every numerical or regulatory output that reaches a board, a regulator, an owner, or an external counterparty. It is deliberately low-tech, because it has to survive in a mid-market shop with no platform budget. Stage one, the checks tab. Every model carries a dedicated tab whose only job is to fail loudly. Balance checks that must net to zero. Cross-foots that must tie. Reconciliations to the source system. Range tests that flag any growth rate, discount rate, or margin outside a defensible band. A red cell on the checks tab blocks the file from shipping. This is the cheapest control with the highest yield: it catches the broken-formula and hardcoded-actual failures mechanically, without a human noticing anything. Stage two, the hostile-reviewer pass. One named person whose assignment is not to confirm the model but to break it. They assume every number is wrong until traced. They probe the assumptions the model buried inline. For high-stakes outputs, a second AI instance, a different model or a fresh context, never the one that authored the file, runs an adversarial review whose findings a human adjudicates. The point is structural disagreement, not a second pair of agreeable eyes. Stage three, provenance on every number. Build the model so each material figure resolves to one of three labeled categories: formula (traceable to other cells), source (tied to a named system and date), or assumption (owner and rationale recorded). A simple provenance column does this. The discipline is that no number ships uncategorized, an unlabeled number is treated as an error, exactly like a red checks-tab cell. Stage four, the no-self-certification rule. The AI does not attest to its own work, and neither does the person who prompted it. The reviewer who signs off is not the author. This is the oldest control in finance, segregation of duties, applied to a new author. It is the rule examiners look for first, and the one AI workflows quietly erase when "the analyst and the agent" collapse into a single keystroke. What we recommend Four concrete moves for a regulated mid-market buyer this quarter. First, classify your AI outputs by blast radius. Numerical synthesis and regulatory language are the high-risk tier and get the full four stages; a meeting summary does not. Spend the control budget where a wrong answer reaches a board or an examiner. Second, ship a standard checks tab as a template. Make it the default starting point for every model in the organization, AI-built or not. The marginal cost is near zero, and it converts the most common AI failures into mechanical, visible ones. Third, write the no-self-certification rule into your AI use policy and name the reviewers. A policy that says "review AI output" and never defines review is the policy an examiner discounts. Define it: hostile pass, provenance check, independent sign-off. Fourth, treat this as one layer of a larger stack, not the whole answer. Output verification sits downstream of model governance, identity, and runtime controls, the full picture is in our field guide, The Five-Layer AI Compliance Stack for Regulated Mid-Market (/blog/five-layer-ai-compliance-stack-mid-market-regulated), and a Securem Diagnostic is where we map it to your specific models and files. The trust layer is the one that decides whether the document you already shipped survives the question that comes after. The model will hand you something that looks finished. Your job is to make "looks finished" and "is defensible" the same sentence, before it ships, not after the examiner asks.