Opus 4.8 and the Routing Discipline: Why Max Effort Can Underperform High
Opus 4.8 leads knowledge-work benchmarks and adds a user-selectable effort control, but max effort can cost 5x the tokens and do worse than high on long tasks. For regulated mid-market buyers, that makes procurement a routing-and-evaluation discipline, and effort level a governance choice the board will ask about.
If your organization is regulated, a behavioral health provider, a PE portfolio company answering to a quarterly board, a healthcare nonprofit with a HIPAA footprint, a SaaS vendor carrying SOC 2 and now state AI law, you have probably standardized on a frontier model the way you standardized on a cloud. You picked a vendor, signed the BAA or the DPA, and moved on. Claude Opus 4.8, which shipped on May 28, 2026, is the moment that posture stops being defensible. Not because the model is weak, it leads knowledge-work benchmarks, but because it ships with a user-selectable effort control, and the new dial breaks the assumption that "use the best model on its strongest setting" is the safe, conservative choice. It often is not. Max effort can consume roughly five times the tokens of a lower setting and still produce worse results on long-horizon tasks than the plain "high" setting. The thing you would reach for to look careful in front of an auditor is the thing that can quietly degrade the work. What changed: the effort dial is now part of the model Opus 4.8 arrived only forty-one days after 4.7, which tells you the release cadence your governance program is now pacing against. The substantive change is not a benchmark number. It is that the model exposes two new control surfaces: dynamic multi-agent workflows, where the model spins up and coordinates sub-agents, and a user-selectable effort level. The stable API id is , so pinning the model is straightforward. The effort level is not a model, it is a parameter, set per request, often defaulted by whoever wired up the integration and never revisited. Here is the counterintuitive part that belongs in your risk notes. On long-horizon work, multi-step analysis, document synthesis, anything where the model plans across many turns, max effort can burn about five times the tokens of "high" and finish behind it. More deliberation produces more chances to wander, more self-generated context to get lost in, and a higher bill for the privilege. The intuition that "more effort is more careful, and careful is what regulators want" is wrong here in a way that costs money and degrades output at the same time. The dial is a governance object, not a power slider. Why brand loyalty is the wrong procurement model The market is making the same point from several directions at once. GPT-5.5 Instant became the ChatGPT default on May 5, exposed as a floating alias, meaning the model behind your prompt can change underneath you with no version bump, the classic model-pinning risk. Google announced Gemini 3.5 and Gemini Omni at I/O on May 1, and released the Gemma 4 open-weight models the same day, a self-hosting path for data that genuinely cannot leave your perimeter. Capable frontier families, each with their own dials, aliases, and self-hosting tradeoffs, and no two of them measured the same way. We have made the underlying argument before. Our briefing on the sabotage risk report warned that a model can be capable and untrustworthy in the same breath, and our work on distilled models warned that a cheaper, smaller model can quietly substitute for the one you evaluated. Opus 4.8 advances that lineage into a sharper claim: even within a single vendor, on a single model, the configuration you ship is a control decision. Brand loyalty, "we're a Claude shop," "we're an OpenAI shop", answers a question no auditor is asking. The question is whether each task is routed to a model and a setting that has been evaluated for that task, and whether you can show the evidence. The principle the cost discipline already named for this is "minimum effective intelligence", route each task to the cheapest model and setting that still yields an accepted result. The effort dial extends it: minimum effective effort, proven by evaluation rather than assumed by reflex. What the audit and the board will ask The infrastructure to answer is now in place, which removes your excuse. Amazon Bedrock added request-level usage attribution on May 20 and brought OpenAI frontier models, Codex, and managed agents into preview on April 28. Microsoft, having renamed Azure AI Foundry to "Microsoft Foundry" effective January 1, shipped project-level cost attribution and brought Managed VNet to GA on May 31. The FinOps Foundation named AI cost management the single most-wanted skill for 2026, with roughly 98% of organizations now managing AI spend as of late January. When per-request attribution exists and you are not using it, "we didn't have visibility" is no longer a finding you get to write. The regulatory frame is converging on the same expectation. The EU AI Act's Digital Omnibus, on May 7, deferred the high-risk Annex III obligations to December 2027, but GPAI enforcement powers still activate August 2, 2026, with fines up to 3% of global turnover, and downstream deployers must collect provider documentation now. The US Treasury's Financial Services AI RMF, published February 19, lays out 230 control objectives across seven domains. Texas TRAIGA has been in force since January 1 with a NIST AI RMF safe harbor, and NIST's preliminary Cyber AI Profile (IR 8596) landed December 16. None of these care which logo is on your model. All of them care whether you chose it deliberately, documented why, and can reproduce the decision. Expect three questions. Which model and effort level handled this regulated workflow, and who set it? What evaluation justified that choice, and when was it last re-run against the current model version? And what did it cost, attributed to that request? If your answer to any of the three is a shrug, the dial chose for you. The routing-and-evaluation artifact The control to put in place is a routing register, one row per regulated AI workflow, reviewed each quarter. The columns: Workflow and data sensitivity (PHI, PII, financial, M&A, none) Model id pinned (e.g. ), never a floating alias like Effort level and who set it Eval that justifies this model and effort, and date last re-run Honesty eval: does the model report uncertainty and refuse cleanly, not just produce a fluent answer Per-request cost attribution source and budget cap Fallback model/setting and self-hosting option (e.g. Gemma 4) if data cannot leave the perimeter The non-obvious column is the honesty eval. Your evaluations must separate honest behavior from merely effective behavior. A max-effort run that produces a confident, wrong, expensive synthesis is worse than a "high" run that flags what it could not verify. Benchmarks reward the first. Auditors and patients pay for the second. If your eval suite only measures accuracy on clean inputs, it will recommend exactly the setting that hurts you. What we recommend Four moves a buyer can make this quarter, the kind we walk through in the fixed-scope Diagnostics we run with regulated mid-market clients. First, pin models by id and ban floating aliases in any regulated workflow; a model that can change underneath you cannot be evaluated. Second, set effort levels deliberately and benchmark "high" against "max" on your own long-horizon tasks before defaulting to the expensive dial, assume nothing. Third, turn on per-request cost attribution in Bedrock or Microsoft Foundry and attach a budget cap, so routing decisions are visible and reversible. Fourth, add the honesty column to your eval suite, so you are buying trustworthy output and not just confident output. For the regulatory shape underneath all of this, our field guide "The State of Mid-Market AI Compliance 2026" (/blog/state-of-mid-market-ai-compliance-2026) maps where these obligations land. The dial is set whether you set it or not, so set it, and write down why.