The AI Billing Shock: Why Soaring Token Costs Are Challenging Mortgage AI Assumptions

Camillo Melchiorre is President and Director of Regulatory Compliance with IndiSoft LLC, Columbia, Md. Contact him if you’re interested in the Mortgage AI Vendor Screening Questionnaire mentioned below.

This editorial is part of a series.

For the last couple of years, technology vendors have promised that artificial intelligence is getting steadily cheaper and that token prices are racing toward zero. For basic, entry-level models, that claim is partly true. Inside mortgage company finance departments, however, the experience is very different. Enterprise AI budgets are not shrinking—they are being overrun.

Camillo T. Melchiorre

Recent enterprise FinOps analyses, including work associated with McKinsey, show that a large majority of companies have already exceeded their AI budgets. High-profile examples reinforce the pattern: Uber reportedly exhausted its entire 2026 AI allocation in the first four months of the year and responded by imposing strict internal usage caps.

The reason is straightforward. While commodity model prices have declined, the premium “reasoning” models that enterprises actually need for complex work have moved in the opposite direction, with token costs rising sharply in many cases. The end of heavy venture-capital subsidies, persistent data-center power constraints, and explicit charges for internal model “thinking time” have made advanced intelligence far more expensive than early projections assumed. When these higher unit prices meet the high token volumes generated by autonomous agents, the result is a double cost shock. Industry analyses indicate that a substantial portion—often a majority—of agentic AI spend is consumed by response refinement and self-correction rather than primary task execution.

The Cost Pressure in Mortgage Origination

According to Mortgage Bankers Association data, the average cost to originate a retail mortgage remains elevated—approximately $11,094 per loan in recent full-year figures, with per-loan production expenses rising further into the $11,800–$12,000 range in early 2026 for many independent mortgage bankers. Lenders have turned to AI in an effort to compress that cost through automated underwriting and income validation. Reported pilot results have shown meaningful operational gains, including substantial reductions in manual processing time and increases in processor capacity.

Those gains, however, are frequently erased by unmanaged token consumption.

Consider a typical agentic workflow used to validate income for a self-employed or Non-QM borrower. The agent does not perform a single static OCR pass. It iterates through hundreds of pages of bank statements, tax returns, and supporting documents, cross-checking data points, flagging inconsistencies, and applying credit-policy rules. Premium reasoning models charge for the internal tokens required to perform these multi-step logical operations. A single underwriting run can consume tens of thousands of tokens. Because borrower data must also travel through encrypted, compliant pathways, every token carries a privacy premium. A large share of the resulting bill is often attributable to the model auditing its own output, correcting parsing errors on imperfect scans, and rewriting summaries before a human underwriter ever sees the file.

Parallels from Healthcare and Corporate Tax

The same pattern has already appeared in healthcare prior-authorization reviews. Insurers deployed agents to examine large volumes of unstructured clinical records against coverage guidelines. A significant share of the resulting compute spend went to self-correction when the models encountered low-quality or inconsistent source documents, generating multi-million-dollar API invoices in some reported cases.

Corporate tax and legal audit practices encountered a related problem—the “context-window trap.” When firms load multi-year document archives into premium models to locate specific transaction trails, the high unit cost of input tokens, applied repeatedly across large static contexts, produces billing spikes that can erase the expected margin on the engagement. Mortgage servicing faces an analogous exposure. Servicers seeking meaningful reductions in per-loan servicing costs are feeding multi-year payment histories, escrow records, and call notes into agentic systems. Regulatory requirements from the CFPB, the GSEs, and state regulators limit the ability to truncate those files, so the full context must often be processed. The resulting token volumes can overwhelm the projected operational savings.

Runaway Logic Loops

A further hidden cost is the tendency of autonomous agents to enter extended reasoning loops when confronted with messy or contradictory data. In commercial credit applications, agents given open-ended instructions to optimize credit limits have been observed consuming large volumes of premium reasoning tokens while repeatedly revising internal models before timing out. The same behavior appears when mortgage agents attempt to reconcile poorly scanned bank statements.

Practical Guardrails for Mortgage Lenders and Servicers

Mortgage organizations can no longer treat AI as an unbounded utility. Operational savings disappear when token costs exceed the labor they replace. Four architectural disciplines are essential:

1. Model cascade. Route routine tasks—initial document classification, basic OCR cleanup, and standard borrower inquiries—to low-cost utility models. Reserve expensive frontier reasoning models for high-stakes underwriting logic, Non-QM income analysis, and compliance-critical determinations.

2. Selective escalation. Design workflows so that only data requiring complex reasoning or regulatory verification is passed to premium models.

3. Hard-coded loop limits. Do not rely solely on prompt instructions. Embed strict iteration caps in the surrounding software. If an agent cannot reconcile asset documentation within a fixed number of cycles (for example, three), the system should terminate the loop and escalate to a human.

4. Data preparation and prompt caching. Clean and compress documents with conventional OCR tools before they reach a premium API. Cache static reference material such as underwriting guidelines so the model does not re-ingest it on every call.

Conclusion

The era of assuming that AI costs will reliably decline is over for the models that perform real mortgage work. Token prices for advanced reasoning systems have risen, agentic workloads amplify consumption, and unmanaged loops and context volumes can erase projected efficiency gains. Mortgage lenders and servicers that treat AI cost control as a core architectural and vendor-management discipline—rather than an afterthought—will be better positioned to capture lasting operational value.

Those that do not will continue to experience the billing shock already visible across the industry.

For anyone interested, I created, with the assistance of several colleagues from the Johns Hopkins University AI Business Strategy Masters Certificate Program, a Mortgage AI Vendor Screening Questionnaire that I will readily share if you message me with your contact information.

Sources

1. Fortune, Forbes, and related reporting on Uber’s exhaustion of its 2026 AI budget within the first four months of the year (May 2026). 

2. Mortgage Bankers Association production cost data (recent releases showing per-loan expenses in the $11,000–$12,000 range for independent mortgage bankers). 

3. Enterprise FinOps analyses, including McKinsey-related surveys on AI budget overruns (2025–2026). 

4. Industry case studies and pilot reports on AI-driven mortgage process efficiency gains. 

5. Parallel observations from healthcare prior-authorization and corporate audit AI deployments (industry analyses, 2025–2026). 

(Views expressed in this article do not necessarily reflect policies of the Mortgage Bankers Association, nor do they connote an MBA endorsement of a specific company, product or service. MBA NewsLink welcomes submissions from member firms. Inquiries can be sent to Editor Michael Tucker or Editorial Manager Anneliese Mahoney.)