MortarBench: Evaluating Mortgage Loan Origination Agents

arXiv:2606.19416 · cs.LG · Submitted 2026-06-17 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "MortarBench: Evaluating Mortgage Loan Origination Agents".

Jane: The paper was written by Fengbin Zhu, Wenqiang Lei, Youcheng Huang, Chao Wang, Shuo Zhang et al. from Association for Computational Linguistics.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Summary: Tom: So, we finished discussing what "MortarBench: Evaluating Mortgage Loan Origination Agents" is designed to test—the entire lifecycle of a mortgage application. Jane, can you walk our listeners through what the paper actually found when they ran these agents through the system?

Jane: The summary really zeroes in on performance metrics. They didn't just check if the agent succeeded or failed; they measured *how* and *why* it struggled in specific parts of the process.

Meng: And that’s key, because a failure isn't just a failure; it might be a failure due to inadequate data extraction from an old document format, or maybe misunderstanding a state-specific regulation.

Lu: I remember reading that they segmented the tasks really granularly—like separating the identity verification steps from the income verification steps, which allows for pinpointing weaknesses.

Tom: So, it's not just a single score? It’s like they gave it grades on different skills, right?

Jane: Exactly. They found that while agents were excellent at basic data intake—like pulling names and addresses—they struggled significantly more when the process required cross-referencing multiple, disparate sources of information.

Lalam: That points to a gap in reasoning capabilities. The AI can read the pieces, but it doesn't seem to be connecting them into a cohesive narrative that satisfies human compliance officers.

Meng: From an engineering standpoint, that multi-source cross-referencing failure is where the real money and risk are exposed. It requires not just reading, but synthesizing contradictory or incomplete information.

Tom: Lu, you mentioned granularity—were there specific areas of the mortgage process they found were surprisingly difficult for these agents to handle?

Lu: Yes, particularly when it came to assessing non-traditional income sources. If a person's income comes from side hustles or freelance gigs that aren't standard W2 forms, the agents seemed to get confused by the lack of a clean, predictable pattern.

Jane: It makes sense because those non-

Paper discussion segment 2: Tom: So, if we can sum up what MortarBench showed us, it’s that evaluating AI agents on real-world mortgage tasks is incredibly complicated and necessary for the industry to trust these tools.

Jane: Exactly! Think about it this way: lending isn't just one question; it involves checking bank statements against property deeds and loan terms all at once, which is a huge hurdle for any new AI system to clear.

Tom: But Jane, you mentioned "huge hurdle"—that’s what struck me; the sheer volume of disparate data types they had to stitch together is staggering, making simple QA models totally inadequate for this kind of process.

Meng: And from an engineering standpoint, that multi-step dependency chain is where most systems fail in testing; you can't just test one piece—you have to test the entire operational pipeline under stress.

Lu: You hit on the core problem, Meng; it’s not just about accuracy on a single data point, it’s about maintaining contextual integrity across multiple documents, which is a fundamental leap in reasoning capability we haven't seen before.

Lalam: Considering that complexity, I see this breakthrough as enabling a massive cultural shift where bureaucratic friction—all those paper trails and manual verifications—can finally start to fade away.

Jane: So it's not just making loans faster; it’s about restoring a degree of simplicity to what used to feel overwhelming for people trying to buy a home.

Tom: It really feels like the next generation of AI needs to be less like a search engine and more like an actual paralegal who reads everything and knows which piece connects to which other piece.

Meng: If we could build agents that reliably handle that level of document synthesis, my startup's immediate focus would shift entirely toward compliance auditing using this framework.

Lu: But I wonder if the underlying principle—the ability to map relationships across diverse data schemas—could be applied far beyond just finance, perhaps to complex scientific literature review?

Lalam: That's a powerful thought, Lu; improving how we synthesize knowledge across specialized fields could change research culture forever, making breakthroughs happen faster than ever before.

Jane: Honestly, it makes you wonder what other highly regulated industries—like medicine or insurance—are waiting for this level of cross-document comprehension to finally modernize their processes.

Tom: Knowing that these agents are so powerful, I’m really curious about how they handle ambiguity when the source documents themselves contradict each other; that's a whole different beast than just finding a match.

Paper discussion segment 3: Jane: Well, think about it like this; most testing materials give you perfectly labeled forms, but in reality, you get faded ink here, a handwritten note there, and maybe three different types of statements all mixed together.

Tom: Exactly! It’s the difference between a textbook problem and finding that one sticky note tucked into the mortgage packet from five years ago.

Lu: What excites me about this push for resilience is that if we can make an AI understand the *intent* behind multiple, conflicting data points—say, a payment date on one form versus a statement date on another—we aren't just building better loan processors; we're building fundamental understanding engines.

Meng: But Lu, understanding intent is where the engineering nightmare begins. When you say "conflicting data," are we talking about simple contradictions, or are we talking about ambiguities that require human judgment, like knowing which signature belongs to the primary guarantor?

Lalam: That ambiguity is precisely where the biggest cultural shift lies. If an AI can reliably flag those points of conflict and explain *why* it's unsure—instead of just guessing—it builds trust in the financial system itself.

Tom: Right, Meng brought up judgment, and Lalam talked about trust; so it sounds like the improvement isn't just accuracy, but transparency about uncertainty.

Jane: So instead of a black box saying "yes" or "no," the system has to say, "Based on this and that, I think yes, but please check the handwritten date here." Does that make sense?

Lu: And if we can codify what a human expert looks for when they see those discrepancies—the pattern of doubt—we’re moving toward an entirely new kind of knowledge graph for finance.

Meng: From an engineering standpoint, that means the model needs to generate not just a list of facts, but a weighted map of plausible explanations based on historical exceptions. That's a huge jump in complexity.

Lalam: It means we could democratize expert knowledge, giving every small lender the analytical power previously only available to massive institutions with teams of senior underwriters.

Tom: Wow, so the potential impact is making specialized expertise accessible everywhere, which is incredible! Given this focus on robust judgment calls, I wonder what happens when we apply this same level of contextual rigor to other high-stakes areas besides mortgages?

Conclusion: Tom: So, wrap-up time—it’s hard to believe we’ve covered everything we have on "MortarBench: Evaluating Mortgage Loan Origination Agents."

Jane: What really sticks with me is how much better this framework is going to make these complicated processes, making them less of a black box and more transparent.

Tom: Exactly! It shows that simply evaluating the performance of AI agents in such a critical financial domain isn't just academic; it’s absolutely necessary for consumer safety.

Lu: You know, thinking about the implications, this isn't just about mortgage applications—it opens up these wild possibilities for every complex, high-stakes automated decision system out there.

Jane: That’s a huge leap from loan origination, Lu; you think it applies to things like insurance underwriting or maybe even medical triage?

Lu: Absolutely! If you can benchmark the reasoning and reliability of an agent in one specific area, that methodology can be scaled to prove capability in almost any field requiring careful judgment.

Meng: I agree with the scalability point, but we have to talk about implementation difficulty; these real-world financial systems are incredibly messy and siloed right now.

Jane: So, does the benchmarking itself solve the problem of integrating those agents into legacy banking infrastructure?

Meng: Not automatically, no. The paper provides the *metrics* for performance, but connecting those ideal agent outputs to decades-old core banking systems requires a massive engineering lift and careful phased roll-out.

Lalam: That speaks directly to culture change, doesn't it? People are comfortable with the known complexity of these systems, so accepting a new AI layer requires not just trust in the tech, but trust in the *process* that technology represents.

Tom: So it’s a shift in institutional confidence as much as it is an improvement in capability.

Jane: And that means we have to guide users and stakeholders through understanding what "benchmarking" actually proves about reliability.

Lu: It suggests that future AI deployments can move beyond simply *doing* tasks and start proving they can *reason* robustly under pressure, which is a huge leap for the field overall.

Meng: And if we're being practical, any successful implementation will need to account for geographical variations in lending laws, not just the core technical performance.

Lalam: Ultimately, advances like those shown in "MortarBench: Evaluating Mortgage Loan Origination Agents" don't just improve efficiency; they raise the baseline expectation of fairness and reliability for AI across society.

Tom: Wow, what a wrap-up—we really appreciate you all joining us today to break down this groundbreaking research.

Jane: Thanks so much to everyone! We’re going to take a quick break and when we come back, we'll be shifting gears entirely and looking at some exciting new developments in multimodal AI.

Fengbin Zhu, Wenqiang Lei, Youcheng Huang, Chao Wang, Shuo Zhang, Jiancheng Lv, Fuli Feng, Tat-Seng Chua

Association for Computational Linguistics

cs.LG

Submitted: 2026-06-17

Updated: 2026-08-25

Code: https://github.com/mtoles/MortarBench

Importance score: 89/100

The gist: The accurate and automated processing of mortgage loan applications is critical for the financial industry, requiring robust methods to extract structured information from highly heterogeneous

Key concepts

MortarBench
A framework designed to test AI agents on the entire mortgage application process. It evaluates performance metrics beyond simple success/failure by measuring *how* and *why* agents struggle in specific parts of lending.
Cross-referencing disparate sources
The ability of an AI to connect information from multiple, varied documents (like bank statements and property deeds) that do not contain the same data. This requires synthesizing contradictory or incomplete information.
Non-traditional income sources
Income earned through side hustles or freelance gigs that do not come from standard W2 forms. The episode notes AI agents struggle with these types of income due to the lack of a clean, predictable pattern.

Terminology

Summary

The accurate and automated processing of mortgage loan applications is critical for the financial industry, requiring robust methods to extract structured information from highly heterogeneous sources. This paper details a comprehensive framework for evaluating agents designed to handle the complexity of loan origination data, which spans identity verification, asset assessment, property valuation, and detailed financial transaction analysis. The ability to reliably interpret and synthesize data from disparate formats—such as bank statements and standardized application forms—is paramount for minimizing errors that can cost millions of dollars.

Data Standardization in Loan Origination

The process relies heavily on standardized data schemas to ensure machine readability across the industry. One key structure is the Uniform Loan Application Dataset (ULAD), which is a standardized, machine-readable data format used in the US mortgage industry. ULAD captures a vast array of information, including borrower-level details (identity, demographics, income), financial standing (assets such as bank accounts and liabilities such as outstanding debts), and property characteristics. Furthermore, transactional records are captured in detailed JSON formats. For instance, a sample bank statement includes structured accounts with specific identifiers (e.g., account: "54712641") and associated transactions, which track details like the description, amount, and date of activity.

Modeling Complex Financial Interactions

The scope of financial data extends beyond traditional banking records to encompass various payment mechanisms. The analysis considers multiple categories of money transfer services, including both US-based options (Zelle, Venmo) and extensive non-US alternatives (Alipay, M-Pesa, Wise). This breadth highlights the necessity for models to process global financial contexts. Additionally, the framework incorporates specialized datasets that model specific financial events; for example, a dataset may track Loan Proceeds or Savings Club deposits alongside other transactions.

Prompt Engineering for Structured Extraction

A core component of the evaluation is the methodology used to guide AI agents through data extraction, detailed in prompt templates. These prompts are designed to move from general understanding to highly specific structured output. The prompting strategy includes several distinct stages:

  1. Initial Scaffold: This shared stage provides a holistic view, presenting the question alongside multiple data sources (e.g., Bank Statement: bank statement , ULAD DU: ulad du ) before requesting an answer based on an answer instruction.

  2. Targeted Extraction: Specific prompts are used to extract structured answers, such as determining if a statement is Boolean or listing relevant accounts and transactions. For example, one prompt instructs the model to Identify the relevant accounts in text (names/descriptions/last4 digits); do not guess or output account IDs.

  3. Confidence Scoring: A sophisticated prompt requires agents to perform detailed reasoning, asking them to List every transaction in the bank statement that is plausibly related... For each one, assign an integer confidence rating from 1 to 5. This forces the model to articulate its assumptions and provide a JSON list of "transaction id": "","confidence":.

Agent Performance Evaluation

The evaluation framework rigorously tests agents across multiple dimensions, distinguishing between model stages (which produce free-form answers) and clean stages (which extract structured final answers). The system measures performance on several key tasks: determining if a transaction list is relevant (Txn list (model)), identifying the specific account type (Account list (model)), and extracting precise dollar amounts. The use of shared prompts ensures that agents are evaluated consistently, while specialized prompts allow for testing nuanced skills, such as filtering transactions based on a confidence threshold T after a cleanup pass.

Improvements for AI systems

I. Architectural Improvements: Transitioning from Prompt Templates to a Multi-Modal Inference Pipeline

The current system relies on highly specialized prompt templates for each extraction task (e.g., Txn list (clean), Account (clean)). The improvement is to abstract these specific instructions into a generalized, modular inference pipeline managed by an orchestrator agent.

  • Improvement: Implement a Financial Reasoning Chain-of-Thought (FinCoT) Module. Instead of asking the LLM to jump directly to the answer type (Boolean, Dollar, List), the system must first force the model through mandatory intermediate reasoning steps that mimic human underwriting thought processes.

  • Capability: The improved AI system can perform Decompositional Reasoning. If asked, Did John use his savings account balance to fund this loan?, the FinCoT module forces the model to generate explicit sub-steps: 1) Identify Loan Source (ULAD field); 2) Identify Funding Account (Bank Statement Name/Type); 3) Compare Amounts (ULAD amount vs. Bank Statement available funds). This structured breakdown drastically reduces hallucination and increases traceability, providing a human-readable audit trail for every conclusion.

II. Data Integration Improvements: Building a Unified Financial Knowledge Graph

The current setup treats the Bank Statement, ULAD, and Text as separate inputs to be queried sequentially. The critical weakness is the lack of explicit relationship mapping across these distinct data types.

  • Improvement: Integrate a Graph Neural Network (GNN) Layer upstream of the LLM prompt generation. This layer processes all input documents (Bank Statement + ULAD + Text) not as separate JSON/Text blocks, but as nodes and potential edges in a single, comprehensive Knowledge Graph.

  • Capability: The improved AI system gains Cross-Domain Relational Inference. It can answer questions that require linking concepts across datasets that were never explicitly linked in the source documents.

  • Example: Given a loan application (ULAD) mentioning Investment Property and a bank statement showing unusual transfers, the system can infer: The transfers noted on Date X likely relate to the collateral valuation for Property Y mentioned in the ULAD, even if no direct text connection exists.

III. Extraction Robustness Improvements: Conflict Detection and Reconciliation

The current confidence scoring (1-5) is subjective and relies on the LLM's internal plausibility assessment. Financial data requires absolute truth verification against multiple sources of record.

  • Improvement: Implement a Triangulation Verification Module. This module treats all extracted facts (amounts, dates, names) as hypotheses that must be validated by at least two independent sources within the input set (Bank Statement to ULAD, Text to Bank Statement).

  • Capability: The system achieves Conflict Resolution and Source Attribution. When a discrepancy is found (e.g., Text implies an amount of 10,000, but the ULAD shows 9,950), the system does not simply choose one answer. Instead, it outputs a structured conflict report:

  • Fact: Loan Amount

  • Conflict: Yes

  • Source A (Text): 10,000 (Reason: User statement)

  • Source B (ULAD): 9,950 (Reason: Official document record)

This prevents the costly error of hallucinating a single correct answer when source data is contradictory.

IV. Efficiency Improvement: Stochastic Prompt Refinement

The current prompt system uses fixed placeholders (question, bank statement). This limits adaptability when new financial instrument types or complex legal jargon are encountered.

  • Improvement: Introduce Stochastic Prompt Generation (SPG). Before executing the final extraction prompt, the system dynamically analyzes the input documents' vocabulary (e.g., identifying terms like amortization schedule, escrow, or specific bond types) and automatically injects relevant domain-specific boilerplate knowledge into the prompt context.

  • Capability: The improved AI system exhibits Zero-Shot Domain Adaptation. It can process novel, unseen financial documents (e.g., a specialized commercial real estate mortgage document) without requiring the development of a new, specific prompt template for that document type.

Sources

Related papers