From Numbers to Judgment: Specialist LLM Agents and Reinforcement Learning for European Listed Real Estate
Pardis Taghavi, Santosh Bhavani
cs.AI, cs.LG
Submitted: 2026-08-11
Updated: 2026-08-13
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: The paper "From Numbers to Judgment: Specialist LLM Agents and Reinforcement Learning for European Listed Real Estate" studies whether the localized numerical operations and integrative judgments of
Terminology
Summary
The paper From Numbers to Judgment: Specialist LLM Agents and Reinforcement Learning for European Listed Real Estate
studies whether the localized numerical operations and integrative judgments of financial analysis benefit from the same form of LLM specialization.
The authors introduce Larix, which maps a 16-lens European listed-real-estate analysis framework to eight lens-aligned specialists. They compare a frontier LLM (Claude Opus 4.8) under monolithic versus specialist-decomposed prompting while holding the model, source evidence, task instructions, output schema, and scoring fixed. Across 19 firms spanning seven regulatory wrappers, decomposition improves the numerical-task aggregate by 15.8 percentage points (from 78.9% to 94.7%) but does not reliably improve, and can reduce, performance on judgment tasks (47.4% in both conditions, a change of 0.0 percentage points). This pattern is stable across four frozen-template dispatches. A single-agent control given the complete framework (Claude-Full) does not reproduce the numerical gain, scoring at exact parity with the generic monolith (66.3% in both conditions) but reducing T1 by 10.5 points while improving T2 by 10.5 points.
The paper then post-trains Qwen3.5-9B with GRPO using task-aligned structured rewards. This raises the development-split score by 12.0 points (from 72.2% to 84.2%) and the judgment aggregate by 14.2 points (from 58.6% to 72.8%), with gains on all four sub-ceiling tasks. The gains transfer to unseen firms (+15.2 points overall; +40.4 on covenant stress), to unseen regulatory wrappers (+4.3), and to later periods (+2.5), with positive transfer on all three anti-memorization splits.
The paper concludes that prompt-level decomposition improves modular numerical execution, whereas targeted parameter adaptation improves integrative financial judgment. The results support a task-dependent design principle: use decomposition to simplify localized, verifiable financial operations, and use targeted parameter adaptation when reliable performance depends on integrating multiple sources of financial evidence.
The evaluation universe contains 25 European listed-real-estate firms spanning eight legal and reporting wrappers. The primary same-model comparison uses a 19-firm continental cohort covering six French SIICs, one Austrian Immobilien-AG, three Belgian GVV/SIRs, four German AGs, one Italian SIIQ, two Spanish SOCIMIs, and two Dutch FBIs, producing 95 task–firm instances per condition across the five-task benchmark. Six Swiss AGs form the eighth wrapper and are used in the cross-wrapper and homogeneous-wrapper analyses.
The benchmark separates source-grounded numerical tasks (T1: regime-specific metric, T3: implied cap rate, T6: payout-regime classification) from judgment-intensive tasks (T2: adjustment identification, T5: covenant stress) that require reconciling multiple disclosures and regime-specific constraints. The numerical tasks T1 and T3 are closed-book probes with no filing injected, while the judgment tasks T2/T5 and the payout-regime task T6 inject company-specific primary-source extracts.
Ground truth combines 45 expert-blinded tuples with 50 author-curated continental tuples carrying per-cell provenance. The external expert received only the primary-source evidence and task definition and did not observe model outputs, specialist prompts, scorer implementation, or the 16-lens framework.
For the RL post-training, the corpus contains 195 rows derived from the benchmark ground truth plus five training-only counterfactual covenant-breach-positive rows, for 200 rows in total. After applying three disjoint evaluation splits (held-out-firm with 35 rows, held-out-wrapper with 25 rows, held-out-period with 45 rows), training uses 70 rows: 65 real rows plus the five counterfactual rows. The remaining 25 real rows form a development split. The authors train a LoRA adapter of rank 32 using GRPO implemented in veRL with vLLM rollouts on a single H100 GPU, with 64 prompts sampled per step and G = 8 candidate responses at temperature 1.0. They planned 90 update steps but stopped at step 30 for budget reasons; the step-20 checkpoint is the model evaluated throughout.
The reward is the deterministic task score used during evaluation, with no separate model-based judge. The rubric weights are unchanged from evaluation, with the boolean breach-classification term carrying weight β = 0.45. Reward-term telemetry over all 15,360 training rollouts shows no degenerate optimization of the β-weighted boolean term: the breach-positive bonus fires on only 5.5% of rollouts, ruling out an always predict breach
collapse.
The paper also reports two harness-hardening findings: field-request alignment (the evaluation prompt must explicitly request every structured field the rubric scores) and rubric sub-classification (rubrics for regime-dependent reasoning require sub-classification rather than coarse categoricals). A numeric-tolerance sensitivity analysis shows that the ±5% (T1) and ±15% (T3) numeric bars are the one load-bearing scoring choice on the numerical tasks.
The paper's four contributions are: (1) a controlled same-model evaluation of lens-aligned specialist decomposition for regime-aware financial analysis, holding source evidence, task instructions, output schemas, and scoring fixed and including a full-framework monolithic control; (2) identification of a task-dependent trade-off where decomposition improves modular numerical analysis but does not reliably improve, and can reduce, performance on tasks requiring integrative financial judgment; (3) development of a structured-reward GRPO post-training procedure showing that a post-trained Qwen3.5-9B improves the judgment tasks that prompt-level decomposition cannot, with gains that transfer to unseen firms and unseen regulatory wrappers; (4) release of a 25-firm, eight-wrapper benchmark of source-grounded numerical and judgment tasks with a corrected, per-cell-provenance rubric, a frozen repeat-dispatch panel, and a hardened scoring harness.
The paper notes several limitations: the evaluation is scoped to the specialist layer with 19 firms across seven wrappers, three lens-aligned specialists, and five tasks, with synthesis and position sizing outside the present design; per-wrapper estimates rest on few firms; the out-of-distribution transfer evidence comes from one preserved checkpoint; the RL corpus is compact (195 real dispatches plus five training-only counterfactual breach-positive examples); and the post-training comparison uses one frontier model family and one 9B open model.
Improvements for AI systems
Improvements to AI systems:
-
Task-adaptive reasoning router: Build a system that classifies each query as either
modular numerical
(e.g., metric extraction, cap-rate calculation) orintegrative judgment
(e.g., covenant stress, adjustment identification) and then automatically selects between (a) decomposing the task into lens-aligned specialist sub-agents or (b) using a single post-trained model with RL-tuned judgment. This prevents the observed 15.8-point numerical gain from being offset by judgment degradation. -
Structured-reward GRPO post-training for judgment tasks: Apply the paper's reward design—using deterministic, rubric-weighted scores with a boolean term (β=0.45) and telemetry to detect degenerate policies—to fine-tune small open models (e.g., 9B) on domain-specific integrative tasks. The improved system can lift judgment accuracy by 14.2 points without reward hacking, as verified by the 5.5% breach-positive firing rate.
-
Regime-aware prompt hardening: Implement two harness fixes from the paper: (a) explicitly request every structured field the rubric scores (field-request alignment), and (b) use rubric sub-classification for regime-dependent reasoning instead of coarse categoricals. This yields more reliable, reproducible outputs across legal wrappers (French SIICs, German AGs, Dutch FBIs, etc.).
-
Anti-memorization transfer validation: Train with disjoint splits (held-out firm, wrapper, period) and report per-split gains. The improved system can demonstrate positive transfer (+15.2 points on unseen firms, +4.3 on unseen wrappers, +2.5 on later periods), proving it learns generalizable financial reasoning rather than memorizing training instances.
-
Numeric-tolerance sensitivity checking: Before deployment, run a sensitivity analysis on scoring tolerances (±5% for T1, ±15% for T3). The improved system can flag when performance is overly dependent on a single tolerance bar, allowing users to adjust confidence thresholds for production use.
What the improved AI system can do:
-
Given a European listed-real-estate firm’s filings, it can automatically decide whether to spawn eight lens-aligned specialists for numerical tasks (achieving 94.7% accuracy on regime-specific metrics and cap rates) or use a single RL-tuned model for judgment tasks (achieving 72.8% on covenant stress and adjustment identification).
-
It can process firms across eight regulatory wrappers (France, Austria, Belgium, Germany, Italy, Spain, Netherlands, Switzerland) with per-cell provenance, producing auditable outputs with explicit field-level justifications.
-
It can be fine-tuned on a small corpus (70 training rows) with GRPO on a single H100 GPU, reaching step-20 checkpoint performance that transfers to unseen firms and wrappers, enabling cost-effective domain adaptation for specialized financial analysis.
-
It can detect and avoid reward hacking (e.g., not collapsing to
always predict breach
) via rollout telemetry, ensuring trustworthy outputs in high-stakes covenant stress testing.
Abstract
We study whether the localized numerical operations and integrative judgments of financial analysis benefit from the same form of LLM specialization. Larix maps a 16-lens European listed-real-estate analysis framework to eight lens-aligned specialists; we compare a frontier LLM under monolithic versus specialist-decomposed prompting while holding the model, source evidence, task instructions, output schema, and scoring fixed. Across 19 firms spanning seven regulatory wrappers, decomposition improves the numerical-task aggregate by 15.8 percentage points but does not reliably improve, and can reduce, performance on judgment tasks, a pattern stable across four frozen-template dispatches; a single-agent control given the complete framework does not reproduce the numerical gain. Post-training Qwen3.5-9B with GRPO using task-aligned structured rewards then raises the development-split score by 12.0 points and the judgment aggregate by 14.2 points, with gains on all four sub-ceiling tasks; the gains transfer to unseen firms (+15.2 points overall; +40.4 on covenant stress) and to unseen regulatory wrappers (+4.3), with positive transfer on all three anti-memorization splits. Prompt-level decomposition thus improves modular numerical execution, whereas targeted parameter adaptation improves integrative financial judgment.
Sources
- Large Language Model based Multi-Agents: A Survey of Progress and Challenges
- FinanceBench: A New Benchmark for Financial Question Answering
- Spec Kit Agents: Context-Grounded Agentic Workflows
- TradingAgents: Multi-Agents LLM Financial Trading Framework
- FinCon: A Synthesized LLM Multi-Agent System with Conceptual Verbal Reinforcement for Enhanced Financial Decision Making
- AlphaAgents: Large Language Model based Multi-Agents for Equity Portfolio Constructions
- Group Sequence Policy Optimization
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection