PortBench: A Correlation-Aware, Full-Pipeline Benchmark for LLM-Driven Portfolio Management
cs.AI, q-fin.PM
Submitted: 2026-05-27
Updated: 2026-09-12
Comments: Project page: https://portbench.github.io/
Project page: https://portbench.github.io
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models (LLMs) have shown strong performance across diverse financial tasks, yet portfolio management (PM) remains poorly benchmarked.
Terminology
Abstract
Large language models (LLMs) have shown strong performance across diverse financial tasks, yet portfolio management (PM) remains poorly benchmarked. Existing benchmarks exhibit two gaps: they are often equity-only and ignore cross-asset correlations; they fail to evaluate the complete PM decision pipeline. We introduce PortBench, a benchmark spanning six heterogeneous asset classes from 2015 to 2025. PortBench comprises a static QA dataset of 6,269 questions across seven task templates and a dynamic five-stage allocation pipeline. To evaluate these layers, we introduce two metrics: a dual-layer correlation score for inter-class hedging and intra-class concentration, and CEPS, which quantifies how reasoning errors compound across pipeline stages. We further evaluate under three stress windows and three risk profiles, and support real-time evaluation to mitigate pretraining contamination on historical markets. Across ten frontier LLMs, strong financial QA performance fails to translate into superior portfolio performance: only 32.5% of 120 evaluations beat equal weighting on Sharpe across four market periods. Our source code is available at this https URL.
Sources
- Standard Benchmarks Fail -- Auditing LLM Agents in Finance Must Prioritize Risk
- StockBench: Can LLM Agents Trade Stocks Profitably In Real-world Markets?
- GLM-5: from Vibe Coding to Agentic Engineering
- MASS: Muli-agent simulation scaling for portfolio construction
- FinanceBench: A New Benchmark for Financial Question Answering
- LLM-Powered Multi-Agent System for Automated Crypto Portfolio Management
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection