ShiJianBench: From Dialogue to Decision for Long-Horizon Evaluation of Investment Advisors
Jie Gong, Maowei Jiang, Zhiwei Liu, Yang Qiao, Wenxi Wu, Mengxi Xiao, Enze Zhang, Ziyan Kuang, Yankai Chen, Caishuang Huang, Meng Zhou, Xiku Du, Xue Liu, Guojun Xiong, Min Peng, Qianqian Xie, Sophia Ananiadou
cs.CL
Submitted: 2026-08-02
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- User Simulation in the Era of Generative AI: User Modeling, Synthetic Data Generation, and System Evaluation
- UserSimCRS v2: Simulation-Based Evaluation for Conversational Recommender Systems
- Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks
- StockBench: Can LLM Agents Trade Stocks Profitably In Real-world Markets?
- Doubly Robust Policy Evaluation and Learning
- A Survey on User Behavior Modeling in Recommender Systems
- FinanceBench: A New Benchmark for Financial Question Answering
- Session-based Recommendations with Recurrent Neural Networks
- Open-FinLLMs: Open Multimodal Large Language Models for Financial Applications
- RecSim: A Configurable Simulation Platform for Recommender Systems
- TradingGPT: Multi-Agent System with Layered Memory and Distinct Characters for Enhanced Financial Trading Performance
- Open FinLLM Leaderboard: Towards Financial AI Readiness
- AgentBench: Evaluating LLMs as Agents
- Fin-R1: A Large Language Model for Financial Reasoning through Reinforcement Learning
- RecSim NG: Toward Principled Uncertainty Modeling for Recommender Ecosystems
- Flipping the Dialogue: Training and Evaluating User Language Models
- RecoGym: A Reinforcement Learning Environment for the problem of Product Recommendation in Online Advertising
- Reliable LLM-based User Simulator for Task-Oriented Dialogue Systems
- Know You First and Be You Better: Modeling Human-Like User Simulators via Implicit Profiles
- BloombergGPT: A Large Language Model for Finance
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering