FinEvolveBench: A Benchmark for Self-Evolving Agents on Low-Repetition Tasks with Implicit Rewards
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "FinEvolveBench: A Benchmark for Self-Evolving Agents on Low-Repetition Tasks with Implicit Rewards".
Jane: Experience-based self-evolution enables language-model agents to improve their behavior by accumulating and updating experience at test time, yet existing evaluations often assume recurring task patterns and explicit success signals.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: The title itself, "FinEvolveBench: A Benchmark for Self-Evolving Agents on Low-Repetition Tasks with Implicit Rewards," really tells us what’s inside; it’s about testing agents that can evolve themselves using feedback when the tasks aren't repetitive.
Jane: And the authors are a team including Zihao Deng, Yining Zhu, Leiming Wang, Jingfei Lu, Junbo Wang, Chuncheng Ran Yu Yang, Dixuan Yang Jikun Shen; it shows a really diverse group tackling this complex problem of how agents learn in non-standard settings.
Lu: Jingfei Lu’s work on this seems particularly relevant because it deals with the mechanism of experience-based adaptation, which is a core concept we need to understand for future AI development.
Meng: From an engineering standpoint, I'm curious how they managed to build a benchmark that simulates such a complex financial stream without making the setup too brittle or overly reliant on one specific data source.
Lalam: The implication here is that we’re moving toward AI systems that don't just react; they actively refine their own strategies based on delayed, contextual feedback, which could make them much more robust in unpredictable environments.
The paper's summary: Tom: So, the core idea of this paper is introducing FINEVOLVEBENCH, which is this benchmark designed specifically for self-evolving agents on low-repetition tasks where rewards are implicit and feedback can be noisy or delayed.
Jane: Essentially, they reconstruct a daily financial information stream covering thirty-one Chinese A-share industry indices and align that with one hundred seventy-seven thousand three hundred twenty-four public news articles to create a rich testing environment for the agents.
Lu: The paper sets up prediction horizons of ten, twenty, and forty trading days over this stream to see if the agents can successfully convert that noisy real-world feedback into usable experience at the exact moment of testing.
Meng: I’m thinking about those prediction horizons; setting them too long or too short could completely change how an agent needs to manage its memory and decision-making process in a live system.
Lalam: The paper describes a method-agnostic replay workflow that ensures every system sees the same date-bounded information and only gets feedback after the relevant market outcome has actually occurred, which is crucial for testing real-time adaptation.
The paper's improvements: Tom: Now, looking at what they suggest as improvements or what they found about different systems, it points out that experience availability alone isn't enough for predictive gains; append-only memory and utility-updated memory didn’t always beat the no-experience pipeline consistently.
Jane: That’s a fair point; the controlled ablation showed that feedback-driven utility updating actually helped in some backbone and horizon settings but could actually hurt in others, which means the mechanism isn't universally perfect.
Lu: They highlight that successful experience reuse requires more than just repeating a fixed prediction rule because superficial similarities in news can lead to different market responses, so agents need a deeper way to learn from those variations.
Meng: It’s interesting how they suggest that the usefulness of experience changes depending on the conditions and across different replay stages; it means an agent's memory needs to be context-aware about when its past knowledge is actually relevant.
Lalam: This leads to the idea that an improved system should be able to dynamically adjust its utility score based on how mature that delayed market outcome is, allowing it to ignore irrelevant experience and focus on what actually matters for the next decision.
Conclusion: Tom: To wrap things up, this work with "FinEvolveBench" shows us that agents can indeed learn from real-world outcomes at test time, but we need sophisticated mechanisms like feedback-driven utility updating to make that learning stick in complex environments.
Jane: So, the big implication is that future AI systems need to be designed not just for pattern recognition, but for continuous, context-aware adaptation based on delayed signals rather than static success criteria.
Lu: I think the future potential lies in creating agents capable of navigating highly unpredictable domains where prior experience only becomes valuable after a significant delay or a change in market conditions occurs.
Meng: For practical application, this suggests that when we build agents for things like dynamic resource allocation, they need to be programmed to understand the temporal value of feedback signals over long horizons.
Lalam: This research helps pave the way for AI that can truly integrate real-world consequences into its learning loop, making them more resilient and adaptive in a constantly shifting landscape.
Zihao Deng, Yining Zhu, Leiming Wang, Jingfei Lu, Junbo Wang, Chuncheng Ran Yu, Dixuan Yang, Jikun Shen
cs.CL
Submitted: 2026-06-05
Updated: 2026-09-28
Importance score: 83/100
The gist: Experience-based self-evolution enables language-model agents to improve their behavior by accumulating and updating experience at test time, yet existing evaluations often assume recurring task
Key concepts
- FINEVOLVEBENCH
- A benchmark designed for self-evolving agents that simulates real-world financial forecasting. It uses 31 Chinese A-share indices and news articles to create a stream of data where rewards are implicit, noisy, and delayed, forcing agents to learn from sparse interactions.
- Replay Workflow
- A standardized four-step protocol used in the benchmark to ensure fair evaluation. It dictates when systems mature their knowledge (update), observe current market data, retrieve past memory for prediction, and then append the new trajectory for future use.
- Market-Adjusted Return (%i,t,h)
- A specific metric calculated by subtracting a broad market index return from an industry's return. This serves as the delayed feedback or evaluator outcome for predictions, helping agents learn how individual sectors perform relative to the overall market.
Terminology
Summary
Experience-based self-evolution enables language-model agents to improve their behavior by accumulating and updating experience at test time, yet existing evaluations often assume recurring task patterns and explicit success signals.
How it works
The paper introduces FINEVOLVEBENCH, a benchmark designed for evaluating self-evolving agents on low-repetition tasks characterized by implicit rewards, noisy feedback, and delayed outcomes. The benchmark reconstructs a daily financial information stream over 31 Chinese A-share industry indices and aligns 177,324 public news articles with market observations. Researchers define prediction horizons (e.g., 10-, 20-, and 40-day trading days) over this stream to test whether agents can convert noisy real-world feedback into reusable experience at test time.
The benchmark employs a method-agnostic replay workflow that ensures every system observes the same date-bounded information and receives feedback only after the relevant outcome becomes observable. The chronological replay protocol standardizes four steps:
-
Mature and update:
after the close, reveal outcomes of earlier predictions whose horizons end on the current trading day.
Eligible systems update onlythe corresponding horizon-specific utility before the current prediction.
-
Observe:
at the 24:00 cutoff, expose the current close, date-bounded news, and all earlier causally available information.
-
Retrieve and predict:
optionally retrieve memory that existed before the current prediction, then produce and freeze the score for each selected horizon.
-
Append and queue:
append the completed prediction trajectory to eligible memory stores
only after prediction, making the new trajectory retrievable on the next trading day.
Benchmark Construction and Data
FINEVOLVEBENCH is built around a specific financial context to simulate real-world decision problems. It includes:
-
Market Universe:
31 Shenwan Hongyuan first-level industry indices.
-
News Corpus:
177,324 time-stamped articles from nine public financial-news sources.
-
Temporal Span: The replay contains
300 prediction days from January 2, 2025 through March 31, 2026,
with market observations extending through June 30, 2026 to compute pending outcomes without exposing them to the agent.
The data construction involves aligning news and market records into trading date batches. A key adjustment is separating an industry’s relative response from broad market movement by calculating:
%i,t,h = r i t,h − r CSI1000 t,h.
This market-adjusted return
serves as the delayed feedback and as the evaluator-side outcome for the prediction score si,t,h.
Experimental Setup and Evaluation Metrics
The experiments test three backbone models: DEEPSEEK-V4-FLASH and QWEN3-35B-A3B. The evaluation focuses on three benchmark questions (RQ1–RQ3):
** RQ1: How do agents with no external experience, append-only memory, and feedbackupdated memory perform on FINEVOLVEBENCH?
**
** RQ2: Are the effects of experience mechanisms stable across backbone models, prediction horizons, and time-series versus cross-sectional evaluation?
**
** RQ3: How do no-experience, append-onlymemory, and utility-updated systems behave across Cold-Start and Exploitation... what additional effect is attributable to feedback-driven utility updating?
**
Performance is measured using two complementary information coefficients:
-
Time-series IC (tsIC): Measures temporal prediction for each target:
Corrt(si,t,h, αi,t,h).
-
Cross-sectional IC (csIC): Measures same-day ranking across targets:
Corri(si,t,h, αi,t,h).
The evaluation period is divided into two stages:
** Dcold = the first one-third of evaluation dates. This measures behavior while experience is limited or being calibrated.**
** Dexploit = the remaining two-thirds of evaluation dates. This measures behavior after a longer history of interactions and matured feedback is available.**
Key Findings
The results yield several diagnostic observations regarding the effectiveness of memory mechanisms:
-
Experience availability is not sufficient for predictive gains:
Append-only Mem0 and utility-updated MemRL do not consistently improve over Pipe.
-
The usefulness of experience changes across conditions and over the replay:
Experience-based systems sometimes improve relative to Pipe during Exploitation, consistent with accumulated experience becoming useful, but these gains do not persist across backbones, horizons, or metrics.
Improvements for AI systems
Based on the FINEVOLVEBENCH benchmark and its findings, here are specific improvements for an existing language model agent:
-
Improve agents by implementing a feedback-driven utility update mechanism that is conditional on delayed, outcome-level feedback.
-
The improved system should be able to distinguish between
useful
andirrelevant/harmful
experience by dynamically adjusting its utility score based on the maturity of the delayed market outcome. -
The system should be able to adapt its prediction strategy across different backbones (e.g., DeepSeek-V4-FLASH vs. QWEN3-35B-A3B) and different prediction horizons (10, 20, 40 days).
-
The improved agent can perform better during the
Exploitation
period by leveraging accumulated experience to improve predictions relative to a baseline (like the no-experience pipeline). -
If it is operating in a Cold-Start phase, the system should focus on learning from the current context rather than relying heavily on prior, potentially irrelevant, experience.
This improved AI system can:
-
Accurately predict market sentiment factors for various industry indices (31 A-share indices) over 10-, 20-, and 40-day horizons.
-
Convert noisy, delayed market outcomes into reusable experience at test time, allowing it to refine its decision-making process sequentially through the replay stream.
-
Demonstrate superior performance on complex tasks where surface recurrence is misleading (i.e., when similar news leads to different market responses).
-
Be robust across different underlying language model architectures by utilizing a mechanism that dynamically adjusts its learning strategy based on the current task context and feedback signal maturity.
Sources
- FinBERT: Financial Sentiment Analysis with Pre-trained Language Models
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- xbench: Tracking Agents Productivity Scaling with Profession-Aligned Real-World Evaluations
- StockBench: Can LLM Agents Trade Stocks Profitably In Real-world Markets?
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- Can ChatGPT Forecast Stock Price Movements? Return Predictability and Large Language Models
- GAIA: a benchmark for General AI Assistants
- Humanity's Last Exam
- ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
- Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory
- FinGPT: Open-Source Financial Large Language Models
- MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory
- LifelongAgentBench: Evaluating LLM Agents as Lifelong Learners
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering