FinEvo-Bench: A Longitudinal Benchmark for Self-Evolving Agents in Professional Financial Workflows

arXiv:2608.06144 · cs.AI · Submitted 2026-08-06 · Read on arXiv

Bo Deng, Kang Zhou, Lifan Guo, Chongyang Tao, Xuanren Chen, Chenggang Xie, Renzhao Liang, Feng Chen, Chi Zhang

Beihang University · Qwen DianJin Team · Alibaba Cloud Computing

cs.AI

Submitted: 2026-08-06

Comments: 22 pages, 4 figures; includes appendices

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 74/100

The gist: The paper introduces FinEvo-Bench, a longitudinal benchmark designed to evaluate self-evolving agents in professional financial workflows.

Terminology

Summary

The paper introduces FinEvo-Bench, a longitudinal benchmark designed to evaluate self-evolving agents in professional financial workflows. The authors motivate the work by noting that most agent benchmarks evaluate tasks independently and cannot measure whether experience from one task helps with later tasks, and that existing self-evolution benchmarks do not jointly cover professional workflows, open-ended deliverables, and multi-aspect evaluation. The benchmark specifically targets whether an agent extracts signals from completed tasks and feedback, updates its persistent state, and applies that state to later tasks.

The paper defines self-evolution as a cross-task process: an agent extracts signals from completed tasks and feedback, updates its persistent state, and applies that state to later tasks. FinEvo-Bench measures both professional performance and self-evolution ability: how effectively an agent turns prior experience into later improvement. The authors emphasize that final artifact quality, domain-specific compliance, and improvement from retained experience answer different questions—a scaffold may achieve a high final score because its backbone is strong, or show a large gain while remaining weaker in absolute terms. Therefore, they report evolved performance and paired within-scaffold gains separately: the former measures final capability; the latter, self-evolution ability.

FinEvo-Bench contains 120 real-case-grounded tasks across 20 business scenes spanning six financial domains. The 120 tasks contain 775 input files, averaging 6.46 per task (range: 2–11), and request open-ended reports, assessments, or recommendations. Each scene contains six substantively distinct cases that share a professional procedure and a manually reviewed rubric for task quality and financial compliance.

Each scene is constructed from three sources: an institution-provided scene description Ds, a reference professional procedure Ps validated in practice, and a candidate pool Cs drawn from institution-provided and publicly documented cases. A domain expert consolidates these into a scene specification:

Ss = Φ(Ds, Ps, Cseligible) = (Os, Es, As, Ys, Ks), where Os is the business objective and scope; Es is the required inputs; As is the professional operations and checks; Ys is the expected deliverable; and Ks is the compliance requirements and professional boundaries. Before release, direct identifiers and sensitive fields are removed or replaced when necessary while preserving data relationships, numerical logic, chronology, and decision conditions.

Cases within a scene must differ substantively in their business situations, input files, analytical focus, judgment conditions, or conclusions; cases with only superficial differences are excluded from the same scene. The six tasks in a scene share a 100-point rubric derived from the professional procedure. "The rubric checks whether an agent uses the input files correctly, completes the necessary analysis, and reaches supported conclusions. It does not require the output to match the reference answer in structure or wording. The rubric also checks financial compliance: fabricated or unsupported data and terms, definitive claims based on insufficient information, decisions beyond the agent's professional role, and guarantees about credit approval, claim outcomes, or investment returns." Two domain experts construct each rubric—one drafts, the other applies it to all six tasks to check applicability and consistency.

For validation, on the 120 deliverables from one complete main-experiment run, the automated rubric judge achieves high absolute agreement with a financial expert (ICC(A, 1) = 0.95; 95% CI: [0.93, 0.97]). The judge's mean and maximum absolute differences from expert scores are 1.6 and 5 points, respectively, on the 0–100 scale. ICC measures agreement in absolute scores, not only consistency in ranking the deliverables.

The evaluation uses three independently shuffled, globally interleaved task streams—rather than contiguous scene blocks—requiring each scaffold to retain and retrieve relevant experience amid unrelated intervening tasks. Each evolving run starts without benchmark-derived experience and processes tasks sequentially in stream order, following an execute–score-and-feedback–reflect-and-consolidate cycle.

Four self-evolving agent scaffolds are compared, all using Qwen3.7-Max with a 1M-token context window, maximum output length of 64K tokens, greedy decoding, and temperature zero:

  • Claude Code and Codex can distill task experience into reusable skills, project-scoped memory, and global memory.

  • Letta maintains editable memory blocks, with memory-management tools that can insert, replace, or rewrite content; every model call receives the full contents of all core memory blocks.

  • GenericAgent distills each completed task into a reusable Markdown experience file, with a prompt-resident title-only L0 index guiding selective loading of full files.

The evaluation protocol pairs each evolving run with a state-reset control: "The paired non-evolving condition resets agent state before every task. Within each run, both conditions share the task order, backbone, decoding configuration, and scoring procedure; only retention of prior-task feedback differs." An independent Claude Code scoring agent backed by Claude Opus 4.6 applies the scene rubric and returns rubric-based feedback.

Importantly, "the scaffold never receives the complete scene rubric or the judge's full item-level record. After finalizing a task, it receives rubric-based feedback summarizing the problems in the current deliverable and their corresponding reasons... providing concrete feedback for reflection while limiting access to the complete evaluation specification, rather than exposing the full rubric for direct optimization."

Four metrics are reported: (1) task quality (mean scene-specific rubric score, 0–100); (2) financial compliance (mean number of triggered compliance issues per task); (3) self-evolution ability measured by paired score gain and compliance-issue reduction relative to the non-evolving condition; and (4) agent-side cost (mean token consumption in units of 104 per task, excluding the independent rubric judge).

Overall performance. "Across the three independently shuffled task streams, retained experience yields positive self-evolution gains for all four scaffolds: relative to the paired non-evolving controls, mean scores rise by 9.33–19.37 points and mean compliance issues fall by 0.12–0.44 per task." Specifically:

  • Letta achieves the highest evolved score (91.65) and fewest compliance issues (0.09 per task).

  • Codex achieves the largest self-evolution gain (+19.37).

  • GenericAgent has the lowest evolved score (83.34) and smallest gain (+9.33), despite its lowest token cost.

The paper notes that no single scaffold is best in absolute performance, self-evolution gain, and cost. Cost differences are substantial: GenericAgent's total cost is less than one-third of Claude Code's, while Letta uses 17.87 × 104 reflection tokens and 32.56 × 104 execution tokens per task because every core memory block enters each model call. GenericAgent uses only 10.21 × 104 and 11.57 × 104, respectively. Final scores and self-evolution gains produce different rankings: Letta ranks first by evolved score, whereas Codex ranks first by gain.

Longitudinal evolution. The paper ranks each scene's six tasks by their order in the global stream and compares gains over ranks 1–3 (Early) versus ranks 4–6 (Late). "Despite fluctuations across individual ranks... every scaffold gains 6.10–8.70 points more over ranks 4–6 than over ranks 1–3. Codex has the highest late gain (23.62), while Letta has the largest early-to-late increase (8.70). Compliance reductions also increase by 0.02–0.10 issues per task. Late gains exceed early gains in each of the three independently shuffled streams, so the trend is not driven by one task order."

Experience utilization. For the three selective-retrieval scaffolds (Claude Code, Codex, GenericAgent), activated tasks score 8.31–9.06 points higher than non-activated tasks. Codex activates experience on 102/120 tasks, Claude Code on 87/120, and GenericAgent on 71/120 (Letta's 120/120 is described as an always-on interface, not selective retrieval). The non-activated subsets still score 3.97–12.31 points above the state-reset averages, suggesting the scaffolds consider their internal knowledge and task-local files sufficient.

Memory vs. skills vs. combination. In Claude Code, the paper compares memory-only, skill-only, and unrestricted memory–skill evolution, with no-evolution and a fixed expert skill as references. Skill-only performs best, scoring 93.71 with 0.05 compliance issues per task... The combined memory–skill setting improves neither quality nor compliance. Execution traces show reflection updates both stores, but later tasks often load only one; incomplete loading may partly explain the absence of an additional gain. Skill-only also uses less reflection overhead (44.03 vs. 60.19 × 104 tokens) than the combination, while memory-only is lowest at 26.18. The fixed expert skill outperforms no evolution but trails all dynamically updated carriers, supporting continued updates beyond an initial professional procedure.

Rubric feedback vs. reference answers. Holding all other settings fixed, rubric feedback is compared to providing a complete reference answer. Rubric feedback raises scores by 3.95–7.93 points, reduces compliance issues by 0.06–0.14, and yields 4–9 more activations for selective-retrieval scaffolds. The explanation: "A reference answer gives one valid solution but may not identify what the current response is missing or separate case-specific choices from reusable procedures. Rubric feedback instead identifies omitted evidence, incomplete analysis, and compliance failures in the current response."

Capability dimensions. Every rubric item is mapped to one of five dimensions: information and evidence use, analysis and calculation, conclusions and recommendations, report quality, and financial compliance. Every scaffold improves across all five dimensions... Report quality has the largest quality gain for each scaffold (6.85–18.17 percentage points), while normalized compliance gains range from 1.93 to 6.74 points. The authors suggest structure, coverage, and expression transfer more readily than evidence use, analysis, and conclusions, which remain tied to case-specific facts. GenericAgent improves least in every dimension.

Cross-scene interference. Using a five-scene sample (30 tasks), the paper compares mixed-scene execution with scene-isolated execution (fresh agent state per scene). Scene isolation lowers Claude Code and Letta by 0.84 and 1.00 points, but raises Codex and GenericAgent by 0.46 and 1.63 points. The score effects are small and mixed in direction, but "scene-isolated execution also triggers 0.01–0.05 more compliance issues per task for every scaffold, consistent with cross-scene transfer of general practices such as evidence checking, cautious wording, and avoiding decisions beyond the agent's role."

The paper lists three contributions:

  • FinEvo-Bench provides 120 real-case-grounded, multi-file tasks with open-ended outputs scored for quality and financial compliance.

  • A paired protocol uses interleaved streams and state-reset controls to measure cross-task self-evolution gains.

  • Shared-backbone experiments compare four scaffolds, experience carriers, feedback forms, capability dimensions, and cross-scene execution.

The paper acknowledges: "FinEvo-Bench covers 20 scenes in six financial domains, uses one backbone, and studies only non-parametric evolution. Results may therefore differ across professions, backbones, and parameter-updating agents. The carrier comparison covers only Claude Code; the cross-scene diagnostic uses five scenes and one ordering. Future work should expand these settings."

The paper concludes: "FinEvo-Bench evaluates four self-evolving agent scaffolds on 120 real-case-grounded financial tasks. The automated rubric judge agrees closely with financial-expert scoring on 120 outputs (ICC(A, 1) = 0.95). Across three shuffled, globally interleaved streams, all four scaffolds outperform their paired state-reset controls: scores rise by 9.33–19.37 points and compliance issues fall by 0.12–0.44 per task. Letta records the highest evolved score (91.65), whereas Codex records the largest self-evolution gain (+19.37); GenericAgent has the lowest cost. Skill-only evolution performs best among Claude Code's carriers, rubric feedback outperforms reference answers, and all five capability dimensions improve. Cross-scene score differences remain small and mixed, while scene isolation triggers 0.01–0.05 more compliance issues per task."

Improvements for AI systems

Improvements I can make to AI systems based on FinEvo-Bench:

  1. Add longitudinal self-evolution memory. The agent extracts lessons from completed tasks, consolidates them into persistent memory or reusable skills, and applies them to later tasks. This raises task scores by 9–19 points and reduces compliance violations by 0.12–0.44 per task.

  2. Provide rubric-based feedback after each task instead of reference answers. Feedback highlighting omitted evidence, incomplete analysis, and compliance failures improves scores by 4–8 points, reduces compliance issues, and increases the agent's ability to transfer lessons to future tasks.

  3. Use selective experience retrieval instead of always-on memory. Activated relevant experiences boost task scores by 8–9 points. A lightweight title-only index with on-demand loading of full experience files achieves strong gains at much lower token cost.

  4. Prefer skill-based evolution over unrestricted memory-plus-skill accumulation. In the Claude Code scaffold, skill-only evolution reached 93.71 quality with 0.05 compliance issues per task; combining memory and skills did not improve results and increased reflection cost. The agent should maintain a small, curated skill store updated dynamically from feedback.

  5. Make the agent explicitly check financial compliance. After producing a deliverable, the agent verifies that it has not fabricated data, made unsupported claims, overstepped its professional role, or guaranteed outcomes. This reduces compliance issues by 0.12–0.44 per task.

  6. Train the agent to separate reusable procedures from case-specific facts. Capability-dimension analysis shows report quality and general practices transfer easily, while evidence use and analysis remain case-specific. The agent should store generalizable heuristics—such as check evidence before making claims and use cautious wording—to support cross-scene transfer.

  7. Evaluate the agent with interleaved, globally shuffled task streams. This forces the agent to retain relevant experience amid unrelated tasks, making its memory more robust and its retrieval more selective. Late-task gains exceed early-task gains by 6–9 points when using this protocol.

  8. Report both absolute performance and self-evolution gain separately. The improved system can distinguish whether a high final score comes from a strong backbone or from genuine learning, enabling better scaffold selection and clearer optimization targets.

  9. Optimize the cost–performance trade-off using measured token usage. The agent can choose a configuration that matches deployment priorities: lowest cost, highest final score, or largest self-evolution gain, based on empirically measured trade-offs.

  10. Use a state-reset control in agent evaluation. The improved system can measure its true self-evolution ability by comparing against a paired run where memory is cleared before every task, isolating the effect of retained experience from task-order and backbone effects.

What the improved AI system can do:

  • Complete professional financial workflows—reports, assessments, recommendations—using multiple input files, with scores near 92/100 and fewer than 0.1 compliance issues per task.

  • Learn from each completed task's feedback and measurably improve on later tasks across shuffled, interleaved workflows.

  • Retrieve only relevant prior experience, saving token cost while maintaining high quality.

  • Avoid fabricated data, unsupported definitive claims, role overreach, and inappropriate guarantees.

  • Transfer general compliance and reporting practices across different business scenes, reducing compliance violations even when task content changes.

  • Provide explainable self-evolution metrics: absolute performance and learning gain are tracked separately.

Sources

Related papers