FINSKILLOPS: A Self-Evolving Multi-Agent System for SEC Filing QA
cs.AI, cs.IR, cs.MA, cs.SE
Submitted: 2026-09-17
Updated: 2026-09-17
Code: https://github.com/chroma-core/chroma
License: http://creativecommons.org/licenses/by/4.0/
The gist: Financial QA systems are typically improved before deployment through better retrieval, prompting, or agent coordination, leaving their reliability behavior fixed thereafter.
Terminology
Abstract
Financial QA systems are typically improved before deployment through better retrieval, prompting, or agent coordination, leaving their reliability behavior fixed thereafter. In practice, new SEC-filing questions repeatedly expose heterogeneous errors in period, entity, evidence use, and calculation. Existing self-improvement methods can turn failures into new behaviors, but offer limited control over where a correction should apply or which previously correct answers it may break. We therefore frame post-deployment improvement as controlled behavioral maintenance: recurring failures should become scoped skill patches, and each patch should earn deployment with- out introducing regressions. We instantiate this view in FINSKILLOPS, a multi-agent system for SEC filing QA. FINSKILLOPS derives reusable skills from evidence-grounded, typed failure diagnoses and governs them through targeted validation, protected-case regression checks, negative controls, and versioned replacement or retirement. Across six financial QA benchmarks, a single frozen skill registry achieves the highest verdict-weighted correctness and reference consistency among the evaluated systems. Evolved skills raise correctness from 3.70 to 4.55 on our enhanced benchmark. In a separate 12-round operational study, only six of 33 proposed skills are promoted, while the monitoring non-correct rate falls from 20.0% to 12.5%. These results establish controlled skill scope, admission, and lifecycle management as the foundation for reliable self-improvement.
Sources
- M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation
- FinDER: Financial Dataset for Question Answering and Evaluating Retrieval-Augmented Generation
- FinanceBench: A New Benchmark for Financial Question Answering
- Fin-RATE: A Real-world Financial Analytics and Tracking Evaluation Benchmark for LLMs on SEC Filings
- MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation
- SkillForge: Forging Domain-Specific, Self-Evolving Agent Skills in Cloud Technical Support
- Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills
- SkillOps: Managing LLM Agent Skill Libraries as Self-Maintaining Software Ecosystems
- VeritasFi: An Adaptable, Multi-tiered RAG Framework for Multi-modal Financial Question Answering
- MinerU: An Open-Source Solution for Precise Document Content Extraction
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- Mixture-of-Agents Enhances Large Language Model Capabilities
- Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward
- SkillOpt: Executive Strategy for Self-Evolving Agent Skills
- AutoSkill: Experience-Driven Lifelong Learning via Skill Self-Evolution
- CoEvoSkills: Self-Evolving Agent Skills via Co-Evolutionary Verification
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection