Propose, Don't Judge: An Anytime-Valid Referee for LLM Agents That Mine Investment Factors
cs.AI, q-fin.PM, q-fin.ST
Submitted: 2026-09-22
Updated: 2026-09-22
Terminology
Sources
- Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops
- Sequential testing of conditionally constrained hypotheses
- Multiple testing in multi-stream sequential change detection
- An online generalization of the (e-)Benjamini-Hochberg procedure
- The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World
- Dynamic $e$-closure for online hypotheses with any-time-valid evidence: closure principles and projective mergers
- R&D-Agent-Quant: A Multi-Agent Framework for Data-Centric Factors and Model Joint Optimization
- CLQT: A Closed-Loop, Cost-Aware, Strategy-Consistent Benchmark for Diagnostic Evaluation of LLM Portfolio-Management Agents
- PACE: Anytime-Valid Acceptance Tests for Self-Evolving Agents
- Admissibility and Complete Classes for False Discovery Rate Control with E-values
- AlphaAgent: LLM-Driven Alpha Mining with Regularized Exploration to Counteract Alpha Decay
- Anytime-valid FDR control with the stopped e-BH procedure
- Improving online FDR procedures via online analogs of e-closure and compound e-values
- AlphaSchema: Exploring the Space of Trading Semantics for LLM-Based Alpha Mining
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection