AgonAlpha: Autonomous Alpha Discovery via Prompt Economy and Scalable Agentic Search

arXiv:2608.11250 · cs.AI, cs.MA, q-fin.CP, q-fin.PM · Submitted 2026-08-04 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "AgonAlpha: Autonomous Alpha Discovery via Prompt Economy and Scalable Agentic Search".

Jane: The paper was written by Weicheng Ye, Youran Sun, Xingyu Ren, Shunyao Yu, Chugang Yi et al. from The Chinese University of Hong Kong and University of Maryland.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Alright, welcome back to the channel, everybody. We've got a paper that's been making the rounds, and I have to say, the title alone got me excited: "AgonAlpha: Autonomous Alpha Discovery via Prompt Economy and Scalable Agentic Search." Jane, you've been looking at this one with me, right?

Jane: Oh, absolutely, Tom. And for anyone tuning in who isn't deep in quantitative finance, let's just say the word "alpha" here isn't the Greek letter. It's the financial term for a trading signal that can beat the market. This paper is about building an AI system that hunts for those signals all on its own.

Tom: And it's not just any AI system. The authors are from the Chinese University of Hong Kong and the University of Maryland, and they've built something they call AgonAlpha. The whole idea is that you don't just ask a language model to spit out a formula. You set up a whole research team, but the team members are AI agents with very specific jobs.

Jane: Right, and that's the part I love. They've got a "proposer" that comes up with the trading ideas, and then a separate "reviewer" whose whole job is to be suspicious. The reviewer gets to re-run the simulations, check the evidence, and even veto the proposer if it catches something fishy.

Tom: It's like having a scientist and a skeptic in the same lab, which is way better than one scientist grading their own homework. And the results? They deployed this on a real platform called WorldQuant BRAIN, and they got some spectacular grades. We'll get into the numbers in a bit, but the architecture alone is worth the listen.

Jane: Definitely. And the name "Agon" comes from the Greek word for contest, which fits perfectly because this whole system is built on that adversarial back-and-forth. It's not just generating formulas; it's generating evidence, arguments, and counter-arguments.

Tom: So stick around, because we're going to break down how this "prompt economy" actually works and why it might change how we think about AI doing research. Jane, what's the one thing you want listeners to take away from the title alone?

Jane: That the unit of discovery isn't a formula anymore. It's a whole artifact — the hypothesis, the evidence, the rationale, the review status. That's a big shift.

Summary: Tom: So we're back, and we're still on "AgonAlpha: Autonomous Alpha Discovery via Prompt Economy and Scalable Agentic Search." Jane just set the stage with the artifact idea, and I want to dig into the summary because this paper is dense with results. Lu, you're our senior researcher — what jumped out at you first?

Lu: The numbers, Tom. They ran this across five different users with different model backends, and they got sixty submissions to the BRAIN platform. Out of those, seventeen came back with a "SPECTACULAR" grade, which is the highest tier on that platform. And the best Fitness score hit nine point five zero, with a Sharpe ratio of three point four eight.

Jane: And for our listeners who don't live in finance, a Sharpe ratio that high is like a basketball player shooting ninety percent from the free-throw line. It's not just lucky; it's consistent. But what's even more interesting to me is how they got there.

Meng: As the engineer on the show, I have to ask about the practical side. They mention a "halving tournament" where they start with sixteen candidate formulas and eliminate half each round. That's sixteen then eight then four then two then one. So they're only running thirty-one simulations instead of eighty. That's a huge cost saving.

Tom: Exactly, Meng. And that's the "prompt economy" part. Every time you ask the AI to do something, it costs money and time. So they built a scheduler that decides which research "lineage" gets the next chunk of budget, and it accounts for work that's still in progress. It's not just a fixed schedule; it's adaptive.

Lu: And that's where the innovation really sits. The scheduler uses a Monte Carlo Tree Search, but it's "pending-aware." That means if a branch of the search tree is already busy running simulations, the system doesn't pile more work onto it. It redirects to a branch that's free. That's a subtle but crucial detail for real-world deployment.

Jane: Right, and the reviewer part is what makes it trustworthy. They caught two cases of "fabrication" — where the proposer's report didn't match the actual simulation results. The reviewer zeroed out the score for those, which means the system can police itself.

Tom: So the summary is: a two-role AI system, a smart budget allocator, and a skeptical reviewer, all working together to find trading signals that actually pass external checks. And it did, seventeen times over.

Meng: And they released everything — every prompt, every decision, every review. That's rare in this space, and it's what makes the results believable.

Improvements: Tom: We're back on "AgonAlpha: Autonomous Alpha Discovery via Prompt Economy and Scalable Agentic Search," and we've covered the basics. Now let's talk about what this paper actually improves over what came before. Jane, you had a great analogy earlier about the scientist and the skeptic — how does that play out in the improvements?

Jane: Well, Tom, the big improvement is that they stop trusting the AI to grade its own work. A lot of previous systems had one language model generate a formula and then score it itself. That's like asking a student to grade their own exam. AlphaBench, which they cite, showed that LLMs are basically random at ranking factors. So AgonAlpha splits the job into two separate roles with fresh contexts.

Lu: And that's the "adversarial" part. The reviewer doesn't just look at the score; it audits the evidence. It checks whether the expression matches the platform records, whether there's any look-ahead bias, whether the economic rationale actually matches the formula. If the reviewer finds a mismatch, it can set the reward to zero, which kills that research path.

Meng: From a systems perspective, the improvement is in the allocation. Older systems used fixed schedules — run this many simulations, then move on. AgonAlpha uses a pending-aware scheduler that watches what's in flight. If a branch is busy, it doesn't double-book it. That's a real engineering win for parallel deployment.

Tom: And they validated that with ten concurrent workers. That's not a toy demo; that's a production-scale test.

Jane: Another improvement is the "artifact" itself. They don't just save the formula. They save the hypothesis, the candidate expressions that failed, the platform evidence, the review notes. So if you want to know why a certain alpha exists, you can trace it back through the whole decision tree.

Lu: That addresses a huge problem in the literature. The paper cites a review of thirty LLM trading papers that found most of them underreport execution details, transaction costs, and artifact availability. AgonAlpha releases everything — prompts, search decisions, platform records, executable expressions. That's the first complete prompt-to-factor trail in the audited literature.

Meng: And it's not just for finance. The architecture is domain-agnostic. The role prompts don't mention stocks or options. So you could point this same system at any problem with a clear evaluation metric — drug discovery, materials science, even logistics.

Jane: That's the exciting part for me. They've shown that a minimal two-role system with a smart scheduler can outperform much more complex multi-agent setups. It's not about adding more agents; it's about giving the ones you have the right authority and the right budget.

Conclusion: Tom: Alright, we're wrapping up our time with "AgonAlpha: Autonomous Alpha Discovery via Prompt Economy and Scalable Agentic Search." Jane, Lu, Meng — give me your one-sentence goodbye to this paper.

Jane: Mine is that this paper proves you don't need a huge team of AI agents to do serious research; you need two well-designed roles and a scheduler that respects the cost of every action.

Lu: And I'd add that the adversarial review with veto power is the missing piece that makes autonomous discovery trustworthy enough to deploy in production.

Meng: For me, it's the pending-aware allocation that makes it scale. Knowing what's in flight and redirecting budget accordingly is what turns a clever idea into a real system.

Tom: And I'll say this — the fact that they released the entire trail, every prompt and every decision, sets a new standard for the field. If you're going to claim your AI found something, you should be able to show your work.

Jane: Exactly. And we should mention that all sixty submissions entered out-of-sample tracking on the BRAIN platform, so the platform itself is still watching these alphas perform in real time. That's the ultimate test.

Tom: Great point. So to our listeners, if you're interested in how AI can do more than chat — how it can actually run a research lab, catch its own mistakes, and find real value in noisy data — this paper is a must-read. We'll be back soon with the next one.

Jane: Until then, keep asking the hard questions, and don't trust a formula without its evidence trail. See you next time.

Tom: Take care, everyone.

Weicheng Ye, Youran Sun, Xingyu Ren, Shunyao Yu, Chugang Yi, Haizhao Yang

The Chinese University of Hong Kong · University of Maryland

cs.AI, cs.MA, q-fin.CP, q-fin.PM

Submitted: 2026-08-04

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 57/100

The gist: The paper presents AgonAlpha, an architecture for autonomous alpha discovery that "searches over frozen research artifacts—hypotheses, executable expressions, platform evidence, rationales, and

Key concepts

AgonAlpha
A system designed to autonomously discover financial trading signals. It operates using an adversarial structure where specialized AI agents act as a 'proposer' (to generate ideas) and a 'reviewer' (to check for errors and veto suspicious results).
Prompt Economy
The cost of running the AI system. The paper uses an adaptive scheduler that manages the budget, ensuring that resources are allocated efficiently to research paths that are still in progress, rather than using a fixed schedule.
Artifact
A complete record of a discovery. Instead of just saving a formula, the the system saves the entire chain: hypothesis, evidence, rationale, and review status. This provides a traceable trail from prompt to final result.
Adversarial Search
The core methodology where two distinct AI roles work together. The 'proposer' creates the signal, and the 'reviewer' rigorously audits the evidence against simulation results, preventing self-grading bias found in previous systems.

Terminology

Summary

The paper presents AgonAlpha, an architecture for autonomous alpha discovery that searches over frozen research artifacts—hypotheses, executable expressions, platform evidence, rationales, and review status—rather than formulas alone. The authors state: "To our knowledge, AgonAlpha is the first alpha-mining system to combine verified artifact search, a fresh-context adversarial reviewer with re-execution and veto authority, and pending-aware parallel budget allocation, together with a complete public evidence trail."

The paper identifies three central challenges from prior literature. First, AlphaBench shows that language models still struggle with zero-shot factor ranking, even though they are relatively effective at generating executable factors. Second, "the performance of published anomalies has also weakened substantially over time. In the post-2005 non-micro-cap sample analyzed by Chen and Welch, the median return is only around 7 basis points per month—an effect size small enough to be explained away by statistical luck or modest transaction costs. Third, a systematic review of 30 LLM-based trading papers reveals a recurring gap in research practice: while model architectures are often described in detail, crucial aspects such as point-in-time data controls, train-test splits, out-of-sample evaluation, transaction costs, turnover assumptions, execution details, and artifact availability are frequently underreported."

The paper identifies four differences from prior formula-level systems:

  1. Search unit: "Existing systems typically treat a formula as the basic unit of discovery. However, a formula alone does not preserve the hypothesis that motivated it, the alternative directions that were explored and rejected, or the objections that were addressed along the way. We instead treat discovery as a lineage of research decisions."

  2. Verification mechanism: "Existing systems often rely on self-evaluation procedures or scalar performance thresholds. While these mechanisms are effective for ranking candidates, they do not independently verify whether the supporting evidence is consistent with the evaluated candidate. We introduce explicit verification procedures to audit the correspondence between generated candidates, evaluation records, and reported results."

  3. Resource allocation: "Existing systems commonly allocate search budgets according to fixed schedules. However, when external platform evaluations require substantial time and multiple research directions proceed concurrently, static allocation cannot adaptively redirect resources toward promising research lineages or account for evaluations already in progress."

  4. Evidence preservation: Existing systems often provide incomplete evidence trails... We therefore emphasize the preservation of complete research artifacts and execution history, enabling results to be reproduced and independently assessed.

AgonAlpha implements three coordinated components: a proposer, a fresh-context reviewer, and a pending-aware scheduler. The proposer commits each candidate as an evidence-bearing artifact; the reviewer independently checks the artifact, may rerun the associated platform evaluation, and can reject unsupported or fabricated claims; and the scheduler allocates each subsequent invocation to a research lineage while accounting for evaluations already in progress.

"The unit of search is an artifact A = (h, f, E, r, v): a hypothesis h, an executable expression f, evidence records E (simulation inputs, metrics, annual results, submission state), an economic rationale r, and a verdict v from the adversarial review. A lineage is a chain of artifacts linked by ancestry; the single search operator is extend(l), which invokes one proposer–reviewer pipeline to append a new artifact to lineage l."

The proposer receives ancestor reports, a work directory, two sampled readings, and the platform interface. It writes a hypothesis, generates 16 candidate formulas, and runs a halving tournament. The proposer submits the winner and freezes the full search as an alpha report. The reviewer "receives only the work directory and platform documentation. It runs in a fresh context on a different model route and has a single task: find grounds for rejection. It may re-run any simulation, and verified fabrication sets the score to zero. The two role prompts total 101 physical lines (57 proposer, 44 reviewer)."

The reviewer audits five dimensions: evidence integrity, sign logic, constant rationale, temporal stability, and selection risk. "Only verified fabrication—mismatched expressions, leaked or look-ahead inputs, invented records—changes the scheduler's score, to zero. Sign, constant, regime, and redundancy findings remain advisory records attached to the artifact. The paper notes: The reviewer never estimates alpha quality—quality is scored by the platform. The reviewer audits evidence integrity: re-execution, record consistency, and look-ahead. This is a consistency-checking task, not the quality-prediction task at which AlphaBench found LLM agents near-random."

Upper level: a pending-aware PW-MCTS across lineages using progressive widening with the condition "done children(n) < k · v(n) + π(n)" where k=1.0 and α=0.5, and a pending-aware UCB formula:


UCB(c) = Qsub(c)/v(c) + C * sqrt(ln(v(p) + π(p))/(v(c) + π(c)))

with C=10.0. The exploitation term values entire lineages using only completed evidence. Pending counts in the exploration term prevent over-allocation to branches whose workers are still busy. A root fallback unconditionally expands the root ρ, preventing pipeline starvation.

Lower level: Each proposer starts with 16 candidates and eliminates half per pass (16 → 8 → 4 → 2 → 1). The planned 31 simulations measure 31 distinct candidate versions versus the counterfactual 80 simulations for all-candidates-through-five-passes, a 61% reduction.

Rewards are computed as population percentiles via mid-rank with fabrication-zeroed artifacts receiving reward 0 regardless of raw Score.

The authors chose WorldQuant BRAIN as an external evaluator. On this production platform, consultants submit alphas for potential compensation, giving the platform economic stakes beyond our paper. The protocol fixed U.S. TOP3000 equities, delay one, and the platform evaluation window January 1, 2019 through December 31, 2023. Humans supplied the system and launched the evaluation, but wrote no factor expression or factor-specific program.

"Five co-authors independently deployed AgonAlpha on WorldQuant BRAIN using separate accounts and different model backends. Each ran the same two-role prompt surface and MCTS scheduler without human-written factor code. Collectively they produced 60 submissions, of which 17 received SPECTACULAR grade. The best observed Fitness and Sharpe ratio were 9.50 and 3.48, respectively."

Five featured alphas are detailed:

  1. A17q5RdR: aligned six-month option demand with construction f opt = M40[B60(IV180 put − IV180 call)], attaining Fitness 3.93, Sharpe 2.52, turnover 5.92%, and self-correlation 0.1827.

  2. 88er8JAl: relative-volume stability with f = −σ40(V/ADV20), recording Fitness 2.55, Sharpe 1.76, return 26.24%, and turnover 9.58%.

  3. LL1mdWz6: persistent short interest with reversal timing with Fitness 2.82, Sharpe 2.32, turnover 7.82%, drawdown 6.91%, and self-correlation 0.4173.

  4. pwlL71Ex: multi-tenor downside-insurance disagreement with Fitness 4.73 and Sharpe 3.03, accompanied by low turnover of 5.08%, self-correlation of 0.4968.

  5. KPE0LnN1: state-conditioned option confirmation attaining Fitness 9.50 and Sharpe 3.48, passes the sub-universe Sharpe check at 1.76 versus 1.51, and remains below the self-correlation ceiling at 0.6321.

"The reviewer audited 24 frozen reports across the full trace and exercised its zero-score authority twice. One intervention addressed a semantic mismatch between a claimed mechanism and its executed operator. The other addressed metrics copied from a different candidate into a final report. Eleven additional artifacts carry persistent risk findings for regime concentration or redundancy."

"Within each pipeline, tournament elimination evaluates 16 + 8 + 4 + 2 + 1 = 31 candidate versions—a 61% reduction from the 80 simulations required to carry all 16 candidates through five passes. A validated deployment instantiated ten concurrent proposer–reviewer pipelines under one pending-aware search tree."

"The accompanying release contains the proposer, reviewer, and dispatcher prompts; prompt history; scheduler code and complete MCTS state; all candidate reports; simulation inputs and responses; rankings; correlation records; submission checks; and final submission responses. The paper claims: This is the first complete prompt-to-factor release in the audited LLM trading literature."

The paper states AgonAlpha is the first system to instantiate all six Agon principles outside of the original Agon factory: Prompt Economy, Minimal Prompts, Future-Facing, Zero-Code, OmniDisciplinary, and Massive Parallelism. AgonAlpha's entire discovery workflow runs on 101 physical lines of role prompt: 57 for the proposer, 44 for the reviewer, plus a short dispatcher. The role prompts contain no market-specific tokens: no 'stock,' no 'equity,' no 'option.'

The paper acknowledges: "A chronological holdout and an external evaluator address different failure modes. A holdout tests later rows under the same author-configured pipeline; BRAIN instead removes the data, simulation, metric, gate, and grade implementations from our control. The paper also notes that the sign and constant dimensions are judgmental; their reliability is evaluated through the reviewer audit record. Additionally, the 48-day mean is the main localized tuning risk for some alphas, and the positive sign on high short interest is an empirical open question."

"AgonAlpha suggests that the core principles of the Agon philosophy can be compressed into a minimal yet complete discovery architecture: two roles and a 101-line prompt are sufficient to support an adversarially verified research workflow. The fundamental search unit is not a formula, but a verified artifact containing the evidence and reasoning behind a candidate. The verifier is granted the authority to independently reproduce evaluations and veto unsupported claims, while the scheduler converts concurrent exploration into a structured search over research lineages. By releasing the complete execution trace, AgonAlpha enables inspection rather than blind trust of every reported result. Because these mechanisms operate independently of the underlying domain, the same two-role interface and MCTS-based scheduling framework can be applied to any setting with a well-defined objective evaluation metric."

Improvements for AI systems

Based on the paper, here are specific improvements I can implement in AI systems:

Improvement: Replace single-formula generation with frozen research artifacts containing: hypothesis, executable expression, evidence records, economic rationale, and review verdict.

What the improved system can do:

  • Maintain full lineage context (why a candidate was generated, what alternatives were rejected)

  • Preserve reasoning trails that enable debugging and reconstruction

  • Search across research decisions, not just formula edits

Improvement: Implement a separate reviewer agent with fresh context, independent model routing, and authority to:

  • Re-run any platform simulation to verify evidence integrity

  • Set scheduler reward to zero upon verified fabrication (mismatched expressions, look-ahead inputs, invented records)

  • Attach advisory warnings (sign logic, constant rationale, temporal stability, selection risk) without blocking submission

Improvement: Implement a pending-aware MCTS scheduler with:

  • Upper level: Progressive widening (branching factor grows as k·(v+π) α), pending-count-aware UCB (ln(v(p)+π(p))/(v(c)+π(c))), backpressure (one in-flight child per non-root node), root fallback to prevent starvation

  • Lower level: Halving tournament within each pipeline (16→8→4→2→1 candidates, 31 simulations vs 80 for all-candidates)

Improvement: Convert raw platform scores to population percentiles via mid-rank, computed at completion and frozen (never recomputed as population grows).

Improvement: Use exactly two role prompts (57-line proposer, 44-line reviewer) plus a short dispatcher, with:

  • No market-specific tokens in role prompts (domain knowledge enters via runtime readings)

  • No model-specific instructions (future-facing, benefits from model upgrades)

  • Zero-code principle (research logic in prompts, not hardcoded pipelines)

Improvement: Use an external, economically-incentivized evaluation platform (WorldQuant BRAIN) rather than author-controlled backtests, and release complete trace: prompts, search decisions, platform records, review text, executable expressions.

Improvement: For dollar-neutral cross-sectional alphas, negate the expression to flip signs of returns and Sharpe while leaving turnover unchanged—derive flipped metrics exactly rather than re-simulating.

Improvement: Move near-duplicate detection (correlation threshold 0.85) into the proposer as a hard pre-ranking gate, before tournament ranking.

Improvement: Require every candidate to include an economic mechanism explanation; reviewer flags unexplained constants, counterintuitive signs, and regime concentration as advisory warnings.

Improvement: Record annual Fitness sequences and flag when best-to-worst annual ratio exceeds 5, attaching dated regime warnings to artifacts.

Improvement: Between tournament passes, revise surviving candidates (not just resample fixed arms), so each simulation measures a distinct candidate version.

Improvement: When UCB descent reaches a busy dead end (all children in-flight), unconditionally expand the root to open a new lineage.


Summary of what the fully improved AI system can do: It can autonomously discover, verify, and submit trading alphas with full provenance, achieving SPECTACULAR grades (Fitness up to 9.50, Sharpe up to 3.48) across multiple users and model backends, while catching its own errors, allocating budget efficiently under concurrency, and providing complete auditability for every result.

Abstract

Language models can propose many plausible trading factors, but an autonomous research system must also allocate its evaluation budget, verify its own evidence, and preserve how each candidate was produced. We present AgonAlpha, an architecture that searches over frozen research artifacts---hypotheses, executable expressions, platform evidence, rationales, and review status---rather than formulas alone. To our knowledge, AgonAlpha is the first alpha-mining system to combine verified artifact search, a fresh-context adversarial reviewer with re-execution and veto authority, and pending-aware parallel budget allocation, together with a complete public evidence trail. Independent deployments on WorldQuant BRAIN produced SPECTACULAR-grade alphas across five users and six model backends, with Fitness reaching 9.50 and Sharpe reaching 3.48, while retaining prompt-to-expression provenance for every submission.

Sources

Related papers