FactorEngine: A Program-level Knowledge-Infused Factor Mining Framework for Quantitative Investment

arXiv:2603.16365 · cs.AI · Submitted 2026-03-17 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "FactorEngine: A Program-level Knowledge-Infused Factor Mining Framework for Quantitative Investment".

Jane: The paper was written by Qinhong Lin, Ruitao Feng, Yinglun Feng, Zhenxin Huang, Yukun Chen et al. from Beijing University of Posts and Telecommunications and Beijing Value Simplex Technology Company Limited and Yangtze Delta Research Institute, University of Electronic Science and Technology of China.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: We're starting things off with a heavy hitter today called "FactorEngine: A Program-level Knowledge-Infused Factor Mining Framework for Quantitative Investment." It sounds like something straight out of a high-tech trading floor.

Jane: It really does, Tom! But if we strip away the jargon, it's basically about building a machine that finds hidden patterns in the stock market to help people make better investments.

Lu: And it isn't just any machine, Jane! The authors from Beijing University of Posts and Telecommunications and Beijing Value Simplex are proposing something that actually thinks like a researcher.

Meng: That sounds ambitious, Lu, but I'm wondering how they actually implement that in a real production environment. Most "thinking" systems are too slow or too messy for actual trading.

Jane: That’s the interesting part, Meng, because the title mentions it's "program-level," which means it writes actual code instead of just guessing numbers.

Lu: Exactly! Imagine a world where the software isn't just following a fixed recipe, but is actually writing its own cookbook as it learns about the market.

Meng: If they can actually write clean, executable Python code like the paper suggests, that would solve a massive headache for engineers who usually have to manually fix these models.

Lalam: It goes beyond just fixing code, though; it represents a shift where human financial wisdom and machine execution finally speak the same language. This could change how we teach finance by turning abstract theories into living, breathing digital logic.

Tom: That's a profound way to look at it, Lalam. We're going to get into exactly how they make that happen in the next part of our show.

Summary: Tom: Now that we've seen the name, let's talk about how FactorEngine actually works, because it uses this clever "macro-micro" approach.

Jane: I loved how they explained this, Tom! They basically separate the "big ideas"—the logic—from the "fine-tuning"—the tiny mathematical details.

Lu: It’s like having a brilliant architect who designs a house and then a separate, incredibly fast construction crew that handles every single screw and nail.

Meng: So you're saying they use the LLM to handle the architectural design of the factor, while something else does the heavy lifting of testing the numbers?

Jane: Precisely! They use Large Language Models to come up with new ideas for code, but then they hand off the tedious math to a local computer using something called Bayesian optimization.

Lu: And don't forget the "Bootstrapping Module," which is my favorite part! It can actually read a boring financial report and turn it into working Python code.

Meng: Wait, so if a researcher writes "momentum is increasing," the system can actually see that text and write a program to track it?

Jane: That's exactly what happens, Meng! It bridges the gap between human words and machine math.

Lalam: This creates a beautiful loop where human insight isn't lost in translation but is instead amplified by AI. It turns the vast ocean of financial text into a structured library of actionable intelligence.

Tom: That's a perfect setup for our next segment, where we look at whether this actually works when you put it to the test.

Improvements: Tom: We've seen the theory, but let's talk about the actual results from "FactorEngine: A Program-level Knowledge-Infused Factor Mining Framework," because the numbers are pretty wild.

Jane: They tested this on real market data in China, specifically the CSI300 and CSI500 markets, and it outperformed almost everything else.

Lu: It didn't just win; it crushed the existing benchmarks like Alpha158! I was looking at how much more diverse their factors were compared to other agents.

Meng: I saw that in the data, too. How much of an improvement are we talking about when we compare it to something like RD-Agent?

Jane: Well, for the version that learns from financial reports—the FE-report setup—they saw a massive jump in things like the Information Coefficient, which is just a fancy way of saying their predictions were much more accurate.

Tom: And they even saw much better annual returns and lower "drawdowns," which means they didn't lose as much money during market dips.

Lu: It’s because the system is constantly evolving! It doesn't just find one good factor and stop; it keeps refining its entire library of code.

Meng: That sounds like it would be very efficient, especially since they used a framework called Polars to make the computations run incredibly fast.

Lalam: The most impressive part is the stability. Even as markets change over years, these factors don't just decay and become useless like older models; they seem to adapt and stay relevant.

Tom: It really feels like we're looking at a new standard for how quantitative research will be done.

Conclusion: Tom: We've covered a lot of ground today, from the high-level architecture to the impressive backtesting results of "FactorEngine: A Program-level Knowledge-Infused Factor Mining Framework."

Jane: It’s such a clever way to combine what humans are good at—reasoning and reading—with what machines are good at—coding and optimizing.

Lu: I'm already thinking about how this could be applied to other fields, like discovering new laws of physics or designing new materials!

Meng: From my side, the practical takeaway is that this makes the whole pipeline much more reliable for real-world deployment.

Lalam: And culturally, it shows a future where AI acts as a collaborator that respects and elevates human expertise rather than just replacing it.

Tom: Well, on that note, we have to wrap this up. Thank you all for joining us!

Jane: Thanks for listening, everyone! We'll see you next time with another incredible paper.

Beijing University of Posts and Telecommunications · Beijing Value Simplex Technology Company Limited · Yangtze Delta Research Institute, University of Electronic Science and Technology of China

cs.AI

Submitted: 2026-03-17

Updated: 2026-09-15

Comments: 10 pages, 7 figures. Accepted at IEEE ICDM 2026

Code: https://github.com/microsoft/qlib

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 85/100

The gist: This paper introduces FactorEngine (FE), a program-level factor discovery framework designed to automate the mining of predictive signals from noisy, non-stationary market data.

Key concepts

Macro-micro approach
A method that separates the logical design of a factor from its mathematical fine-tuning. Large Language Models handle the high-level architectural design and code creation, while Bayesian optimization is used to perform the heavy lifting of testing and optimizing the specific numerical details.
Bootstrapping Module
A feature that converts human financial knowledge into machine-readable logic. It can read qualitative information from financial reports and automatically translate those text-based insights into executable Python code, allowing the system to turn abstract theories into structured, actionable digital intelligence.
Information Coefficient
A metric used to measure the accuracy of a model's predictions. In testing on Chinese markets like the CSI300 and CSI500, FactorEngine showed significant improvements in this coefficient compared to other agents, indicating its ability to more accurately predict market trends.

Terminology

Summary

This paper introduces FactorEngine (FE), a program-level factor discovery framework designed to automate the mining of predictive signals from noisy, non-stationary market data. By casting factors as Turing-complete code, FE addresses the limitations of existing symbolic and neural approaches, providing a scalable, interpretable, and computationally efficient method for quantitative investment.

The core challenges and approach

The authors identify three critical challenges in current alpha mining: (1) bounded expressiveness due to symbolic factor reliance, (2) limited factor diversity and stability, and (3) inefficient evolution pipelines. To address these, FactorEngine (FE) introduces a program-level factor discovery framework that treats factors as Turing-complete code. This allows for complex control flows, conditional logic, and iterative computation, making the factors more adaptable to rapidly changing market conditions than traditional symbolic or neural methods.

The three key separations

To optimize both effectiveness and efficiency, FE realizes three key separations: (i) logic separation between program logic/idea evolution and parameter optimization, (ii) search strategy separation between LLM-driven directional search and automated Bayesian search, and (iii) resource separation between LLM utilization and local computation resources. By decoupling these, LLM agents focus on logic discovery while local computation with Bayesian search automates parameter optimization, effectively overcoming efficiency bottlenecks and reducing evolution costs.

The Bootstrapping and Evolution modules

The framework operates through a closed-loop pipeline featuring a knowledge-infused bootstrapping module and an Evolution Module. The bootstrapping module employs a closed-loop multi-agent extraction–verification–code-generation pipeline to transform unstructured financial reports into executable factor programs. The Evolution Module then performs macro–micro co-evolution via a four-stage pipeline:

  • Program Selection: Using the Upper Confidence Bound for Trees (UCT) criterion to select promising nodes from a tree-structured search space.

  • Idea Generation: Leveraging chains of experience (CoE) to guide LLMs in heuristic searches within high-dimensional code spaces, allowing the system to learn from failures.

  • Implementation: Executing micro mutations through automated Bayesian search to fine-tune parameters like window sizes and decay factors.

  • Analysis: Utilizing feedback propagation to transform empirical outcomes into actionable guidance for subsequent evolution.

Performance and diversity

Extensive backtests on real-world OHLCV data show that FE achieves state-of-the-art predictive and portfolio performance. The framework produces factors with substantially stronger predictive stability and portfolio impact, such as higher IC/ICIR (and Rank IC/ICIR) and improved AR/Sharpe compared to baseline methods. Notably, FE demonstrates a 58% improvement in Information Coefficient (IC) and a 126% increase in excess annual return compared to Alpha158 when using factors derived from financial reports. Additionally, FE enhances factor diversity through a multi-island evolution configuration that improves sampling diversity and reduces redundancy.

Improvements for AI systems

1. Automated Scientific Theory Synthesis (ASTS)

  • Improvement: Apply the Bootstrapping Module and Macro-Micro Co-evolution loop to scientific discovery rather than financial data.

  • Capability: The system will ingest massive corpora of unstructured scientific literature (PDFs), use a multi-agent pipeline to extract mathematical laws and chemical/physical properties, and transform them into executable simulation code (e.g., molecular dynamics or fluid mechanics scripts). It will then use LLM-guided macro mutations to propose new physical hypotheses/models and Bayesian optimization to fine-tune the parameters of those simulations, enabling the autonomous discovery of new materials or drug candidates with grounded scientific rationales.

2. Hybrid Symbolic-Neural Co-evolutionary Architectures

  • Improvement: Implement Macro-Micro Separation and Programmatic Representation for Neural Architecture Search (NAS).

  • Capability: Instead of searching for weights within a fixed architecture, the system will evolve Hybrid Programs where an LLM designs the high-level symbolic logic (e.g., control flow, conditional branches, or graph topology) and a Bayesian optimizer handles the continuous parameters (e.g., neural weights or sensor gains). This produces AI systems that combine the high performance of deep learning with the interpretability and auditability of symbolic logic, specifically for mission-critical applications like autonomous robotics or medical diagnostics.

3. Multi-Island Chain-of-Experience (CoE) Reasoning Agents

  • Improvement: Integrate Multi-Island Evolution and Chain of Experience (CoE) into agentic reasoning frameworks.

  • Capability: A multi-agent system designed for complex, long-horizon task planning that runs multiple independent reasoning islands to explore different strategic paths. Each island maintains a structured Chain of Experience that records not just successful outcomes, but also the specific logical trajectories that led to failures. Periodically, successful reasoning sub-graphs (successful logic modules) are migrated between islands. This prevents the agent from falling into local reasoning optima and allows it to learn from mistakes by internalizing why certain planning paths failed.

4. Programmatic Software Optimization & Refinement

  • Improvement: Utilize the Logic Revision vs. Parameter Optimization separation for automated software engineering.

  • Capability: An AI system that treats complex software modules as Turing-complete programs to be evolved. The LLM agent performs Macro Mutations by proposing high-level structural refactoring or algorithmic changes (logic evolution), while a local Bayesian search performs Micro Mutations by optimizing low-level implementation details (e.g., memory allocation sizes, loop unrolling factors, or cache-line alignments). This results in software that is optimized for specific hardware constraints without human intervention.

Abstract

We study alpha factor mining, the automated discovery of predictive signals from noisy, non-stationary market data-under a practical requirement that mined factors be directly executable and auditable, and that the discovery process remain computationally tractable at scale. Existing symbolic approaches are limited by bounded expressiveness, while neural forecasters often trade interpretability for performance and remain vulnerable to regime shifts and overfitting. We introduce FactorEngine (FE), a program-level factor discovery framework that casts factors as Turing-complete code and improves both effectiveness and efficiency via three separations: (i) logic revision vs. parameter optimization, (ii) LLM-guided directional search vs. Bayesian hyperparameter search, and (iii) LLM usage vs. local computation. FE further incorporates a knowledge-infused bootstrapping module that transforms unstructured financial reports into executable factor programs through a closed-loop multi-agent extraction-verification-code-generation pipeline, and an experience knowledge base that supports trajectory-aware refinement (including learning from failures). Across extensive backtests on real-world OHLCV data, FE produces factors with substantially stronger predictive stability and portfolio impact-for example, higher IC/ICIR (and Rank IC/ICIR) and improved AR/Sharpe, than baseline methods, achieving state-of-the-art predictive and portfolio performance.

Sources

Related papers