Tree-of-Ideas: Automated Research Ideation via Cross-Trajectory Reasoning over Scholarly Evolution

arXiv:2608.10740 · cs.AI · Submitted 2026-08-11 · Read on arXiv

Xun Li, Yiying Yang, Pengtao Li, Xiao Yao, Suyu Liu, Xiaoyang Ye, Ziyu Lu, Yuan Yao, Yangning Li, Yinghui Li, Wenhao Jiang

Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ) · University of Melbourne · University of Sydney · Shenzhen University · Nanyang Technological University · Tsinghua University · Tencent Youtu Lab, Tencent

cs.AI

Submitted: 2026-08-11

Updated: 2026-08-12

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: The paper addresses a fundamental limitation in automated research ideation systems.

Terminology

Summary

The paper addresses a fundamental limitation in automated research ideation systems. As stated in the abstract: Effective research ideation requires moving beyond a static understanding of prior work to trace how research problems and solutions evolve across the literature. The authors identify that Existing methods either treat papers as unstructured context or model scholarly evolution as isolated citation chains, overlooking interactions among research trajectories.

The paper argues that "Research ideation can therefore be viewed as an evolutionary reasoning problem: a useful system should identify not only which papers are relevant, but also how problems, methods, and unresolved questions develop across the literature. Furthermore, some opportunities become visible only when multiple research trajectories are considered together."

The authors note that Retrieval-based systems provide relevant papers or summaries as context, but typically leave the relationships among them implicit and Chain-based systems organize papers into chronological or citation trajectories, but represent a branching research landscape through isolated linear paths. Even when citation graphs are available, they are often used primarily to organize or retrieve papers rather than to reason about how knowledge changes across citation relations.

The paper proposes Tree-of-Ideas (ToI), an automated research ideation framework that reasons over branching scholarly trajectories. The framework consists of three stages:

  1. EvoGraph (Stage 1): Given a research topic and a literature corpus, EvoGraph retrieves and expands relevant papers to construct a topic-centered citation graph and identify frontier papers.

  2. EvoTrace (Stage 2): EvoTrace reconstructs a local evolution tree for each frontier paper, tracing methodological transitions, unresolved research gaps, and supporting provenance within and across branches. The formal definition states: "Given a frontier paper r and the topic-centered citation graph Gq, EvoTrace constructs an annotated local evolution tree Tr = (Vr, Er, ϕ, ψ), where Vr and Er denote the retained papers and citation relations, ϕ describes the evolutionary transition associated with each edge, and ψ records the research gaps associated with each paper."

  3. EvoAgent (Stage 3): "EvoAgent extracts convergent and cross-link signals, retains directions supported by multiple trajectories, and iteratively generates, reviews, and filters ideas to produce a diverse set of literature-grounded research proposals."

The overall process is formalized as: (q, L) −−−−−→ Gq −−−−−−→ Tr −−−−−−→ I where Gq is a topic-centered citation graph, r is a frontier paper selected from the graph, and Tr is the corresponding evolution tree containing multiple research trajectories.

EvoTrace "performs relational inference over citation-linked paper pairs. Given a predecessor v and a citing paper u, it infers how u advances v, whether the gaps in v are addressed, narrowed, inherited, or transformed, and what new gaps arise. The framework models ϕ at the relation level so that the same paper can represent different advances over different predecessors."

The process works by: "Starting from (r), EvoTrace recursively follows relevant citations toward earlier work. Although the graph is traversed backward from the frontier paper, each inferred transition is oriented in the evolutionary direction from earlier to later work."

EvoAgent identifies two classes of signals:

Convergent signals include three forms:

  1. Shared gaps, where the same or related gaps persist across multiple trajectories

  2. Common actions, where different trajectories repeatedly adopt the same research action or solution pattern

  3. Frontier races, where distinct approaches pursue the same frontier objective in parallel

Cross-link signals take the form of capability–gap bridges, where a capability developed along one trajectory can potentially address an unresolved gap exposed by another.

The framework enforces a critical constraint: A direction is retained only when its evidence spans at least two distinct trajectories; otherwise, it is treated as a single-lineage continuation rather than a cross-trajectory opportunity.

All automatic methods, including ToI, use DeepSeek-V4-Flash as the backbone model. The evaluation covers six research topics with five ideas generated per topic. The literature space is constructed from the Semantic Scholar API and a pre-built citation graph of papers from major AI venues (NeurIPS, ICML, ACL, EMNLP, ICLR, CVPR, AAAI, etc.). Starting from topic-relevant frontier papers, the system recursively expand[s] their references to a maximum depth of three.

The paper compares against:

  • Direct Prompting: LLM generates ideas from the topic description alone, without literature context

  • RAG: Standard retrieval-augmented generation—paper abstracts are retrieved and concatenated as context for idea generation

  • CoI-Agent: Constructs linear research-trend chains from anchor papers, then generates ideas conditioned on chain context

  • AI-Scientist: A multi-step agentic framework that reads, summarizes, and generates ideas with self-evaluation

  • ResearchAgent: Iteratively generates and refines research ideas over scientific literature using knowledge graphs

  • Real Paper†: Core ideas extracted from representative published papers, serving as a human-level reference rather than a competing system

Ideas are evaluated on five dimensions using a 10-point scale: Novelty, Significance, Groundedness, Feasibility, and Expected Effectiveness.

ToI achieves the best overall performance among all automatic methods, with an average score of 6.27, compared with 5.36 for the strongest baseline, ResearchAgent. The results show:

  • Novelty: 6.36 (highest among automatic methods)

  • Significance: 5.71

  • Groundedness: 7.00 (highest among automatic methods)

  • Feasibility: 6.59

  • Effectiveness: 5.68

The paper notes: "The particularly strong results in Novelty and Groundedness support our central hypothesis that reconstructing scholarly evolution and reasoning across multiple trajectories enables the discovery of research ideas that are both less obvious and more firmly connected to prior work."

ToI also reaches an aggregate score comparable to the Real Paper reference (6.29). In contrast, Direct Prompting obtains a substantially lower Groundedness score (3.83), highlighting the importance of literature-grounded context for credible automated research ideation.

CoI-Agent and ToI achieve the strongest overall performance among the automatic methods, with average scores of 6.71 and 6.62, respectively. ToI ranks first in Novelty (6.78) and Groundedness (7.17), and ties for first in Significance (6.39).

The paper notes: ToI obtains a lower Feasibility score (6.28), suggesting that ideas derived from cross-trajectory synthesis may be more ambitious and therefore more difficult to operationalize immediately. At the aggregate level, ToI performs comparably to the Real Paper reference (6.55).

Performance consistently decreases as structural context is removed: full ToI (6.04) outperforms chain-only (5.85) and frontier-only (4.77). The largest reductions occur in Novelty and Groundedness, indicating that multiple evolutionary paths expose both more diverse research opportunities and stronger historical support.

The w/o Gap variant retains the same evolution-tree structure but removes the gap annotations inferred along each trajectory. This reduces the average score from 6.27 to 5.55, with consistent degradation across all five evaluation dimensions, which indicates that recovering evolutionary relations alone is insufficient for effective ideation.

Removing explicit signal discovery decreases the average score from 6.27 to 5.62. Both signal types contribute: Convergent-only and Cross-link-only reaching 5.71 and 5.81, respectively, but both remain below the full model. This demonstrates that convergent and cross-link signals capture complementary opportunities that are most effective when used jointly.

The paper traces the full ToI pipeline on the Diffusion Models topic. Starting from Simple and Critical Iterative Denoising (2025), EvoTrace reconstructs two mainline branches: one develops Critic-guided denoising, while the other introduces score-based corrector steps. Although both improve the reverse process, they retain the fixed mask-absorbing forward schedule inherited from earlier discrete diffusion models. EvoAgent identifies this as a convergent gap and proposes NF-Diff (Neural Forward Diffusion): parameterizes the forward transition kernel via a small hypernetwork conditioned on the current noisy graph, outputting per-element distributions over the discrete vocabulary.

The paper emphasizes: "The convergent gap is only visible when both mainline branches are held simultaneously; a single-chain method following either branch alone would not identify the fixed forward process as the shared structural bottleneck."

The paper highlights three main contributions:

  1. We present Tree-of-Ideas (ToI), an end-to-end framework for automated research ideation through reasoning over branching scholarly evolution

  2. We develop EvoTrace, which reconstructs gap-centered scholarly evolution trajectories from citation relations

  3. We introduce EvoAgent, which reasons over the gap-centered trajectories through two complementary signals

The paper acknowledges: "Our evaluation is conducted at the idea stage without downstream implementation or empirical validation; therefore, highly rated ideas may not necessarily lead to successful research outcomes. In addition, ToI relies on the completeness and quality of the constructed citation graph, and missing or weakly connected trajectories may limit the evolutionary evidence available for cross-trajectory reasoning."

The paper concludes: "We introduced Tree-of-Ideas (ToI), which frames automated scientific ideation as cross-trajectory reasoning over scholarly evolution. By reconstructing evolution trees and tracing gaps across multiple research trajectories, ToI uses citation structure as explicit reasoning context rather than flat retrieval evidence. Experiments across six AI topics show that ToI achieves the strongest overall performance among automatic methods, particularly in Novelty and Groundedness, while ablations confirm the value of multi-path evolutionary structure. These results demonstrate the potential of scholarly evolution as a structured basis for generating grounded and original research ideas."

Improvements for AI systems

Improvements to AI systems:

  1. Evolutionary citation graph construction (EvoGraph): Build topic-centered citation graphs with recursive reference expansion (up to depth 3) to identify frontier papers, enabling systems to reason over branching scholarly landscapes rather than flat document lists.

  2. Gap-centered evolution tracing (EvoTrace): Infer relational transitions between citation-linked papers—whether gaps are addressed, narrowed, inherited, or transformed—and annotate each edge with the evolutionary change and associated research gaps, allowing the same paper to represent different advances over different predecessors.

  3. Cross-trajectory signal extraction (EvoAgent): Detect convergent signals (shared gaps, common actions, frontier races) and cross-link signals (capability–gap bridges) across multiple evolution branches, retaining only directions supported by at least two distinct trajectories to filter out single-lineage continuations.

  4. Iterative idea generation with multi-path filtering: Generate, review, and filter research proposals conditioned on evolution trees, enforcing that each idea is grounded in evidence spanning multiple trajectories, thereby increasing novelty and groundedness simultaneously.

  5. Structured reasoning over scholarly evolution: Replace flat retrieval context with explicit evolutionary reasoning—treating citation structure as reasoning context, not just retrieval evidence—to expose opportunities visible only when multiple research trajectories are considered jointly.

What the improved AI system can do:

  • Generate research ideas with higher novelty and groundedness than chain-based or retrieval-based systems, matching human-level reference quality (ToI: 6.27 vs. Real Paper: 6.29 in model evaluation).

  • Identify shared structural bottlenecks across parallel research branches (e.g., fixed forward process in diffusion models) that single-chain methods would miss.

  • Produce ideas with strong historical support by tracing how problems and solutions evolve across the literature, reducing hallucinated or weakly grounded proposals.

  • Distinguish cross-trajectory opportunities from single-lineage continuations, ensuring proposed directions are genuinely novel syntheses rather than incremental extensions.

  • Provide provenance for each idea by linking it to specific gaps and capabilities across multiple evolution branches, improving interpretability and verifiability.

  • Adapt to new topics by automatically constructing evolution trees from citation graphs, enabling scalable ideation across diverse research domains.

Sources

Related papers