OmniScientist: An Omni-Modal Omni-Discipline AI Scientist

arXiv:2608.13558 · cs.AI, cs.CL · Submitted 2026-08-13 · Read on arXiv

Bobo Li, Hao Fei, Tianjie Ju, Mong-Li Lee, Wynne Hsu

National University of Singapore · University of Oxford

cs.AI, cs.CL

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: 30 pages, 13 figures, 19 tables. Project page: https://omni-scientist.github.io/

Code: https://github.com/Omni-Scientist/OmniScientist

Project page: https://omni-scientist.github.io

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 92/100

The gist: The paper introduces OmniScientist, described as "an end-to-end, omni-modal AI scientist that conducts multidisciplinary research directly from heterogeneous raw evidence." The system is designed to

Terminology

Summary

The paper introduces OmniScientist, described as an end-to-end, omni-modal AI scientist that conducts multidisciplinary research directly from heterogeneous raw evidence. The system is designed to address a critical gap in existing AI scientist systems: Existing systems are increasingly workflow-complete, while remaining evidence-incomplete. The authors argue that current AI-scientist pipelines typically expose data through text, code, labels, or summaries prepared in advance, which means the agent therefore inherits a human-chosen representation before inquiry begins.

The paper identifies that Scientific evidence reaches researchers in forms with markedly different internal structure and organizes scientific evidence into 4 discipline-independent families:

  1. Perceptual - Images, video, micrographs, radar, astronomical and remote-sensing imagery, the visual form of scientific plots, audio, and 3-D structure

  2. Symbolic - Natural-language documents, formulae, variables, rules, sequences, knowledge graphs, logical and causal relations, mathematical models

  3. Quantitative-statistical - Tables, measurements, distributions, curves, correlations, significance tests, regression results

  4. Procedural / dynamic - Experimental steps, code execution, agent traces, simulations, dynamic evolution, protocols

The system consists of a perception layer and 3 autonomous agents for ideation, experiment, and writeup operating within a deterministic pipeline. The architecture allows observations to shape research questions, experimental decisions, and final claims throughout the research lifecycle.

The perception layer supplies this grounding by organizing observations hierarchically. Artifacts are categorized into evidence families, and specific modalities define exact representations. The framework prioritizes native numeric analysis by extracting key properties, such as FFT peaks or trend points, directly from the raw artifacts and invokes visual rendering... only when spatial or structural patterns are essential.

The ideation stage requires the agent to formulate a concrete, novel, and falsifiable question that can be answered computationally from the supplied data. It uses a ReAct loop to establish grounding through an inventory of the materials, contextualize these findings by searching the literature through OpenAlex, and develop at least 5 candidate ideas.

During experimentation, the agent autonomously translates the finalized idea into a methodological design and implements it through iterative code generation. The system uses a controlled run python environment that manages subprocess execution and figure capture. The agent incorporates a comprehensive suite of at least 4 analyses into the experimental design, combining a main hypothesis test with essential controls such as baselines, ablation studies, mechanism probes, or sensitivity sweeps.

The writeup stage carries 5 structural specifications that fix the skeleton and the length of each venue style - machine learning, biomedical, earth & space, physics, and chemistry formats. Each section is expanded only from the slice of the structured experiment record it needs.

The system implements three deterministic checks enforced in code:

  1. Idea Check: validates structural completeness by requiring a clear research question, hypothesis, experiment sketch, and falsification criterion and verifies the thoroughness of the generation process by checking for the 5 self-filtered candidates and at least 3 focused literature searches.

  2. Rigour Check: grounds the experiment in reality by confirming that the agent genuinely accessed the requested dataset and generated figures matching the raw execution trace and enforces a strict multiple-comparison correction that accounts for every test attempted during the debugging loop.

  3. Claim Check: ensures every number in the drafted text is matched against the grounded set derived from the experiment record.

The system was evaluated on a 36-case demonstration suite spanning 5 discipline families, all 4 evidence families, and modalities ranging from images and waveforms to audio, video, 3-D structures, trajectories, tables, formulae, and graphs. The suite includes datasets such as EuroSAT, Galaxy Zoo, STEAD, Kather CRC, Chest X-ray, PlantVillage, SemanticKITTI, and PDEBench.

  • The system completes the full path from raw data to a compiled manuscript in all 36 cases

  • achieves a mean overall paper score of 6.3 with the reference reasoning backbone (Claude Sonnet 5)

  • paper quality remains broadly consistent across evidence modalities

  • factual accuracy consistently leads the evaluation metrics across all reasoning engines

The paper compares the full system against a blind variant that receives precomputed scalar features and never accesses the raw observation. Results show:

  • Full perception improves every evaluated dimension

  • the largest gains in multimodal grounding and scientific significance

  • its papers win 85% of the head-to-head judgments

  • The gains appear in the scientific substance of the papers: the questions selected, the analyses executed, and the claims supported by the resulting evidence

The system was tested across 9 reasoning backbones including Claude Sonnet 5, GPT-5.6, GLM-5.2, Kimi K2.7, Qwen3.5 (9B, 27B, 122B), and Gemma-4 (26B, 31B). The strongest alternate backbones, such as GLM and Kimi, fall within a remarkably similar performance range.

The system audited roughly 1,500 three-component broadband seismograms from the STEAD catalogue and identified a clear onset-and-decay envelope within a trace explicitly labelled as noise. This led to the finding that 21.7% (163 of 750) of the noise-labelled traces carry coherent, polarised, cross-component transient bursts, bounded by a 95% confidence interval of [18.8, 24.9].

The system observed that pneumonic lung fields were not uniformly brighter but unevenly mottled, with dense patches sitting beside clear ones. It converted this into the spatial dispersion of a sliding-window local Shannon-entropy map and found "Spatial heterogeneity separates normal from pneumonic lung fields with a large effect size (Cohen's d > 1.2) and achieved a peak area under the curve (AUC) of 0.851 on the held-out set."

  1. Lifecycle-wide perception is essential: direct perception improves all 7 evaluation dimensions and wins 85% of head-to-head judgments

  2. Cross-disciplinary applicability: Adding a discipline requires a new specification file, with no change to the core pipeline and no domain-specific research code

  3. Robustness across backbones: End-to-end generation achieves high quality and is consistent across top LLM backbones

  4. Provenance enforcement: factual accuracy consistently ranks at the top across all models, attributed to the shared provenance requirement, which holds each reasoning engine to the same rigorous evidence standards

  5. Perception broadens research questions: the perception-driven system anchors its research questions on attributes exclusive to the raw multimodal records while the blind baseline invariably restricts its hypothesis formulation to precomputed scalar features

The paper concludes that OmniScientist establishes a foundational blueprint for future AI scientists and the broader development of automated empirical research.

Improvements for AI systems

Improvements to AI Systems:

  1. Native multimodal evidence grounding: Replace text-only or precomputed-feature inputs with a hierarchical perception layer that ingests raw perceptual, symbolic, quantitative, and procedural artifacts directly, extracting numeric properties (e.g., FFT peaks, trend points) before resorting to visual rendering.

  2. Autonomous research-question formulation from raw data: Enable the agent to inventory raw materials, search literature via APIs (e.g., OpenAlex), and generate ≥5 falsifiable candidate hypotheses grounded in observed artifact attributes, rather than inheriting human-chosen representations.

  3. Code-enforced scientific rigor: Implement deterministic checks that (a) validate structural completeness of ideas (question, hypothesis, experiment sketch, falsification criterion), (b) confirm genuine dataset access and figure-execution trace matching, and (c) apply strict multiple-comparison corrections accounting for every test attempted during debugging.

  4. Provenance-locked claim generation: Force every numeric claim in the final manuscript to be matched against a grounded set derived from the experiment record, eliminating hallucinated statistics.

  5. Discipline-agnostic pipeline via specification files: Allow new scientific domains to be added by supplying a venue-style specification file (skeleton, section lengths, formatting) without altering core pipeline code or adding domain-specific research logic.

  6. Iterative experiment design with mandatory controls: Require the agent to autonomously design experiments including a main hypothesis test plus ≥4 controls (baselines, ablations, mechanism probes, sensitivity sweeps) within a controlled Python execution environment.

  7. Perception-driven hypothesis broadening: Use raw multimodal observations to anchor research questions on attributes invisible to scalar-feature baselines (e.g., spatial heterogeneity, waveform envelopes, cross-component correlations), expanding the hypothesis space beyond precomputed features.


What the Improved AI System Can Do:

  • Conduct end-to-end research from raw heterogeneous data (images, waveforms, video, 3-D structures, tables, graphs, code traces) to a compiled, venue-formatted manuscript with no human preprocessing.

  • Automatically audit large scientific datasets (e.g., 1,500 seismograms) and discover previously missed artifacts (e.g., coherent signals in noise-labelled traces) with quantified confidence intervals.

  • Translate visual observations into novel quantitative features (e.g., local Shannon-entropy maps for radiographs) and validate them with large effect sizes and held-out AUC scores.

  • Maintain consistent paper quality across diverse evidence modalities and reasoning backbones (from 9B-parameter models to frontier LLMs), with factual accuracy as the top-ranked metric.

  • Add new scientific disciplines by simply writing a specification file, enabling rapid deployment across fields like seismology, radiology, astronomy, and physics without code changes.

  • Guarantee that all claims, figures, and statistics trace back to actual executed experiments, making outputs auditable and reproducible.

Sources

Related papers