Bioinfoysis Technical Report
cs.AI, cs.MA
Submitted: 2026-09-03
Updated: 2026-09-13
Project page: https://report.bioinfoysis.com
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language model agents have shown promise in bioinformatics, but most existing systems focus primarily on producing final answers, treating planning, tool use, and code execution as transient
Terminology
Abstract
Large language model agents have shown promise in bioinformatics, but most existing systems focus primarily on producing final answers, treating planning, tool use, and code execution as transient interactions. This design is poorly suited to long-horizon bioinformatics tasks, where conclusions must remain connected to the data, computations, and intermediate evidence that support them. We introduce Bioinfoysis, a multi-agent harness that represents each request as a persistent, artifact-grounded analysis run. Bioinfoysis combines global planning with step-wise, evidence-driven replanning: the planner maintains an executable checklist and revises pending steps using structured handoffs returned after each worker execution. These handoffs bind intermediate results to their responsible agent, checklist step, and plan generation, preventing stale evidence from being silently reused after replanning. A controlled runtime validates generated scripts, tables, and figures before they are used in downstream analysis or reporting, while role-specific context, persistent memory, and governed bioinformatics skills support reliable execution over long analysis trajectories. We evaluate Bioinfoysis on BixBench and two question-answering tracks of LAB-Bench 2. On BixBench, Bioinfoysis achieves state-of-the-art accuracy of 82.4%. Across four underlying language models, Bioinfoysis increases average accuracy from 27.81% to 64.13% on SeqQA2 and from 3.13% to 31.25% on DbQA2. These results demonstrate that reliable bioinformatics automation depends not only on model capability, but also on the harness that governs planning, execution, memory, and evidence flow. We hope that the emergence of Bioinfoysis will play a driving and leading role in the development of the bioinformatics community. Our demo website can be seen in https://report.bioinfoysis.com/.
Sources
- AI for Auto-Research: Roadmap & User Guide
- Qwen3 Technical Report
- OpenAI GPT-5 System Card
- SignBot: Learning Human-to-Humanoid Sign Language Interaction
- aiXiv: A Next-Generation Open Access Ecosystem for Scientific Discovery Generated by AI Scientists
- BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments
- BixBench: a Comprehensive Benchmark for LLM-based Agents in Computational Biology
- BioAgent Bench: An AI Agent Evaluation Suite for Bioinformatics
- Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models
- Phantom Guardrails: When Self-Improving Agent Harnesses Fix Failures That Never Happened
- ReAct: Synergizing Reasoning and Acting in Language Models
- When Safe Skills Collide: Measuring Compositional Risk in Agent Skill Ecosystems
- SP-Mind: An Autonomous Reasoning Agent for Spatial Proteomics Analysis
- Understanding the Challenges in Iterative Generative Optimization with LLMs
- MemMark: State-Evolution Attribution Watermarking for Agent Long-Term Memory Systems
- LAB-Bench: Measuring Capabilities of Language Models for Biology Research
- Rethinking the AI Scientist: Interactive Multi-Agent Workflows for Scientific Discovery
- K-Dense Analyst: Towards Fully Automated Scientific Analysis
- Embedding Foundation Model Predictions in Discrete-Choice Models with Structural Guarantees
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection