Nomad: Autonomous Exploration and Discovery
cs.AI
Submitted: 2026-03-31
Updated: 2026-08-27
License: http://creativecommons.org/licenses/by/4.0/
The gist: We introduce Nomad, a system for autonomous data exploration and insight discovery.
Terminology
Abstract
We introduce Nomad, a system for autonomous data exploration and insight discovery. Given a corpus of documents, databases, or other data sources, users rarely know the full set of questions, hypotheses, or connections that could be explored. As a result, query-driven question answering and prompt-driven deep-research systems remain limited by human framing and often fail to cover the broader insight space. Nomad addresses this problem with an exploration-first architecture. It constructs an explicit Exploration Map over the domain and systematically traverses it to balance breadth and depth. It generates and selects hypotheses and investigates them with an explorer agent that can use document search, web search, and database tools. Candidate insights are then checked by an independent verifier before entering a reporting pipeline that produces cited reports and higher-level meta-reports. We also present a comprehensive evaluation framework for autonomous discovery systems that measures trustworthiness, report quality, and diversity. Using corpora of selected UN and WHO reports and arXiv papers on LLM agents, we show that Nomad produces reports with strong numeric grounding, higher overall quality and actionability than baselines, and more diverse insights over several runs. Nomad is a step toward autonomous systems that not only answer user questions or conduct directed research, but also discover which questions, research directions, and insights are worth surfacing in the first place.
Sources
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection
- LongWriter: Unleashing 10,000+ Word Generation from Long Context LLMs
- From Local to Global: A Graph RAG Approach to Query-Focused Summarization
- Understanding DeepResearch via Reports
- Robin: A multi-agent system for automating scientific discovery
- HippoRAG: Neurobiologically Inspired Long-Term Memory for Large Language Models
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- WebDancer: Towards Autonomous Information Seeking Agency
- LongCite: Enabling LLMs to Generate Fine-grained Citations in Long-context QA
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection