AquiLLM: Evaluating Faithfulness in Open-Weight RAG-LLM Systems for Scientific Research
cs.AI
Submitted: 2026-09-15
Updated: 2026-09-26
Comments: 14 pages, 1 figure
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: Scientific research increasingly relies on large, heterogeneous data sources, motivating interest in retrieval-augmented generation (RAG) systems that provide natural language access to scientific
Terminology
Abstract
Scientific research increasingly relies on large, heterogeneous data sources, motivating interest in retrieval-augmented generation (RAG) systems that provide natural language access to scientific knowledge and research workflows. Researchers are exploring the viability of these systems as natural language interfaces for document search and for generating analysis code and pipeline components. At the same time, concerns about data privacy and control over research infrastructure have motivated interest in open-weight models and open-source deployments hosted within research institutions. In astronomy, this development follows a long history of computational infrastructure development, from archival databases and SQL-based systems to LLM-assisted research tools. This paper presents a domain-expert evaluation of faithfulness for AquiLLM, an open-weight, offline RAG-LLM platform designed to support scientific research groups in the use and preservation of tacit and formal knowledge. We define faithfulness as the extent to which generated responses remain grounded in retrieved scientific context without unsupported claims or omissions. We report results from an astronomy case study evaluating AquiLLM across retrieval and scientific analysis tasks. AquiLLM performs most reliably on explicit retrieval-oriented questions grounded in the RAG collection, while faithfulness degrades for queries requiring synthesis or ambiguity resolution. These results highlight both the promise and limitations of open-weight RAG-LLM systems for scientific research and demonstrate the importance of domain-expert evaluation beyond standard benchmark leaderboards.
Sources
- Natural Language Interfaces to Databases - An Introduction
- Language agents achieve superhuman synthesis of scientific knowledge
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- AquiLLM: a RAG Tool for Capturing Tacit Knowledge in Research Groups
- Holistic Evaluation of Language Models
- A review of faithfulness metrics for hallucination assessment in Large Language Models
- The World Wide Telescope: An Archetype for Online Science
- The SDSS SkyServer, Public Access to the Sloan Digital Sky Server Data
- LLM4SR: A Survey on Large Language Models for Scientific Research
- AstroLLaMA: Towards Specialized Foundation Models in Astronomy
- Large language models in materials science and the need for open-source approaches
- Reason and Verify: A Framework for Faithful Retrieval-Augmented Generation
- Benchmarking LLM Faithfulness in RAG with Evolving Leaderboards
- Qwen3 Technical Report
- AquiLLM: An Architecture for Supporting Tacit Knowledge Capture in Research Groups
- Euclid. I. Overview of the Euclid mission
- The Hyper Suprime-Cam SSP Survey: Overview and Survey Design
- Elements of effective machine learning datasets in astronomy
- Photometric Redshifts for Cosmology: Improving Accuracy and Uncertainty Estimates Using Bayesian Neural Networks
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection