Does Finetuning with Scientific Data Increase Hallucinations? A Multi-domain Factuality Evaluation of LLMs
cs.CL
Submitted: 2026-06-19
Updated: 2026-08-28
Code: https://github.com/ryabhmd/SciFactCheck
Terminology
Sources
- LLMs as Science Journalists: Supporting Early-stage Researchers in Communicating Their Science to the Public
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Transforming Science with Large Language Models: A Survey on AI-assisted Scientific Discovery, Experimentation, Content Generation, and Evaluation
- Qwen2.5-Coder Technical Report
- The 17% Gap: Quantifying Epistemic Decay in AI-Assisted Survey Papers
- Camels in a Changing Climate: Enhancing LM Adaptation with Tulu 2
- OpenFActScore: Open-Source Atomic Evaluation of Factuality in Text Generation
- Docling: An Efficient Open-Source Toolkit for AI-driven Document Conversion
- The Curious Case of Factuality Finetuning: Models' Internal Beliefs Can Improve Factuality
- HalluCitation Matters: Revealing the Impact of Hallucinated References with 300 Hallucinated Papers in ACL Conferences
- MinerU: An Open-Source Solution for Precise Document Content Extraction
- Qwen3 Technical Report
- WildHallucinations: Evaluating Long-form Factuality in LLMs with Real-World Entity Queries
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering