COGTRL: Training LLMs for Scientific Discovery Assistance using Cognitive Traces via Reinforcement Learning
cs.CL
Submitted: 2026-08-31
Updated: 2026-09-06
Comments: Accepted EMNLP 2026
Code: https://github.com/shri071/CoGTRL
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large Language Models (LLMs) trained on extensive scientific research are increasingly integrated as assistants for scientific discovery.
Terminology
Abstract
Large Language Models (LLMs) trained on extensive scientific research are increasingly integrated as assistants for scientific discovery. However, most research papers omit the fine-grained cognitive process of examining constraints, failed alternatives, and iterative decisions required to achieve the desired goal. Such cognitive processes are vital for real-world scientists working toward specific goals under constraints. In this paper, we show that LLMs, when trained to produce such cognitive traces, perform better as scientific discovery assistants than when trained solely on scientific literature. We propose COGTRL, a trajectory-level reinforcement learning framework that trains LLMs to emulate cognitively grounded reasoning by jointly optimizing cognitive traces and the scientific steps produced in an interleaved manner. Across two 3B-parameter models and two scientific domains (AI and Materials Science), COGTRL improves method quality by an average of 7.85 points over comparable 3B model baselines and achieves competitive performance relative to 70B parameter models. Moreover, analysis by domain experts shows a preference for methods generated by COGTRL over the baselines.
Sources
- LitLLM: A Toolkit for Scientific Literature Review
- The Impact of Large Language Models on Scientific Discovery: a Preliminary Study using GPT-4
- Fara-7B: An Efficient Agentic Model for Computer Use
- SciBERT: A Pretrained Language Model for Scientific Text
- ChemCrow: Augmenting large-language models with chemistry tools
- SmileyLlama: Modifying Large Language Models for Directed Chemical Space Exploration
- Evaluating Large Language Models Trained on Code
- MatExpert: Decomposing Materials Discovery by Mimicking Human Experts
- SciAgents: Automating scientific discovery through multi-agent intelligent graph reasoning
- Accelerating scientific discovery with Co-Scientist
- OptAgent: Optimizing Query Rewriting for E-commerce via Multi-Agent Simulation
- Training Large Language Models to Reason in a Continuous Latent Space
- GPT-4o System Card
- LLMatDesign: Autonomous Materials Discovery with Large Language Models
- Towards Leveraging Large Language Models for Automated Medical Q&A Evaluation
- Hypothesis Generation for Materials Discovery and Design Using Goal-Driven and Constraint-Guided LLM Agents
- PaperQA: Retrieval-Augmented Generative Agent for Scientific Research
- DrugAgent: Automating AI-aided Drug Discovery Programming through LLM Multi-Agent Collaboration
- Understanding R1-Zero-Like Training: A Critical Perspective
- AlphaEvolve: A coding agent for scientific and algorithmic discovery
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering