Rethinking Domain Specialization for Open-Ended Scientific Reasoning in Astronomy Language Models
astro-ph.IM, cs.AI, cs.CY, cs.LG
Submitted: 2026-09-15
Updated: 2026-09-15
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models
- MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
- OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
- Measuring Massive Multitask Language Understanding
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
- AstroMLab 1: Who Wins Astronomy Jeopardy!?
- BERTScore: Evaluating Text Generation with BERT
- Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment
- Self-Preference Bias in LLM-as-a-Judge
- EAIRA: Establishing a Methodology for Evaluating AI Models as Scientific Research Assistants
- AstroMLab 4: Benchmark-Topping Performance in Astronomy Q&A with a 70B-Parameter Domain-Specialized Reasoning Model
- Investigating Data Contamination in Modern Benchmarks for Large Language Models
Related papers
- A signal dedispersion algorithm for imaging-based transient searches
- AVICA: A fully automated CASA pipeline for large volume VLBI data calibration
- Spectral Map Making with SPHEREx
- Long-Integration Magnetar Burst Observatory (LIMBO): Instrument Summary and Early FRB Rate Constraints
- Towards independent event horizon imaging of the supermassive black holes in M87 and the Milky Way
- A PINK update: Improvements to the CELEBI fast radio burst data reduction and analysis pipeline