SCICONVBENCH: Benchmarking LLMs on Multi-Turn Clarification for Task Formulation in Computational Science
cs.AI, physics.comp-ph
Submitted: 2026-05-18
Updated: 2026-09-17
Code: https://github.com/csml-rpi/SciConvBench
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- MetaOpenFOAM: an LLM-based multi-agent framework for CFD
- ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
- User Simulation with Large Language Models for Evaluating Task-Oriented Dialogue
- Fine-tuning a Large Language Model for Automating Computational Fluid Dynamics Simulations
- Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
- ClarQ-LLM: A Benchmark for Models Clarifying and Requesting Information in Task-Oriented Dialog
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- MirrorBench: A Benchmark to Evaluate Conversational User-Proxy Agents for Human-Likeness
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- Aligning Language Models to Explicitly Handle Ambiguity
- MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models
- LLMs Get Lost In Multi-Turn Conversation
- ContraDoc: Understanding Self-Contradictions in Documents with Large Language Models
- From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
- MatTools: Benchmarking Large Language Models for Materials Science Tools
- AgentBench: Evaluating LLMs as Agents
- SciAgent: Tool-augmented Language Models for Scientific Reasoning
- GAIA: a benchmark for General AI Assistants
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection