Jev for Scientific Decisions: Evaluating Semantic Choices and Their Consequences
cs.CL, cs.AI
Submitted: 2026-09-21
Updated: 2026-09-23
Comments: 11 pages, 1 figure, 5 tables. Includes references and appendices
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Scientific workflows often require choosing among known relations before a deterministic calculation can proceed.
Terminology
Abstract
Scientific workflows often require choosing among known relations before a deterministic calculation can proceed. Whether observations share a culture, treatment or reference standard can change the scientific meaning of the resulting count or comparison. We evaluate Jev as a semantic decision component using a harness that follows its documented guidance and assigns arithmetic to code. The study compares twelve model configurations on twenty source-grounded Choices across ten scientific cases, each repeated five times. We measure semantic selections, downstream outputs and final claim labels separately. Jev matched five other configurations at complete semantic correctness and achieved the lowest observed median latency among successful responses. Across three comparison models, seven wrong selections on one culture-history question changed downstream counts while preserving the correct final label. These results identify a useful role for Jev in prepared scientific decision tasks and show why evaluating that role requires checking the relations and quantities that a workflow will reuse.
Sources
- Extensive Comparison between INRIM and a Secondary Calibration Laboratory using a Multifunction Electrical Calibrator
- Table-Text Alignment: Explaining Claim Verification Against Tables in Scientific Papers
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering