From Atomic Evidence to Logical Composition: Structured Compositional Reasoning over Compound Answer Options
Obed Junias, Maria Leonor Pacheco
University of Colorado Boulder
cs.CL, cs.AI
Submitted: 2026-08-13
Updated: 2026-08-14
Comments: 21 pages, 6 figures, 10 tables
Code: https://github.com/obedjunias19/structured-compositional-reasoning
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: Large language models (LLMs) perform well across a wide range of tasks but exhibit persistent weaknesses in logical reasoning.
Terminology
Summary
Large language models (LLMs) perform well across a wide range of tasks but exhibit persistent weaknesses in logical reasoning. The paper identifies a specific failure mode: models often fail when answer options require combining atomic judgments under explicit logical operators, even when they judge the individual atoms correctly. This produces a compositionality gap, where a model solves the component subproblems correctly yet fails to combine them.
The paper studies compound answer options connected by AND, OR, and NEITHER/NOR. The authors note that difficulty tracks the operator rather than the content, suggesting the issue stems from how logical possibilities are represented and combined, not from missing knowledge. This pattern aligns with mental-model theories of human reasoning, which predict performance ordering from strongest on conjunction to weakest on negated compositions.
The paper presents a structured framework that decomposes each compound option into atomic answers and scores contrastive hypotheses about each one, so the model never sees a compound option as a whole. The framework consists of several stages:
Option Decomposition: Each compound option is deterministically parsed into a triplet (a1, ◦, a2), and atomic answers appearing anywhere in the instance are collected into a set UC. Atoms shared across options are scored once, so a proposition receives a single judgment wherever it appears.
Contrastive Hypothesis Construction: For each atomic answer, the framework constructs a pair of opposing natural language hypotheses: h+ stating the answer satisfies the context, and h− stating it does not. This makes the result a comparison between two readings of the same atom.
Confidence Elicitation: The positive and negative hypotheses are presented as choices A and B within a single prompt, and the model selects which is more plausible. The raw evidence scores are normalized log probabilities over the two alternatives. Paired multiple-choice provides the strongest evidence compared to alternatives including independent true–false scoring, generation sampling, and verbalized confidence.
Score Calibration: The paper evaluates two standard post-hoc calibration methods (Platt scaling and isotonic calibration) and introduces a novel relative calibration
method. Relative calibration supplies both a score's magnitude and its standing within the instance, using features including the within-instance standardized score, rank among atomic scores, and distance from the highest positive score. A logistic regression maps these features to calibrated scores.
Globally Constrained Inference: An operator-constrained integer linear program (ILP) combines the calibrated scores into a single prediction. Binary variables represent inferred statuses of atomic answers and validity of compound options. Linear inequalities encode the operator semantics exactly:
-
AND: both atoms must hold
-
OR: at least one atom must hold
-
NEITHER/NOR: neither atom may hold
The objective maximizes total evidence for the atomic assignment, and exactly one compound option is selected.
The framework is evaluated on two benchmarks:
LOGICAL-COMMONSENSEQA: A commonsense reasoning benchmark with 19,996 instances (11,996 train / 6,000 dev / 2,000 test), evenly distributed across four settings: AND, OR, NEITHER/NOR, and MIXED. The test set is divided into human-validated (HV) and non-validated (NV) subsets of 1,000 each.
LOGICAL-SATA: A new reading-comprehension benchmark constructed from SATA-Bench, containing 5,400 instances (2,400 train / 1,000 dev / 2,000 test). It pairs independently annotated answers into compound options of the same form as LOGICAL-COMMONSENSEQA.
Structured inference substantially outperforms direct prompting across both benchmarks:
On LOGICAL-COMMONSENSEQA-HV, the strongest direct-prompting configuration achieves 48.3 macro-F1, while paired multiple-choice evidence with globally constrained inference achieves 75.8, an improvement of 27.5 points. Relative calibration further increases performance to 77.0.
On LOGICAL-SATA, paired multiple-choice structured inference obtains 72.2 macro-F1 compared to 47.0 for the strongest direct-prompting baseline, with relative calibration improving to 75.6.
The largest gains occur on NEITHER/NOR: macro-F1 rises from 14.0 to 76.8 on LOGICAL-COMMONSENSEQA and from 12.6 to 73.4 on LOGICAL-SATA. On AND, gains are smaller: 70.8 to 72.4 on LOGICAL-COMMONSENSEQA and 70.9 to 73.6 on LOGICAL-SATA.
Calibration improves both atomic evidence reliability and downstream compound prediction. Relative calibration performs best on both Brier score and log loss metrics. On LOGICAL-SATA, relative calibration achieves Brier score 0.1449 and log loss 0.4438, compared to uncalibrated 0.1896 and 0.8141. On LOGICAL-COMMONSENSEQA-HV, relative calibration achieves Brier score 0.1464 and log loss 0.4567.
Relative calibration's largest downstream gains occur in the MIXED setting, where Macro-F1 increases by 4.9 on LOGICAL-COMMONSENSEQA and 11.2 points on LOGICAL-SATA. This is because in MIXED instances, options impose opposing demands—AND and OR require atoms to be accepted while NEITHER/NOR requires rejection—so systematic bias favors one operator over another.
When the inference layer is provided with gold atomic statuses, accuracy reaches 1.00 on both benchmarks, confirming the ILP encodes operator semantics exactly. Atomic accuracy is 0.830 on LOGICAL-COMMONSENSEQA-HV and 0.824 on LOGICAL-SATA, against compound accuracy of 0.758 and 0.723.
Qualitative analysis reveals different error sources: On LOGICAL-COMMONSENSEQA, errors often involve broad interpretations of open-ended commonsense questions or insufficient attention to modifiers. On LOGICAL-SATA, the model often selects the label matching a passage's main topic while rejecting other labels that also apply. The logical operators determine how errors propagate: OR can tolerate an incorrect atomic judgment when another atomic answer remains supported, whereas AND and NEITHER/NOR can be invalidated by a single incorrect assignment.
The paper contributes: (1) a structured framework eliciting contrastive evidence for individual atomic answers combined through operator-constrained ILP inference; (2) a relative calibration method scoring each atom by both confidence and standing among other atoms; (3) LOGICAL-SATA, a new reading-comprehension benchmark for compound answer reasoning; (4) evaluation across two benchmarks showing largest gains on operators that degrade most under standard prompting.
The results suggest that some failures on logical reasoning tasks reflect difficulties in composing local judgments rather than only absence of relevant knowledge. Future work could extend the framework to options with more than two atomic answers, structures such as implication and exclusive disjunction, nested expressions, and less structured text. A further direction is replacing inference with a probabilistic formulation allowing atomic uncertainty to propagate to compound predictions.
Improvements for AI systems
Improvements to AI systems:
-
Decompose compound options into atomic judgments before answering. Instead of prompting the model to evaluate a compound statement like
A AND B
as a whole, the system first extracts each atomic proposition (A, B) and scores them independently via contrastive hypothesis pairs (e.g.,A is true
vs.A is false
). This eliminates the compositionality gap where the model knows the atoms but fails to combine them. -
Use paired multiple-choice evidence elicitation rather than free-form generation or single true/false scoring. For each atomic answer, present two opposing natural-language hypotheses as options A and B in a single prompt, and take the normalized log-probability over the two choices. This yields more reliable evidence than independent scoring or verbalized confidence.
-
Apply relative calibration per instance, not just global calibration. For each atomic score, feed features like its within-instance z-score, rank among all atomic scores in the instance, and distance from the highest positive score into a logistic regression to produce a calibrated probability. This corrects systematic bias where some operators (e.g., NEITHER/NOR) are systematically under- or over-scored relative to others.
-
Enforce logical consistency via operator-constrained integer linear programming (ILP). After calibrating atomic scores, use binary variables for each atom’s inferred truth and each compound option’s validity, with linear constraints encoding AND (both atoms true), OR (at least one true), and NEITHER/NOR (both false). Maximize total evidence subject to exactly one option being selected. This guarantees the final prediction is logically coherent, even if individual atom scores are noisy.
-
Share atomic evidence across options. When the same atomic answer appears in multiple compound options, score it once and reuse that score in all constraints. This prevents redundant or conflicting judgments for the same proposition.
-
Handle negated compositions explicitly. For NEITHER/NOR options, construct hypotheses that state the atom does not satisfy the context, and use the ILP to require both atoms be false. This directly addresses the operator where standard prompting fails most (e.g., macro-F1 rising from 14 to 77).
What the improved AI system can do:
-
Answer multiple-choice questions with compound options (AND, OR, NEITHER/NOR) with high accuracy, even when the model fails on the same questions under direct prompting. For example, on LOGICAL-COMMONSENSEQA, macro-F1 improves from 48.3 to 77.0; on LOGICAL-SATA, from 47.0 to 75.6.
-
Reason robustly on negated logical operators, where it previously performed near chance (14.0 macro-F1 on NEITHER/NOR) and now achieves 76.8, because it decomposes the negation into atomic rejections and combines them with exact logical constraints.
-
Provide calibrated confidence scores for each atomic judgment, enabling downstream systems to know not just what is true but how likely it is, relative to other atoms in the same instance. This reduces overconfidence and improves decision-making in mixed-operator settings (e.g., +11.2 macro-F1 on MIXED in LOGICAL-SATA).
-
Guarantee logical consistency in predictions—the final answer always respects the operator semantics (verified by 100% accuracy when given gold atomic statuses), so the system never outputs an option that violates AND/OR/NEITHER logic.
-
Transfer to new reading-comprehension benchmarks without retraining, as demonstrated on the newly constructed LOGICAL-SATA, where the same framework yields large gains over direct prompting.
-
Scale to more complex logical structures in the future, such as implications, exclusive disjunction, nested expressions, and multi-atom compounds, by extending the ILP constraints and decomposition rules.
Abstract
Large language models often fail when answer options require combining atomic judgments under explicit logical operators, even when they judge the individual atoms correctly. We study compound options connected by AND, OR, and NEITHER/NOR, introducing a framework that decomposes each option into atomic answers and scores contrastive hypotheses about each one, so the model never sees a compound option. An operator-constrained integer linear program then composes the calibrated scores into a single prediction. We evaluate on LOGICAL-COMMONSENSEQA and introduce LOGICAL-SATA, a reading-comprehension benchmark derived from SATA-Bench. Our framework improves Macro-F1 from 48.3 to 77.0 on the human-validated LOGICAL-COMMONSENSEQA split and from 47.0 to 75.6 on LOGICAL-SATA, with the largest gains on NEITHER/NOR.
Sources
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Measuring Massive Multitask Language Understanding
- Decomposed Prompting: A Modular Approach for Solving Complex Tasks
- Language Models (Mostly) Know What They Know
- Training language models to follow instructions with human feedback
- SATA-BENCH: Select All That Apply Benchmark for Multiple Choice Questions
- ReClor: A Reading Comprehension Dataset Requiring Logical Reasoning
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering