EVAR: Evidence-Validated Hypothesis Admission for Budget-Aware Narrative Reasoning
cs.CL
Submitted: 2026-08-30
Updated: 2026-08-30
Comments: Accepted to the Main Conference of EMNLP 2026. 16 pages, 3 figures
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Lost in the Middle: How Language Models Use Long Contexts
- True Detective: A Deep Abductive Reasoning Benchmark Undoable for GPT-3 and Challenging for GPT-4
- S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning
- Faith and Fate: Limits of Transformers on Compositionality
- Self-Refine: Iterative Refinement with Self-Feedback
- QuALITY: Question Answering with Long Input Texts, Yes!
- Did Aristotle Use a Laptop? A Question Answering Benchmark with Implicit Reasoning Strategies
- Reflexion: Language Agents with Verbal Reinforcement Learning
- CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing
- Beyond Survival: Evaluating LLMs in Social Deduction Games with Human-Aligned Strategies
- Large Language Models Cannot Self-Correct Reasoning Yet
- GPT-4 Doesn't Know It's Wrong: An Analysis of Iterative Prompting for Reasoning Problems
- Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
- PLAYER*: Enhancing LLM-based Multi-Agent Communication and Interaction in Murder Mystery Games
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting
- Avalon's Game of Thoughts: Battle Against Deception through Recursive Contemplation
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering
- Tree of Thoughts: Deliberate Problem Solving with Large Language Models
- Exploring Collaboration Mechanisms for LLM Agents: A Social Psychology View
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering