RuleWeaver: Benchmarking Rule-Centered Scenario Reasoning for Large Language Models
cs.CL
Submitted: 2026-08-27
Updated: 2026-08-27
Code: https://github.com/openai/tiktoken
Terminology
Sources
- One Mind, Many Tongues: A Deep Dive into Language-Agnostic Knowledge Neurons in Large Language Models
- Transformers as Soft Reasoners over Language
- GLM-5: from Vibe Coding to Agentic Engineering
- LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models
- FOLIO: Natural Language Reasoning with First-Order Logic
- CUAD: An Expert-Annotated NLP Dataset for Legal Contract Review
- WikiHow: A Large Scale Text Summarization Dataset
- Bias-Restrained Prefix Representation Finetuning for Mathematical Reasoning
- Learning How to Remember: A Meta-Cognitive Management Method for Structured and Transferable Agent Memory
- GPT-4 Technical Report
- LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language Models
- Multi-LogiEval: Towards Evaluating Multi-Step Logical Reasoning Ability of Large Language Models
- Large Language Models Meet Symbolic Provers for Logical Reasoning Evaluation
- Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought
- TaskBench: Benchmarking Large Language Models for Task Automation
- ProofWriter: Generating Implications, Proofs, and Abductive Statements over Natural Language
- Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
- A Survey of Large Language Models
- AR-LSAT: Investigating Analytical Reasoning of Text
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering