Policy-as-logic for robust reasoning over rules

arXiv:2608.11905 · cs.AI, cs.LG, cs.SC · Submitted 2026-08-12 · Read on arXiv

Rahul Nair, Bastian Lipka, Elizabeth Daly

IBM Research · IBM

cs.AI, cs.LG, cs.SC

Submitted: 2026-08-12

Updated: 2026-08-13

Comments: RobustifAI Workshop at IJCAI-ECAI '26

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: Policy-as-Logic for Robust Reasoning over Rules proposes a hybrid symbolic approach called Policy-as-logic (PaL) that expresses policies in formal logic and separates fact extraction using language

Terminology

Summary

Policy-as-Logic for Robust Reasoning over Rules proposes a hybrid symbolic approach called Policy-as-logic (PaL) that expresses policies in formal logic and separates fact extraction using language models from reasoning using answer set solvers. The paper states: "We present a hybrid symbolic approach that expresses policies in formal logic and at inference time exploits the representation power of language models for fact extraction to ground predicates, and an answer set solver for reasoning such that responses are interpretable, auditable, and as we show, accurate and robust under input perturbations."

The method involves four steps: semantic parsing (using an LLM to convert policy text into an Answer Set Program and a schema), extraction (translating input queries to structured JSON facts using the schema), grounding (converting facts to grounded atoms), and solving (using Clingo to determine answer sets, followed by interpretation back to decisions). The paper notes: Through the strict separation of fact extraction using LLMs and reasoning using classical solvers, we aim to improve robustness of pipelines involving policies.

Experiments cover four domains: Airline baggage fees, Income tax, NBA Transactions (from RuleArena), and HR content moderation (from PolyGuard). The first three involve objective, knowledge-based criteria, while HR involves subjective, belief-based criteria. The paper compares PaL against policy-as-prompt (0-shot and 1-shot) and policy-as-code baselines across four LLMs: GPT-OSS-120B, Qwen-2.5 72B, Llama-3.3 70B, and Granite-4.1 8B.

Key results show: On Airline, our method achieves accuracy numbers between 0.94 and 1.00 across all models, while baselines don't exceed 0.38 for policy-as-prompt and 0.40 for policy-as-code. For Tax, all baselines remain below 0.10 while PaL scores 0.31 across all models. On NBA, the pipeline outperforms baselines, though the gap is smaller. For robustness under six perturbation types, On Tax, every baseline across all four models drops to 0.00 robustness... PaL maintains robustness close to the accuracy across all four domains.

On HR, the pattern reverses: The pipeline offers no systematic advantage over the baseline. PaL and the 0-shot baseline score between 0.93 and 0.96. The paper explains: The HR rules are simple... so the solver contributes no logical reasoning beyond what the LLM does in a single forward pass.

Token efficiency is a major advantage: our method needs fewer tokens by an order of magnitude in most domains. For example, on Airline, PaL uses 1,175 total tokens per query versus 11,684 for 0-shot and 12,903 for 1-shot.

The paper concludes: "We have presented policy-as-logic formulation that had broad applicability in practice for domains where there are written rules. Preliminary evidence from a few representative domains suggests significant performance gains where automated decision pipelines leverage structured reasoning. In our experiments, gains are predominantly in domains where rules are based on knowledge rather than beliefs."

Improvements for AI systems

Improvements to AI Systems:

  1. Hybrid Neuro-Symbolic Decision Engine: Integrate a dual-component architecture where an LLM handles only fact extraction (converting unstructured input to structured JSON) and a symbolic solver (e.g., Clingo) performs all logical inference. This replaces end-to-end LLM reasoning for policy compliance, eliminating hallucinated rule application.

  2. Schema-Constrained Fact Extraction: Use the LLM to generate a domain-specific schema (predicates and types) from policy text, then force all subsequent extractions to conform to that schema. This prevents the LLM from inventing new predicates or misinterpreting entities, improving grounding accuracy.

  3. Perturbation-Resilient Reasoning: For knowledge-based domains (e.g., tax, baggage fees), the system automatically detects when input perturbations (rephrasing, numeric changes, negation) occur. Since reasoning is symbolic, the solver recomputes answers deterministically, maintaining accuracy even when the LLM's fact extraction is slightly altered—unlike pure LLM pipelines that degrade to 0% accuracy.

  4. Token-Efficient Policy Execution: Replace long-context policy prompts (thousands of tokens) with a compact Answer Set Program (ASP) representation. The LLM only processes the query and schema, reducing token usage by 10x (e.g., 1,175 vs. 11,684 tokens per query). This enables cheaper, faster deployment in high-volume decision systems.

  5. Auditable Decision Trails: The solver's answer sets provide explicit, traceable logical proofs for each decision (e.g., "denied because baggage weight > 50 lbs AND class = economy"). This allows the AI system to output human-readable justifications for every action, meeting compliance and audit requirements.

  6. Adaptive Reasoning Mode Selection: The system automatically switches between symbolic reasoning (for objective, rule-based policies) and pure LLM inference (for subjective, belief-based policies like HR moderation). This is triggered by analyzing rule complexity—if rules are simple and non-hierarchical, the solver adds no value, so the system falls back to 0-shot prompting to save compute.

  7. Cross-Model Robustness Standardization: The hybrid pipeline normalizes performance across different LLMs (from 8B to 120B parameters), achieving near-identical accuracy (0.94–1.00) regardless of the underlying model. This allows organizations to use smaller, cheaper models without sacrificing decision quality.

What the Improved AI System Can Do:

  • Process complex, multi-condition policy documents (e.g., tax codes, airline fee schedules) and return accurate, legally compliant decisions with 95–100% accuracy, even under adversarial rephrasing or numeric perturbations.

  • Generate a complete, auditable reasoning chain for every decision, showing exactly which rule and facts led to the outcome.

  • Handle 10x more queries per dollar by using minimal tokens per request, while maintaining deterministic correctness.

  • Automatically adapt to new policy domains by generating a fresh logic program from the policy text, without retraining the LLM.

  • Operate reliably across heterogeneous LLM backends, making it deployable in resource-constrained environments (e.g., on-premise with an 8B model) without performance loss.

Abstract

In many practical applications of generative AI systems, from tax rules to airline baggage allowance, responses to natural language queries must respect written policies or rules. We present a hybrid symbolic approach that expresses policies in formal logic and at inference time exploits the representation power of language models for fact extraction to ground predicates, and an answer set solver for reasoning such that responses are interpretable, auditable, and as we show, accurate and robust under input perturbations. Specifically, we show this separation of extraction and reasoning steps outperforms policy-as-prompt and policy-as-code methods in most cases with 10x reduction in token usage. The results point to the value of structured reasoning and symbolic solvers in conjunction with generative models to make robust decisions involving objective criteria.

Sources

Related papers