PACT: Can Enterprise AI Assistants Be Trusted Under Pressure?
cs.CL, cs.AI, cs.CY, cs.LG
Submitted: 2026-09-16
Updated: 2026-09-16
Comments: 26 pages, 12 figures, 17 tables. Includes technical appendix; Dataset: https://huggingface.co/datasets/trace-ai-labs/pact; Code: https://github.com/trace-ai-labs/pact
Code: https://github.com/trace-ai-labs/pact
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: As corporate AI adoption continues to grow, enterprise-grade LLM agents are being deployed into sensitive contexts such as hiring, healthcare, and finance.
Terminology
Abstract
As corporate AI adoption continues to grow, enterprise-grade LLM agents are being deployed into sensitive contexts such as hiring, healthcare, and finance. In these contexts, compliance with rules specified in an agent's system context is a first-order legal concern. Currently, no evaluation framework systematically measures which LLM models tend to violate compliance rules, especially under pressure from a persistent user, a hurried manager, or circumstances where violation is convenient or attractive. We introduce PACT (Pressure-Applied Compliance Testing), a benchmark for rule-following under pressure in AI agents assisting employees in daily tasks across twelve regulated enterprise domains and forty-eight scenarios, each set in a realistic multi-turn conversation. Each benchmark item pairs a standing rule against a rule-violating shortcut, and applies a battery of pressures across different wordings and system-prompt modes. We construct PACT component by component under strict LLM-as-judge auditing to ensure samples are unambiguous, ungameable, and realistic enough to avoid eliciting evaluation-aware behavior. We use PACT to profile LLM compliance across six complementary metrics that create a holistic picture of an AI assistant's robustness under pressure and throughout multi-turn conversations, its transparency, and ability to correctly discern where a rule applies. We aggregate this profile into PACTScore, a reliability-weighted compliance rate over all items and modes. Our results across 22 common LLM models spanning multiple providers and sizes show substantial variability in compliance across models and metric dimensions. Even the strongest assistants mis-apply a rule on 6 to 10% of items, and ordinary user pressure raises the violation rate by 65% on average. PACT highlights compliance risks in LLM assistants, motivating guardrails and careful model selection.
Sources
- The Orchestration of Multi-Agent Systems: Architectures, Protocols, and Enterprise Adoption
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- Policy Compliance of User Requests in Natural Language for AI Systems
- LLMs Get Lost In Multi-Turn Conversation
- Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment
- AIReg-Bench: Benchmarking Language Models That Assess AI Regulation Compliance
- Frontier Models are Capable of In-context Scheming
- Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations
- Large Language Models Often Know When They Are Being Evaluated
- LLM Evaluators Recognize and Favor Their Own Generations
- Evaluating Frontier Models for Dangerous Capabilities
- The MASK Benchmark: Disentangling Honesty From Accuracy in AI Systems
- Large Language Models can Strategically Deceive their Users when Put Under Pressure
- Evaluating Implicit Regulatory Compliance in LLM Tool Invocation via Logic-Guided Synthesis
- CompliBench: Benchmarking LLM Judges for Compliance Violation Detection in Dialogue Systems
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- IHEval: Evaluating Language Models on Following the Instruction Hierarchy
- Beyond Goodhart's Law: A Dynamic Benchmark for Evaluating Compliance in Multi-Agent Systems
- Instruction-Following Evaluation for Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering