Plan Pointers and Record-Directive Form in Budgeted Verification of Inherited Agent Memory
cs.IR, cs.AI, cs.CL
Submitted: 2026-09-03
Updated: 2026-09-21
Terminology
Sources
- SafeCommit: Certifying When Memory-Grounded Agents May Safely Act
- Inference-Time Budget Control for LLM Search Agents
- To Copy or Not to Copy: Copying Is Easier to Induce Than Recall
- PlanGuard: Defending Agents against Indirect Prompt Injection via Planning-based Consistency Verification
- Decision-Aware Memory Cards: Counterfactual-Inspired Context Selection and Compression for Tool-Using LLM Agents
- LLMs can be easily Confused by Instructional Distractions
- Failing to Falsify: Evaluating and Mitigating Confirmation Bias in Language Models
- The Task Shield: Enforcing Task Alignment to Defend Against Indirect Prompt Injection in LLM Agents
- IH-Benchmark: A Conflict-Centered Benchmark for Instruction-Hierarchy Robustness in LLM Applications
- Will the Agent Recuse, and Will It Stop? Measuring LLM-Agent Compliance with In-Band Governance Signals at the Access Door and Mid-Flight
- When Stale Constraints Go Unchecked: Budgeted Verification Failures in Inherited Agent Memory
- Why Do AI Agents Break Rules? How Framing, Context, and Social Signals Shape Compliance
- Blind Refusal: Language Models Refuse to Help Users Evade Unjust, Absurd, and Illegitimate Rules
- Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components
- Ask the World Before Acting: Environment Probing for Calibrated Agent World Models
- Compliance vs. Sensibility: On the Reasoning Controllability in Large Language Models
- From Lossy to Verified: A Provenance-Aware Tiered Memory for Agents
Related papers
- The Price of Isolation: Estimating the Ecosystem Cost of Symmetric Two-Sided A/B Testing
- SCAR: Semantic Continuity-Aware Retrieval for Efficient Context Expansion in RAG
- MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora
- RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation
- Right Family, Wrong Skill: Evaluating Risk Exposure in Agent Skill Retrieval
- UltRAG: a Universal Simple Scalable Recipe for Knowledge Graph RAG