From Alignment to Access Control: A Framework for GenAI Policy Enforcement
cs.CR, cs.AI
Submitted: 2026-09-22
Updated: 2026-09-22
Code: https://github.com/ibm-granite/granite.trust.policy-tools
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Concrete Problems in AI Safety
- LCGuard: Latent Communication Guard for Safe KV Sharing in Multi-Agent Systems
- Constitutional AI: Harmlessness from AI Feedback
- ShieldAgent: Shielding Agents via Verifiable Safety Policy Reasoning
- ELEPHANT: Measuring and understanding social sycophancy in LLMs
- Improved Supervised Fine-Tuning for Large Language Models to Mitigate Catastrophic Forgetting
- Activated LoRA: Fine-tuned LLMs for Intrinsics
- Cognitive Firewall: A Proactive, Zero-Trust, Multi-Gate Framework for LLM Safety
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Agent Meltdowns: The Road to Hell Is Paved with Helpful Agents
- Towards Robust Legal Reasoning: Harnessing Logical LLMs in Law
- SOPBench: Evaluating Language Agents at Following Standard Operating Procedures and Constraints
- Rethinking Machine Unlearning for Large Language Models
- IH-Benchmark: A Conflict-Centered Benchmark for Instruction-Hierarchy Robustness in LLM Applications
- Granite Guardian
- Measuring Agents in Production
- Steering Llama 2 via Contrastive Activation Addition
- FragFuse: Bypassing Access Control of Large Language Model Agents via Memory-Based Query Fragmentation and Fusion
- GAVEL: Towards Rule-Based Safety Through Activation Monitoring
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs