Security papers — 2026-10-05
The focus today is on making large language models harder to re-identify when they are used by autonomous agents, which is important because if we cannot track who is using these powerful systems, it complicates security and accountability. Researchers looked at methods like Constant-Rate Certified Deletion to anonymize LLM outputs against agentic re-identification while keeping useful model performance.
A related effort explored the fragility of trigger-tag mechanisms used for misuse detection in open-weight models, showing that these methods are easily bypassed by adversarial prompts. This vulnerability connects directly to another piece of work that investigated passing the test on prompt injection detectors for LLM agents, suggesting a need for more robust ways to verify agent behavior.
We also examined techniques for improving model robustness against input sequence variations, which is important because agents often feed slightly altered inputs into these systems. Furthermore, there is research into RMCW, a deletion-robust watermark based on Reed-Muller codes designed specifically for language models to embed traceable information securely. This work builds upon the idea of embedding data but focuses on making that embedding resilient to deletion attacks.
Finally, we looked at speculative decoding and prefix scheduling as a way to improve the efficiency of generating text, which is a technical refinement that supports the overall goal of deploying these models more reliably.
The work on Extended Differential Cryptanalysis of Kuznyechik is the most pressing because it directly challenges the security assumptions underpinning current cryptographic primitives. Researchers attempted to extend existing differential cryptanalytic techniques to this specific system, and what emerged suggests a more robust path toward identifying weaknesses in its structure. This finding builds upon earlier work that explored how these differential paths interact within the protocol's state transitions.
Then there is the research into mitigating watermark forgery in generative models, which is important because it addresses a tangible security risk in deploying AI systems. The study involved introducing randomized key selection into the generation process to make it harder for attackers to forge watermarks. This approach seems promising, though further testing is needed to confirm its efficacy across different model architectures.
Another piece of work focuses on adaptive quantum-safe cryptography for 6G vehicular networks, which matters because securing future communication infrastructure against quantum threats is paramount. The researchers optimized the cryptographic parameters based on context within the network environment to enhance security while maintaining performance. This optimization effort connects to the threat modeling work concerning emerging AI-agent protocols, as both seek to build resilient systems for future deployments.
The comparative analysis of security threats for emerging AI-agent protocols, looking at MCP, A2A, Agora, and ANP provides a necessary framework for understanding where vulnerabilities lie in these new agent interactions. This analysis helps guide the development of more secure communication standards. Hop-Decayed Influence examines new vulnerabilities in GraphRAG pipelines when LLMs are involved in structural auxiliary indexing.
Finally, X-NegoBox presents an explainable privacy-budget negotiation framework for secure peer-to-peer energy data exchange, which is significant for decentralized systems. This framework allows participants to manage their privacy levels explicitly during data sharing. This concept of managing information flow is related to the intent-hiding jailbreaks research, which uses information theory to analyze compositional attacks on agent protocols.
The most pressing work involves establishing information equivalence across different privacy accounting frameworks, which is crucial because without it, we cannot reliably measure the true cost of data usage in complex systems. This effort builds upon earlier explorations into mitigating private data leakage within large language models by introducing a whiteout mechanism that attempts to obscure sensitive training data during inference.
A significant development is the creation of SideKernel, a usable microVM sandbox specifically designed for AI coding agents running on macOS, which provides a safe environment for these autonomous systems to operate without risking the host system. This work connects directly to how we might later analyze hardware Trojans using CITADEL, which aims to find malicious insertions in devices powered by large language models.
Furthermore, research into SoK stablecoins in the quantum era suggests that current cryptographic standards will require substantial updates as quantum computing capabilities mature, impacting how we secure digital assets. This is complemented by work on Pincer, which establishes resource authorization for agents by utilizing a digital twin to manage their access rights effectively.
Finally, practical security enhancements are being refined through work moving from TS-SUF-2 to TS-SUF-4 for FROST2 threshold signatures, offering tangible improvements in the security protocols used for those specific cryptographic operations.
The most crucial development today involves building a defense-in-depth framework and reference architecture for securing autonomous AI agents running on Kubernetes, because as these agents become more capable, their security posture directly impacts system integrity. We explored AgentTrap, which is designed to counter stateful feedback deception used against autonomous penetration testing agents by introducing a mechanism to detect when an agent is being tricked into believing a successful test run. This work builds upon the foundational concept of securing computer-use agents against branch steering attacks, which addresses how malicious actors can manipulate the agent's decision-making paths.
Next, we looked at digital twin-assisted mapping of industrial control system telemetry to the ATT&CK for ICS framework, which is important because it allows us to map real operational data directly to known adversarial techniques. This approach uses evidence-driven dependency reasoning to figure out how different components in an ICS environment rely on each other, providing a clearer picture of potential attack vectors. This mapping effort connects with research on security-aware dependency analysis for LLM agents, which seeks to move beyond simple predefined sinks by analyzing the actual dependencies within a language model agent's operational flow.
We also examined EvoRiskBench, an evolving benchmark for runtime security risks in workspace agents, which is significant because it provides a standardized way to test how these agents behave under various real-world security pressures. This benchmark helps us understand the emergent risks that appear when these agents operate outside of controlled environments. Finally, LiBRA addresses image watermark removal through detection-aware image watermark removal via bidirectional latent optimization, which is a specialized technique for handling data integrity issues within agent training or deployment pipelines.
The most pressing development concerns the defense framework for agentic unmanned aerial vehicle swarms, which addresses the critical need to ensure these systems operate safely by focusing on the perception-reasoning interface. This work introduced a defense-in-depth strategy specifically targeting vulnerabilities at that interface. This framework builds upon existing ideas by focusing on persona guardrails, creating a production-grade defense mechanism for agentic systems. It aims to control how agents behave in real operational environments.
CorrectGuard provides eyes-off correctness estimation for black-box security guardrails, which is vital because it allows us to assess the reliability of these defenses without needing access to the internal workings of the system. This feeds into understanding threat-preserving representation sensitivity in agent-security benchmarks, which explores how agents react when their representations are deliberately manipulated.
PrivDev maps static analysis data types to a domain-specific policy verification language, which is a foundational step for building secure agentic systems. This mapping helps define what data structures are permissible within the system's operational constraints.
Finally, PoCoFL introduces policy-compliant federated learning, which suggests a way to train models collaboratively while strictly adhering to specific policies across different agents. This method complements the security guardrail work by ensuring that collective learning remains compliant with predefined rules.
Today's papers
- LLM Anonymization Against Agentic Re-Identification This paper proposes a method to protect LLMs from being re-identified when used by autonomous agents. [paper] [episode]
- From Positionwise Confidence to Prefix Scheduling: Verifier Skipping in Speculative Decoding This work improves speculative decoding by skipping verifiers based on positionwise confidence. [paper] [episode]
- Evaluating and Improving the Robustness of Large Language Models to Input Sequence Variations This paper systematically evaluates how robust large language models are against changes in their input sequences. [paper]
- RMCW: A Deletion-Robust Watermark Based on Reed--Muller Codes for Language Models This research introduces a watermark method that is resistant to deletion attacks for language models. [paper]
- SecJev: Bringing Security Expertise to System One Decision Models This paper brings security knowledge into the decision-making process of System One models. [paper]
- The Fragility of Trigger-Tag Mechanisms for Misuse Detection in Open-Weight LLMs This study examines why trigger-tag methods used to detect misuse in open language models are easily bypassed. [paper]
- Passing the Test You Trained On: Re-evaluating Prompt-Injection Detectors for LLM Agents This paper retests prompt injection detectors to see if they still work effectively against LLM agents. [paper]
- Constant-Rate Certified Deletion This paper provides a method for deletion that guarantees a constant rate of information loss. [paper]
- Extended Differential Cryptanalysis of Kuznyechik This research uses extended differential cryptanalysis to analyze the security properties of Kuznyechik. [paper] [episode]
- Mitigating Watermark Forgery in Generative Models via Randomized Key Selection This work suggests using randomized keys to prevent the forgery of watermarks in generative models. [paper] [episode]
- Adaptive Quantum-Safe Cryptography for 6G Vehicular Networks via Context-Aware Optimization This paper develops quantum-safe cryptography optimized for 6G vehicular networks based on context. [paper] [episode]
- Security Threat Modeling for Emerging AI-Agent Protocols: A Comparative Analysis of MCP, A2A, Agora, and ANP This paper compares different threat models for various AI agent protocols. [paper] [episode]
- X-NegoBox: An Explainable Privacy-Budget Negotiation Framework for Secure Peer-to-Peer Energy Data Exchange This work presents a framework that allows users to negotiate privacy budgets in peer-to-peer energy data exchanges. [paper] [episode]
- Intent-Hiding Jailbreaks: An Information-Theoretic Framework for Compositional Attacks This paper uses information theory to create a framework for understanding and preventing compositional attacks that hide intent. [paper]
- MIRROR: Multipath Quorum Integrity for LLM Multi-Agent Communication This research introduces a system to ensure the integrity of communication among multiple LLM agents using multipath quorum. [paper]
- Hop-Decayed Influence: New Vulnerabilities of Structural Auxiliary Indexing in GraphRAG Pipelines with LLM This paper identifies new vulnerabilities in graph retrieval augmented generation pipelines that use structural auxiliary indexing with LLMs. [paper]
- Unifying Privacy Accounting: Information Equivalence and Information Loss This paper unifies the concepts of information equivalence and information loss for privacy accounting. [paper]
- Mitigating Private Data Leakage in LLMs with Whiteout This work proposes a technique called whiteout to reduce private data leakage from language models. [paper]
- SoK: Stablecoins in the Quantum Era This paper discusses the future implications of stablecoins as quantum computing becomes available. [paper]
- SideKernel: A Usable microVM Sandbox for AI Coding Agents on macOS This paper describes a usable virtual machine sandbox for AI coding agents running on macOS. [paper]
- CITADEL: CWE-Guided Insertion of Hardware Trojans via Analysis of DFG-Enabled LLMs This work analyzes how hardware Trojans can be inserted into devices using DFG-enabled LLMs based on CWEs. [paper]
- Out of Sync, Out of Sight: Phantom State Attacks against IIoT Intrusion Detection This paper explores phantom state attacks that can bypass intrusion detection systems in Industrial Internet of Things. [paper]
- Pincer: Resource Authorization for Agents using a Digital Twin This paper proposes a resource authorization method for agents that leverages a digital twin. [paper]
- From TS-SUF-2 to TS-SUF-4: Practical Security Enhancements for FROST2 Threshold Signatures This paper details practical security improvements applied to FROST2 threshold signatures from version 2 to version 4. [paper]
- Containing the Autonomous Operator: A Defense-in-Depth Framework and Reference Architecture for Securing AI Agents on Kubernetes This paper presents a defense-in-depth framework for securing AI agents running on Kubernetes. [paper]
- AgentTrap: Stateful Feedback Deception against Autonomous Penetration Testing Agents This work introduces AgentTrap to counter stateful feedback deception from autonomous penetration testing agents. [paper]
- Digital Twin-Assisted Mapping of ICS Telemetry to ATT&CK for ICS with Evidence-Driven Dependency Reasoning This paper uses a digital twin to map industrial control system telemetry to the MITRE ATT&CK framework. [paper]
- Beyond Predefined Sinks: Security-Aware Dependency Analysis for LLM Agents This research focuses on security-aware dependency analysis for LLM agents that goes beyond predefined sinks. [paper]
- Securing Computer-Use Agents Against Branch Steering Attacks This paper discusses how to secure computer-use agents against branch steering attacks. [paper]
- EvoRiskBench: An Evolving Benchmark for Runtime Security Risks in Workspace Agents This paper introduces EvoRiskBench as an evolving benchmark for runtime security risks in workspace agents. [paper]
- LiBRA: Detection-Aware Image Watermark Removal via Bidirectional Latent Optimization This work proposes a method to remove image watermarks that is aware of detection mechanisms using bidirectional latent optimization. [paper]
- Reversing the Clock: Layout-Aware Recovery of Design Intent from Clock Distribution Networks This paper describes how to recover design intent from clock distribution networks by being layout-aware. [paper]
- Asymptotic Analysis of Trading Fees in CFMM This paper provides an asymptotic analysis of trading fees within the CFMM framework. [paper]
- Defense-in-Depth at the Perception-Reasoning Interface of LLM-Centric Agentic UAV Swarms This work proposes a defense strategy for the perception and reasoning interfaces of LLM agent swarms on UAVs. [paper]
- Persona Guardrail: A Production-Grade Defense Framework for Agentic Systems This paper presents a production-grade defense framework designed to guard the persona of agentic systems. [paper]
- CorrectGuard: Eyes-Off Correctness Estimation for Black-Box Security Guardrails This work develops CorrectGuard to estimate correctness without needing access to the internal workings of black-box security guardrails. [paper]
- PrivDev: Mapping Static-Analysis Data Types to DPV This paper maps static analysis data types to a dynamic privacy verification system. [paper]
- A Secure dToF LiDAR SoC with Dual-Domain Fingerprinting and Event-Driven AFE Circuit Achieving Sensor-Level Attack Resilience This paper describes a secure Time of Flight LiDAR system using dual-domain fingerprinting for sensor resilience. [paper]
- Threat-Preserving Representation Sensitivity in Agent-Security Benchmarks This paper investigates how to create benchmarks that preserve threat information when evaluating agent security. [paper]
- PoCoFL: POlicy-COmpliant Federated Learning This work introduces Policy-Compliant Federated Learning to ensure privacy policies are followed during the learning process. [paper]
The papers
- Security Threat Modeling for Emerging AI-Agent Protocols: A Comparative Analysis of MCP, A2A, Agora, and ANP — This paper presents a systematic security analysis of four emerging AI agent communication protocols—Model Context Protocol (MCP), Agent2Agent (A2A), Agora, and Agent Network Protocol (ANP)—to establish a protocol-centric risk assessment framework for secure deployment. [episode]
- X-NegoBox: An Explainable Privacy-Budget Negotiation Framework for Secure Peer-to-Peer Energy Data Exchange — The X-NegoBox framework introduces an explainable negotiation system designed to manage adaptive differential privacy budgets during secure peer-to-peer energy data exchange, addressing limitations in existing static privacy policies by enabling transparent, context-aware decisio [episode]
- LLM Anonymization Against Agentic Re-Identification — Agentic LLMs with web search change the anonymization problem because rich contextual details can become cross-referenceable evidence, yet those same details often carry significant downstream analytic value. [episode]
- Adaptive Quantum-Safe Cryptography for 6G Vehicular Networks via Context-Aware Optimization — Powerful quantum computers may be able to break communication security for vehicles in 6G networks, necessitating new post-quantum cryptography methods that often introduce latency challenges. [episode]
- Extended Differential Cryptanalysis of Kuznyechik — This research introduces an inner c-differential cryptanalysis technique to analyze block ciphers, addressing structural limitations that previously prevented practical application of c-differential uniformity in real-world scenarios. [episode]
- From Positionwise Confidence to Prefix Scheduling: Verifier Skipping in Speculative Decoding — Speculative diffusion decoding (SDD) can be optimized by introducing verifier skipping, a lossy policy that commits a selected draft prefix directly to save verification costs. [episode]
- Mitigating Watermark Forgery in Generative Models via Randomized Key Selection — Watermarking enables GenAI providers to verify whether content was generated by their models, and this work proposes a defense against forgery attacks by randomizing key selection for each query, which provably resists forgery independent of the number of watermarked samples coll [episode]
- SoK: Stablecoins in the Quantum Era —
- SideKernel: A Usable microVM Sandbox for AI Coding Agents on macOS —
- CITADEL: CWE-Guided Insertion of Hardware Trojans via Analysis of DFG-Enabled LLMs —
- Out of Sync, Out of Sight: Phantom State Attacks against IIoT Intrusion Detection —
- Pincer: Resource Authorization for Agents using a Digital Twin —
- From TS-SUF-2 to TS-SUF-4: Practical Security Enhancements for FROST2 Threshold Signatures —
- RMCW: A Deletion-Robust Watermark Based on Reed--Muller Codes for Language Models —
- Containing the Autonomous Operator: A Defense-in-Depth Framework and Reference Architecture for Securing AI Agents on Kubernetes —
- AgentTrap: Stateful Feedback Deception against Autonomous Penetration Testing Agents —
- Digital Twin-Assisted Mapping of ICS Telemetry to ATT&CK for ICS with Evidence-Driven Dependency Reasoning —
- Beyond Predefined Sinks: Security-Aware Dependency Analysis for LLM Agents —
- SecJev: Bringing Security Expertise to System One Decision Models —
- Securing Computer-Use Agents Against Branch Steering Attacks —
- The Fragility of Trigger-Tag Mechanisms for Misuse Detection in Open-Weight LLMs —
- EvoRiskBench: An Evolving Benchmark for Runtime Security Risks in Workspace Agents —
- LiBRA: Detection-Aware Image Watermark Removal via Bidirectional Latent Optimization —
- Reversing the Clock: Layout-Aware Recovery of Design Intent from Clock Distribution Networks —
- Asymptotic Analysis of Trading Fees in CFMM —
- Defense-in-Depth at the Perception-Reasoning Interface of LLM-Centric Agentic UAV Swarms —
- Persona Guardrail: A Production-Grade Defense Framework for Agentic Systems —
- Passing the Test You Trained On: Re-evaluating Prompt-Injection Detectors for LLM Agents —
- CorrectGuard: Eyes-Off Correctness Estimation for Black-Box Security Guardrails —
- PrivDev: Mapping Static-Analysis Data Types to DPV —
- A Secure dToF LiDAR SoC with Dual-Domain Fingerprinting and Event-Driven AFE Circuit Achieving Sensor-Level Attack Resilience —
- Threat-Preserving Representation Sensitivity in Agent-Security Benchmarks —
- Constant-Rate Certified Deletion —
- PoCoFL: POlicy-COmpliant Federated Learning —
- Intent-Hiding Jailbreaks: An Information-Theoretic Framework for Compositional Attacks —
- MIRROR: Multipath Quorum Integrity for LLM Multi-Agent Communication —
- Hop-Decayed Influence: New Vulnerabilities of Structural Auxiliary Indexing in GraphRAG Pipelines with LLM —
- Unifying Privacy Accounting: Information Equivalence and Information Loss —
- Mitigating Private Data Leakage in LLMs with Whiteout —
- Evaluating and Improving the Robustness of Large Language Models to Input Sequence Variations —
Important terms
- Constant-Rate Certified Deletion
- A method used to anonymize large language model outputs against re-identification attacks while still maintaining good model performance. It helps keep users private when using these powerful AI systems.
- RMCW Watermark
- This is a deletion-robust watermark based on Reed-Muller codes specifically for language models. It securely embeds traceable information so that it survives deletion attempts.
- AgentTrap
- A defense mechanism designed to detect when autonomous agents are being tricked into believing a test run was successful. This counters stateful feedback deception used against testing agents.
- Defense-in-depth Framework
- A comprehensive security architecture for securing AI agents running on Kubernetes. It layers multiple security measures to ensure the integrity of the entire system.