Security papers — 2026-09-29
The work that matters most is developing server-enforced watermarking within U-shaped split federated learning setups. This technique embeds invisible markers directly into model updates during training, allowing later verification if AI-generated content came from a specific source.
This concept treats watermarking as a monitoring primitive instead of an afterthought. Researchers are examining how this works with other agentic systems, specifically Proteus, which is designed to be a self-evolving red team for agent skill ecosystems. This helps determine if these agents can bypass security assumptions when operating autonomously.
Another significant piece of research addresses the growing problem of synthetic media misinformation and the difficulty in detecting it as AI-generated multimodal content gains traction. Furthermore, investigators are looking into the privacy risks in patient-facing medical AI systems where Retrieval Augmented Generation or RAG chatbots expose backend vulnerabilities.
Finally, a large-scale benchmark is assessing the security landscape of large cloud language model services by checking how easily traffic analysis attacks can expose sensitive information. This work connects the need for source tracking with real-world risks posed by autonomous agents and data leakage in critical applications.
The most critical development today involves SkillDRE, which systematically tests how agent skills can be evolved through both pre-execution and runtime feedback. This matters because it shows a pathway for adversarial manipulation of an agent's capabilities before and during its actual operation.
SkillDRE explores this by using a dual-stage red-team evolution process to probe skill sets. It specifically examines how an agent's performance changes when it receives feedback both before starting a task and while the task is running, which reveals hidden vulnerabilities in the skill acquisition process.
CyberClear provides a benchmark for assessing LLM agent systems against advanced persistent threat attack chains by focusing on provenance tracking. Understanding where an attack comes from allows defenders to build better defenses against sophisticated threats.
Hearsay investigates the trustworthiness of records generated by deployed agents, asking whether an auditor can rely on what the agent actually writes when it performs a task. This directly impacts how we can verify the integrity of automated decision-making processes.
REFINE introduces a resilient framework for intelligent enterprise alert triage within security operations centers, aiming to improve how security teams handle incoming alerts. This work is significant because it focuses on making the triage process robust against unexpected or malicious inputs.
Trust the Brand, Lose Control looks at how identity hijacks the orchestration of LLM agents. This is a key concern because if an attacker gains control over an agent's identity, they can potentially redirect its intended actions.
Ask Without Telling examines a method where local small language models consult cloud large language models without exposing the actual task intent to the cloud model. This technique offers a way to leverage powerful external intelligence while maintaining some level of operational privacy.
Learning to Refer addresses privacy concerns by employing client-resolved generation for language models, which helps ensure that generated content respects user boundaries. This method is crucial for deploying these models in sensitive environments where data leakage is a major risk.
The most pressing work involves understanding how agents can leak sensitive information through their browser usage, which exposes user behavior in a way that traditional security measures might miss. A study on AgentTell explored this by measuring behavioral side-channel leakage in browser-use agents, showing they can reveal information about the underlying system or user actions. This suggests a new vector for covert data exfiltration and connects to checking leakage witnesses versus certifying bounded non-leakage.
Another significant piece of research addresses the reliability of detection methods when an agent's memory is being maliciously poisoned, specifically looking at retrieval observability bounds on provenance detection. Researchers measured how well these detectors cover different poisoning scenarios and found that a standalone detector often fails to provide accurate results. This means we need better ways to verify if an agent's memory is trustworthy, which relates to exploring LLMs for attack investigations.
Then there is the work on evasion attacks targeting cost-utility-based adversarial training for online AutoML in IoT networks. This demonstrates how attackers can bypass security measures designed to make these systems robust. This shows that even well-trained models can be tricked into making suboptimal or insecure decisions when facing targeted manipulation, contrasting with DegreeSpar which focuses on structured degree sparsity for efficient secure transformer inference.
The most significant development today involves TokenScanner, which aims to detect backdoors and discover triggers within text-to-image low-rank adaptations by performing a full vocabulary scan. This is important because it addresses growing security concerns surrounding generative AI models where hidden vulnerabilities could be exploited.
This work builds upon the concept of residual transferability in neural image watermarking, which explores how much information from one image can be transferred to another through a watermark. Understanding this transfer is key to measuring inference exposure. Furthermore, E3C presents a tool for evaluating communication and computation costs in authentication and key exchange protocols, offering practical metrics for assessing the efficiency of secure exchanges.
TokenScanner's deep vocabulary scan complements this by looking specifically at the textual prompts used to generate images, tying into how information might leak through different generative pathways. Similarly, Armadillo introduces robust single-server secure aggregation for federated learning. This is a vital step toward building more trustworthy decentralized machine learning systems and relies on input validation to maintain security while allowing model training across distributed data without centralizing sensitive information.
The most critical piece of work today involved developing a simulation study to attribute sensor deviations in oilfield digital twins to potential causes like degradation, weather, or direct attack. This matters because accurately diagnosing the source of a fault is essential for maintaining operational integrity and safety in these complex systems. The researchers explored probabilistic attribution methods within this framework to test how different failure modes influence the likelihood of a specific deviation being caused by wear versus an external event.
Following that, there was work on hardware-rooted physical unclonable functions for device-level traceability in knowledge distillation. This technique uses hardware randomness to create unique fingerprints for devices, which is important for ensuring that distilled models retain verifiable lineage back to their original physical components. This contrasts with the simulation work by focusing on device identity rather than environmental or operational fault diagnosis.
Another area of focus was a compact shielded CSV, a lightweight client-side validation blockchain designed to be post-quantum secure and private. This addresses security concerns by providing a decentralized ledger for validating data locally, which is significant given the increasing threat landscape. This contrasts with the traceability work by focusing on secure data handling rather than model provenance.
The research also touched upon VulContextBench, which serves as a benchmark for retrieving security context in coding agents. This tool helps evaluate how well these agents can understand and utilize necessary security information when performing tasks, linking conceptually to the simulation study's need to correctly interpret system states.
Finally, there was a neurophysiological framework examining how deepfakes exploit cognitive engagement and implicit visual evaluation. This work is significant because it moves beyond technical detection methods to look at the human vulnerability exploited by synthetic media, providing a different kind of context for understanding digital threats compared to infrastructure-focused studies.
The most significant development today involves understanding how AI orchestration at the expression layer can be exploited. Researchers explored weird machine compositors, which are systems that combine different computational elements to create novel outputs; they found ways these compositors can be manipulated to produce unintended results. This manipulation is important because it opens avenues for subtle control over complex AI behaviors.
A related piece of work looked at API secrets and how they interact with large language models, analyzing the threat of API credentials becoming part of the LLM's vocabulary. They empirically evaluated a vault-mediated execution boundary to see if this handling could be mitigated, suggesting a path toward safer agent systems. This contrasts with other security concerns, such as those in provenance-based intrusion detection where auditing evaluation is key to identifying intrusions based on data lineage.
Further down the line, there was work on verifiable credentials used for privacy-preserving federated analytics. This means developing methods to analyze shared data without exposing individual user information, building upon the need for secure handling discussed earlier when considering API secrets and LLM interaction. Separately, research into application agnostic side-channel emanations from FPGA clock distribution networks examined how hardware itself can leak information during computation.
Finally, there is a study on HESP, which separates what an alert triage agent should probe from when it should stop probing within a local LLM environment. This addresses the practical deployment of these AI systems by providing guardrails for their operation.
The most significant piece of work today involves COGNIT-Guard because it tackles the critical need for reliable decision making in autonomous systems by implementing calibrated standalone guardrails. This system uses a heterogeneous CPU and NPU confidence cascading mechanism to handle explicit latency and false-positive constraints, which is vital when these systems are operating in real-time environments.
SecProbe addresses agent security decisions by adaptively evaluating coding agents against known cybersecurity vulnerabilities. This work moves beyond simple checks by assessing how well these agents perform when encountering actual exploits. Following that, the research on Carpet-Bombing detection focuses on using per-packet uniformity testing to detect this type of attack. This is important because it offers a way to catch malicious flooding before it overwhelms the system.
Another area of focus is evaluating System One models for agent security decisions by examining their reliability and calibration when making selective automation choices. This helps us understand how trustworthy these models are when they are tasked with making high-stakes automated judgments, contrasting with the work on individual-level unlearning in vision-language models, which deals with forgetting specific personal data from these large models.
The paper on anytime-valid leakage detection on ML-KEM EM traces is also relevant because it provides a method for detecting subtle information leakage during cryptographic operations. This is closely related to optimizing and securing the modern watermarking channel for images, as both explore ways to ensure integrity or detect unauthorized access within data transmission pathways. Finally, the poster on ProofWeave aims for a privacy-minimised evidence plane anchored by continuous agentic assurance, suggesting a future direction for providing verifiable guarantees in complex agentic workflows.
The most pressing development is the traffic analysis attack against Introduction Protocol and Onion Services, which demonstrates a vulnerability where network traffic patterns reveal sensitive information about the underlying services. This finding is significant because it directly impacts the security of decentralized communication methods by showing how metadata can be exploited.
This work builds upon earlier concerns regarding agentic security auditing, specifically examining how to maintain continuous assurance for those auditors at software delivery decision gates. The research suggests that sustained participation in these channels is a key mechanism for ensuring robust security oversight within communities.
Another critical area involves the implementation of data diodes using commodity hardware and open source software. This provides a physical enforcement layer for data flow control, which is important because it offers a tangible way to isolate systems and prevent unauthorized outbound communication.
Furthermore, the study on SoK cryptocurrency mixing and anonymity details the architectures, threat models, and operational aspects related to cryptocurrency mixing services. This contributes to understanding the security implications of privacy-enhancing technologies in digital finance.
The research into dithered Gaussian mechanisms for randomness-efficient differential privacy addresses how to introduce controlled noise into systems while maintaining a level of privacy. This is relevant because it offers a practical way to balance utility and anonymity in data processing pipelines.
Finally, physics-attested federated learning focuses on securing collaborative anomaly detection within critical water infrastructure by leveraging physical laws for verification. This work is important because it moves beyond purely mathematical security assurances to incorporate verifiable physical constraints into machine learning models.
Today's papers
- Server-Enforced Watermarking in U-Shaped Split Federated Learning This paper proposes a way to mark data securely during federated learning to prevent misuse. [paper]
- The Synthetic Media Shift: Tracking the Rise, Virality, and Detectability of AI-Generated Multimodal Misinformation This research tracks how much fake media is being made by AI and how easy it is to spot. [paper]
- When RAG Chatbots Expose Their Backend: An Anonymized Case Study of Privacy and Security Risks in Patient-Facing Medical AI This study looks at the privacy dangers when medical chatbots reveal their underlying systems. [paper]
- Proteus: A Self-Evolving Red Team for Agent Skill Ecosystems This paper introduces a self-improving red team to test the skills of AI agents. [paper]
- Watermarking Should Be Treated as a Monitoring Primitive This paper argues that watermarking should be viewed as a way to monitor data rather than just protect it. [paper]
- Who Owns This Agent? Tracing AI Agents Back to Their Owners This work focuses on how to track the original owner of an AI agent. [paper]
- The End of Trust: How Agentic AI Breaks Security Assumptions This paper explores how autonomous agents can break traditional security assumptions in systems. [paper]
- A Large-Scale Benchmark and Risk Assessment of Traffic Analysis Attacks on Cloud LLM Services This study benchmarks and assesses the risks of traffic analysis attacks against cloud large language models. [paper]
- SilentCall: Hidden Tool-Call Backdoors in Open-Weight Agents, and How to Catch Them This paper reveals hidden backdoors in open AI agents that use tools. [paper]
- SkillDRE: Dual-Stage Red-Team Evolution of Agent Skills via Pre-Execution and Runtime Feedback This framework uses two stages to evolve the skills of an agent through testing before and during its operation. [paper]
- CyberClear: A Benchmark for LLM Agent Systems on APT Attack Chain Provenance This paper provides a benchmark to test how well LLM agents can detect advanced persistent threat attack chains. [paper]
- Hearsay: Can an Auditor Trust the Record a Deployed Agent Harness Writes? This research questions whether an auditor can trust the records created by deployed AI agents. [paper]
- REFINE: A Resilient Evolution Framework for Intelligent Enterprise Alert Triage in Security Operations Centers This paper presents a framework for evolving security alert triage systems in enterprise security operations centers. [paper]
- Trust the Brand, Lose Control: How Identity Hijacks LLM Agent Orchestration This study examines how identity hijacking can lead to a loss of control over an AI agent's orchestration. [paper]
- Ask Without Telling: Local SLMs Consult Cloud LLMs Without Revealing Task Intent This paper shows how local small language models can use cloud models without revealing the specific task they are performing. [paper]
- Learning to Refer: Client-Resolved Generation for Privacy-Aware Language Models This work describes a method for making language models generate text while keeping user data private by resolving things on the client side. [paper]
- AgentTell: Behavioural Side-Channel Leakage in Browser-Use Agents This paper investigates how agents that use web browsers can leak information through their behavior. [paper]
- Digital Agriculture Sandbox for Collaborative Research This paper proposes a safe environment for researchers to collaborate on digital agriculture projects. [paper]
- Retrieval Observability Bounds on Provenance Detection for Agent Memory Poisoning: Measured Coverage and a Falsified Standalone Detector This study measures how well we can detect memory poisoning in agents by observing retrieval systems. [paper]
- Evasion Attacks on Cost-Utility-Based Adversarial Training for Online AutoML in IoT Networks This paper looks at ways to evade security training when using cost-utility methods for online machine learning on IoT devices. [paper]
- DegreeSpar: Structured Degree Sparsity for Efficient Secure Transformer Inference This technique uses structured sparsity to make secure transformer inference faster and more efficient. [paper]
- Silent Failures in Agentic Security Evaluation: A Validated Harness for Tool-Call Mediation Under Indirect Prompt Injection This paper provides a test harness to find silent failures when agents use tools under indirect prompt injection attacks. [paper]
- Benchmarking and Exploring the Capabilities of LLMs for Attack Investigations This research benchmarks how well large language models can be used to investigate security threats. [paper]
- Checking Leakage Witnesses versus Certifying Bounded Non-Leakage This paper compares different methods for checking if information leaks from a model against guaranteeing that no information leaks. [paper]
- Reading Is Not Leaking: Local, Auditable Measurement and Reduction of Inference Exposure from Public Footprints This work shows how to measure and reduce the exposure of information when using public language models locally. [paper]
- Residual Transferability in Neural Image Watermarking This paper studies how much watermarking can persist in neural images after transfer between different models. [paper]
- E3C: A Tool for Evaluating Communication and Computation Costs in Authentication and Key Exchange Protocol This tool helps evaluate the costs associated with communication and computation in authentication protocols. [paper] [episode]
- TokenScanner: Detecting Backdoors and Discovering Triggers in Text-to-Image LoRAs via Full Vocabulary Scanning This tool scans text-to-image models to find hidden backdoors using a full vocabulary scan. [paper]
- Semantic Matching of Behavioral Primitives for MUD-Based IoT Device Identification This paper matches the behaviors of devices in a MUD to identify them semantically. [paper] [episode]
- Armadillo: Robust Single-Server Secure Aggregation for Federated Learning with Input Validation This paper presents a secure aggregation method for federated learning that includes input validation. [paper] [episode]
- Stateful Agent Backdoors: Constructing Cross-Session Attack Programs This research focuses on creating backdoors in agents that can be activated across multiple sessions. [paper] [episode]
- Federated Sovereign Transport Protocol (FSTP): Verifiable Coordination Without Disclosure This protocol allows for verifiable coordination in federated systems without revealing sensitive information. [paper] [episode]
- Attributing Sensor Deviations to Degradation, Weather, or Attack in Oilfield Digital Twins: A Simulation Study of Probabilistic Attribution and Cost-Based Decisions This study simulates how to attribute sensor changes in digital twins to different causes like weather or attacks. [paper]
- Hardware-Rooted PUF Fingerprinting for Device-Level Traceability in Knowledge Distillation This paper uses hardware features to create fingerprints for tracing devices during knowledge distillation. [paper]
- Compact Shielded CSV: Post-Quantum, Private, Lightweight Client-Side Validation Blockchain This is a lightweight blockchain solution that provides post-quantum privacy and client-side validation. [paper]
- VulContextBench: A Benchmark for Security Context Retrieval in Coding Agents This paper creates a benchmark to test how well coding agents retrieve security context. [paper]
- You Can't Spot a Deepfake?And Neither Can Your Brain Nor Eyes: A Neurophysiological Framework for Deepfake Exploitation of Cognitive Engagement and Implicit Visual Evaluation This research uses neuroscience to study how deepfakes exploit human cognitive engagement. [paper]
- Something to Talk About: Social Media as a Lens on Healthcare Ransomware Events This paper examines social media as a source of information regarding healthcare ransomware incidents. [paper]
- BMA: Backchain Memory Attacks Create Unauthorized Control Paths in LLM Agents This paper shows how backchain memory attacks can create unauthorized control paths in LLM agents. [paper]
- Coordinated Electromagnetic Side-Channel Attacks for Voter--Ballot Linking: A Case Study of the Brazilian E-Polling System This study analyzes coordinated electromagnetic attacks used to link voters to ballots during e-polling. [paper]
- Weird Machine Compositors: Exploiting AI Orchestration at the Expression Layer This paper explores how to exploit AI orchestration at the level where expressions are created. [paper]
- Where the Numbers Come From: Auditing Evaluation in Provenance-Based Intrusion Detection This paper focuses on auditing how numbers are generated in intrusion detection systems based on provenance. [paper]
- On the Usage of Verifiable Credentials in Privacy-Preserving Federated Analytics This paper discusses using verifiable credentials for privacy-preserving federated analytics. [paper]
- Never Emitted: Reporter Attribution in GitHub's Machine-Readable Vulnerability Records This research looks at how to attribute reports to the original reporter in GitHub vulnerability records. [paper]
- Application Agnostic EM Side-Channel Emanations of the FPGA Clock Distribution Network This paper analyzes electromagnetic emanations from FPGA clock distribution networks across different applications. [paper]
- API Secrets Should Never Become Tokens in the LLM's Vocabulary: A Threat Analysis of API Credential Handling in LLM Agent Systems and an Empirical Evaluation of a Vault-Mediated Execution Boundary This paper analyzes the risks of using API secrets as tokens and evaluates vault-mediated execution boundaries. [paper]
- HESP: Separating What to Probe from When to Stop in Local LLM Alert-Triage Agents This paper proposes a method for local agents to separate what they should probe from when they should stop. [paper]
- Information Blackhole: Exploring Backdoor Mechanism in 3D Point Cloud Reconstruction This research investigates backdoor mechanisms that can cause information blackholes during 3D point cloud reconstruction. [paper]
- COGNIT-Guard: Calibrated Standalone Direct-Decision Guardrails with Heterogeneous CPU-NPU Confidence Cascading under Explicit Latency and False-Positive Constraints This paper describes a guardrail system that uses different hardware components to make calibrated decisions under strict constraints. [paper]
- SecProbe: Adaptive Evaluation of Coding Agents on Cybersecurity Vulnerabilities This paper presents an adaptive method for evaluating coding agents against cybersecurity vulnerabilities. [paper]
- Estimation is Not Enough: Carpet-Bombing Detection via Per-Packet Uniformity Testing This research proposes detecting carpet bombing attacks by testing the uniformity of individual packets. [paper]
- Evaluating System One Models for Agent Security Decisions: Reliability, Calibration, and Selective Automation This paper evaluates system one models used for agent security decisions based on reliability and calibration. [paper]
- What Does It Mean to Forget a Person? Individual-Level Unlearning in Vision-Language Models This work explores how to perform individual-level unlearning in vision-language models. [paper]
- The Price of Peeking: Anytime-Valid Leakage Detection on ML-KEM EM Traces This paper discusses detecting leakage from ML-KEM electromagnetic traces using anytime validation. [paper]
- Optimizing and Securing the Modern Watermarking Channel for Images This research focuses on improving and securing the modern watermarking channel for images. [paper]
- Poster: Towards ProofWeave: A Privacy-Minimised, Integrity-Anchored Evidence Plane for Continuous Agentic Assurance This paper proposes a privacy-preserving evidence plane to ensure continuous assurance in agentic systems. [paper]
- Implementing Data Diodes Using Commodity Hardware and Open Source Software This paper shows how to build data diodes using standard hardware and open source software. [paper]
- Continuous Assurance of Agentic Security Auditors for Software Delivery Decision Gates This research discusses providing continuous assurance for agents acting as security auditors at software delivery gates. [paper]
- Sustained Participation as a Security Resource: The Bounded Participation Channel This paper explores the concept of sustained participation as a bounded channel for security resources. [paper]
- SoK: Cryptocurrency Mixing and Anonymity - Architectures, Threat Models, Operational Aspects and Security This paper covers the architectures, threats, and security aspects of cryptocurrency mixing for anonymity. [paper] [episode]
The papers
- Decoding One Safety Trigger Token for Balancing Safety and Usability in Large Language Models — Large Language Models (LLMs) are vulnerable to jailbreak attacks that manipulate them into generating harmful content despite safety alignments, and this work proposes D-STT, an algorithm that identifies and explicitly decodes learned "safety trigger tokens" to activate the model [episode]
- MOSAIC-Bench: Measuring Compositional Vulnerability Induction in Coding Agents — Coding agents often pass per-prompt safety review yet ship exploitable code when their tasks are decomposed into routine engineering tickets. [episode]
- AutoDojo: A Generative Benchmark for Evaluating Prompt Injection Defenses in LLM Agents — Indirect prompt injection (IPI) poses a major security threat to LLM-powered agents, and this paper introduces AutoDojo, an adaptive extension of AgentDojo that optimizes IPI against existing defenses using a cheap, black-box attack. [episode]
- HardSecBench: Benchmarking the Security Awareness of LLMs for Hardware Code Generation — Large language models (LLMs) are increasingly used for hardware and firmware code generation, but existing studies primarily evaluate functional correctness while largely overlooking security. [episode]
- Armadillo: Robust Single-Server Secure Aggregation for Federated Learning with Input Validation — Armadillo presents a secure aggregation system designed for federated learning that achieves disruptive resistance against adversarial clients by integrating input validation techniques. [episode]
- MemPoison: Bypassing Selective Memory Mechanisms to Plant Backdoors in LLM Agents — Large language model (LLM) agents increasingly leverage long-term memory to support persistent and autonomous task execution, but this capability introduces a new attack surface: memory poisoning, where adversaries can inject malicious information to influence future behavior. [episode]
- Federated Sovereign Transport Protocol (FSTP): Verifiable Coordination Without Disclosure — This research introduces FSTP, a synchronization boundary and transport layer for federated networks designed to enforce data confinement structurally, addressing the gap in existing protocols where data confinement relies on operator policy rather than protocol structure. [episode]
- Semantic Matching of Behavioral Primitives for MUD-Based IoT Device Identification — Accurate identification of Internet of Things (IoT) devices is crucial for security and policy enforcement, especially as runtime communication patterns evolve over time. [episode]
- SkillBloat: Token Amplification Attacks via Skill Injection in LLM Coding Agents — Agent skills extend coding agents with task-specific instructions, scripts, and resources, but they also create a trusted instruction channel that can be abused beyond conventional security attacks. [episode]
- Stateful Agent Backdoors: Constructing Cross-Session Attack Programs — Existing backdoor attacks on Large Language Model-based agents remain stateless, executing fixed behaviors confined to a single session; this paper proposes a stateful agent backdoor that extends the attack lifecycle across multiple sessions under permission isolation by maintain [episode]
- SoK: Cryptocurrency Mixing and Anonymity - Architectures, Threat Models, Operational Aspects and Security — This survey provides a comprehensive review of mixing proposals and existing implementations, beginning by summarizing a set of review criteria for mixing services, focusing on control structures, obfuscation primitives, and robustness. [episode]
- SilentWood: Efficient Private Inference Over Gradient-Boosting Decision Forests — Gradient boosting decision forests offer higher accuracy and lower training times than decision trees for large datasets, but naively extending private inference protocols to these forests leads to impractical running times. [episode]
- A traffic analysis attack against Introduction Protocol and Onion Services — Tor onion services rely on long-lived introduction circuits to support anonymous rendezvous between clients and services, and although Tor incorporates defenses against traffic analysis, "the introduction protocol retains deterministic routing structure that can be exploited by a [episode]
- Dithered Gaussian Mechanism for Randomness-Efficient Differential Privacy — We present a novel alternative to previous discrete noise mechanisms, which protects against floating-point vulnerabilities without requiring separate privacy accounting, and directly inherits the privacy guarantees of the Gaussian mechanism via postprocessing. [episode]
- Gravity Falls: A Comparative Analysis of Domain-Generation Algorithm (DGA) Detection Methods for Mobile Device Spearphishing — Mobile devices are frequently targeted by eCrime threat actors using SMS spearphishing links that employ Domain Generation Algorithms (DGA) to rotate hostile infrastructure, but research has largely overlooked how well detectors generalize to smishing-driven domain tactics outsid [episode]
- E3C: A Tool for Evaluating Communication and Computation Costs in Authentication and Key Exchange Protocol — Calculating computational and communication costs for authentication and key exchange protocols is crucial for designing lightweight protocols suitable for resource-constrained environments like IoT systems, where minimizing computation pressure is essential. [episode]
- alpha-Wasserstein Mechanism for R' e nyi Pufferfish Privacy — This paper introduces an α-Wasserstein mechanism for achieving (α, ϵ)-Rényi Pufferfish Privacy using Laplace and Gaussian noise, demonstrating that this framework provides exact privacy guarantees without requiring additional relaxations. [episode]
- Every Bit, Everywhere, All at Once: A Binomial Multibit LLM Watermark — Multibit watermarking has emerged as a leading approach for detecting AI-generated content, and this work introduces a fundamentally new approach to encoding complex payloads by directly encoding every bit at every token position using binomial encoding. [episode]
- SilentCall: Hidden Tool-Call Backdoors in Open-Weight Agents, and How to Catch Them —
- Checking Leakage Witnesses versus Certifying Bounded Non-Leakage —
- BMA: Backchain Memory Attacks Create Unauthorized Control Paths in LLM Agents —
- DegreeSpar: Structured Degree Sparsity for Efficient Secure Transformer Inference —
- Residual Transferability in Neural Image Watermarking —
- SkillDRE: Dual-Stage Red-Team Evolution of Agent Skills via Pre-Execution and Runtime Feedback —
- CyberClear: A Benchmark for LLM Agent Systems on APT Attack Chain Provenance —
- Hearsay: Can an Auditor Trust the Record a Deployed Agent Harness Writes? —
- REFINE: A Resilient Evolution Framework for Intelligent Enterprise Alert Triage in Security Operations Centers —
- Reading Is Not Leaking: Local, Auditable Measurement and Reduction of Inference Exposure from Public Footprints —
- VulContextBench: A Benchmark for Security Context Retrieval in Coding Agents —
- Trust the Brand, Lose Control: How Identity Hijacks LLM Agent Orchestration —
- Ask Without Telling: Local SLMs Consult Cloud LLMs Without Revealing Task Intent —
- Silent Failures in Agentic Security Evaluation: A Validated Harness for Tool-Call Mediation Under Indirect Prompt Injection —
- Learning to Refer: Client-Resolved Generation for Privacy-Aware Language Models —
- You Can't Spot a Deepfake?And Neither Can Your Brain Nor Eyes: A Neurophysiological Framework for Deepfake Exploitation of Cognitive Engagement and Implicit Visual Evaluation —
- AgentTell: Behavioural Side-Channel Leakage in Browser-Use Agents —
- Coordinated Electromagnetic Side-Channel Attacks for Voter--Ballot Linking: A Case Study of the Brazilian E-Polling System —
- On the Usage of Verifiable Credentials in Privacy-Preserving Federated Analytics —
- Never Emitted: Reporter Attribution in GitHub's Machine-Readable Vulnerability Records —
- Application Agnostic EM Side-Channel Emanations of the FPGA Clock Distribution Network —
- Estimation is Not Enough: Carpet-Bombing Detection via Per-Packet Uniformity Testing —
- API Secrets Should Never Become Tokens in the LLM's Vocabulary: A Threat Analysis of API Credential Handling in LLM Agent Systems and an Empirical Evaluation of a Vault-Mediated Execution Boundary —
- Evaluating System One Models for Agent Security Decisions: Reliability, Calibration, and Selective Automation —
- Weird Machine Compositors: Exploiting AI Orchestration at the Expression Layer —
- HESP: Separating What to Probe from When to Stop in Local LLM Alert-Triage Agents —
- What Does It Mean to Forget a Person? Individual-Level Unlearning in Vision-Language Models —
- Where the Numbers Come From: Auditing Evaluation in Provenance-Based Intrusion Detection —
- Information Blackhole: Exploring Backdoor Mechanism in 3D Point Cloud Reconstruction —
- The Price of Peeking: Anytime-Valid Leakage Detection on ML-KEM EM Traces —
- COGNIT-Guard: Calibrated Standalone Direct-Decision Guardrails with Heterogeneous CPU-NPU Confidence Cascading under Explicit Latency and False-Positive Constraints —
- SecProbe: Adaptive Evaluation of Coding Agents on Cybersecurity Vulnerabilities —
- The Cost of Stability: Deanonymizing Onion Services Long-Lived Introduction Circuits —
- Can Prompt Anonymity Protect Your Identity From LLM Providers? —
- Near-Duplicate Families Break Exact-Record Membership Inference —
- A2A-CaseVerify: Merkle-Linked Case-Evidence Verification for Cross-Organization A2A Workflows —
- A2A-ForensicTrace: Offline Verification of Tamper-Evident A2A Runtime Evidence —
- TrackFlood: Relocating Latency Attacks from NMS-Free Detectors to Real-Time Trackers —
- The Privacy Fallacy of Crowdsourced Fine-Tuning: Extracting Proprietary Data via Topic-Based Poisoning —
- GateDrain: Availability Attacks and Admission-Side Defense for Confidence-Gated Edge-Cloud Inference —
- DeMark: A Query-Free Black-Box Attack for Quality-Preserving Audio Watermark Removal —
- PerceptFence: Content-Mediation Architecture and Deterministic Coverage for Screen-Share AI Assistants —
- LoRo-Mark:Provably Lossless and Robust Agent Watermarking —
- Long-Term Operational Planning Using Scenario-Based System Load Forecasting —
- Certified Multi-Source Integrity for Structured Agent Actions —
- RADNPO: Reference-free Adaptive Negative Preference Optimization for LLM Unlearning —
- ReproBench: Benchmarking LLM Agents on Reproducing Vulnerability From Scratch —
- Breaking Windows Malware Detection: A Comprehensive Evaluation of Problem-Space Adversarial Robustness —
- CoDeL: Co-Evolutionary Defense against Indirect Prompt Injection in LLM-based Agents —
- How to Tame a Multi-Headed Hydra? Adaptive Multi-Category Safety Steering for Large Language Models —
- SEAD: A State-Based Perspective on Attack and Defense in Tool-Using Agents —
- AuxMark: Defending Against Unauthorized Agent Distillation via Auxiliary Behavioral Watermarking —
- StallGrid: Measuring Internet-exposed Engagement in Protocol-Native OT Tarpits —
- Optimizing and Securing the Modern Watermarking Channel for Images —
- CoSec: Benchmarking Agent Security in Communities —
- Physics-Attested Federated Learning: Securing Collaborative Anomaly Detection in Critical Water Infrastructure —
- JEV as a Judge for Agent Trace Security: An Empirical Comparison with Generative LLM Judges —
- JevVibe: Efficient Classification-Guided Secure Code Generation —
- GAZEleak: Passcode Inference Against Eye-tracking XR Devices Through External Observation —
- Efficient TCitH-Based Alternatives to SLH-DSA: Cross-Layer ASIC Design of Mirath —
- LENS: The Sum Is Worse Than the Parts for Set-Level Poisoning in Retrieval-Augmented Generation —
- Trajectory-Level Security Debt in LLM Coding Agents —
- OT-PCA: New Key-Recovery Plaintext-Checking Oracle Based Side-Channel Attacks on HQC with Offline Templates —
- Poster: Towards ProofWeave: A Privacy-Minimised, Integrity-Anchored Evidence Plane for Continuous Agentic Assurance —
- Implementing Data Diodes Using Commodity Hardware and Open Source Software —
- Continuous Assurance of Agentic Security Auditors for Software Delivery Decision Gates —
- Sustained Participation as a Security Resource: The Bounded Participation Channel —
- "Nothing to See Here'': Unintended Disclosure through Revision Traces of LLM Deliverables —
- AI-Based Vulnerability Assessment Capability and Cyber Attack Graph Analysis —
- Indistinguishability of Sum of Permutations: A Fourier Analytic Route to Classical and Quantum Security —
- LLM-Assisted Automatic Security Proofs for Cryptographic Protocols: How Far Are We? —
- Beyond Scalar Probes: Exploiting Vector-Valued Outputs in ReLU Networks For Signature Extraction —
- Detecting False Data Injection and Unstable Operation in Smart Grid via System-Aware Graph Boundary Learning —
- INTCC: A Framework for Interactive Confidential Computing —
- The Compiler May Read It, the Agent May Not: Keeping Part of a Research Code Away from a Coding Agent —
- SEABench: Benchmarking Endogenous Misalignment In Self-Evolving Agents —
- Server-Enforced Watermarking in U-Shaped Split Federated Learning —
- Digital Agriculture Sandbox for Collaborative Research —
- The Synthetic Media Shift: Tracking the Rise, Virality, and Detectability of AI-Generated Multimodal Misinformation —
- When RAG Chatbots Expose Their Backend: An Anonymized Case Study of Privacy and Security Risks in Patient-Facing Medical AI —
- Proteus: A Self-Evolving Red Team for Agent Skill Ecosystems —
- Watermarking Should Be Treated as a Monitoring Primitive —
- Who Owns This Agent? Tracing AI Agents Back to Their Owners —
- The End of Trust: How Agentic AI Breaks Security Assumptions —
- Benchmarking and Exploring the Capabilities of LLMs for Attack Investigations —
- Retrieval Observability Bounds on Provenance Detection for Agent Memory Poisoning: Measured Coverage and a Falsified Standalone Detector —
- A Large-Scale Benchmark and Risk Assessment of Traffic Analysis Attacks on Cloud LLM Services —
- TokenScanner: Detecting Backdoors and Discovering Triggers in Text-to-Image LoRAs via Full Vocabulary Scanning —
- Hardware-Rooted PUF Fingerprinting for Device-Level Traceability in Knowledge Distillation —
- Attributing Sensor Deviations to Degradation, Weather, or Attack in Oilfield Digital Twins: A Simulation Study of Probabilistic Attribution and Cost-Based Decisions —
- Evasion Attacks on Cost-Utility-Based Adversarial Training for Online AutoML in IoT Networks —
- Something to Talk About: Social Media as a Lens on Healthcare Ransomware Events —
- Compact Shielded CSV: Post-Quantum, Private, Lightweight Client-Side Validation Blockchain —
Important terms
- Server-enforced watermarking
- Embedding invisible markers directly into model updates during training to allow later verification of content source, treating watermarking as a monitoring primitive.
- SkillDRE
- A systematic test for evolving agent skills using pre-execution and runtime feedback. It reveals hidden vulnerabilities in how an agent acquires new capabilities.
- TokenScanner
- A tool that performs a full vocabulary scan to detect backdoors and triggers within text-to-image models, addressing security concerns in generative AI.
- COGNIT-Guard
- A system using heterogeneous hardware for calibrated guardrails. It handles latency and false positives in real-time autonomous systems, ensuring reliable decision making.