Security papers — 2026-09-29

The work that matters most is developing server-enforced watermarking within U-shaped split federated learning setups. This technique embeds invisible markers directly into model updates during training, allowing later verification if AI-generated content came from a specific source.

This concept treats watermarking as a monitoring primitive instead of an afterthought. Researchers are examining how this works with other agentic systems, specifically Proteus, which is designed to be a self-evolving red team for agent skill ecosystems. This helps determine if these agents can bypass security assumptions when operating autonomously.

Another significant piece of research addresses the growing problem of synthetic media misinformation and the difficulty in detecting it as AI-generated multimodal content gains traction. Furthermore, investigators are looking into the privacy risks in patient-facing medical AI systems where Retrieval Augmented Generation or RAG chatbots expose backend vulnerabilities.

Finally, a large-scale benchmark is assessing the security landscape of large cloud language model services by checking how easily traffic analysis attacks can expose sensitive information. This work connects the need for source tracking with real-world risks posed by autonomous agents and data leakage in critical applications.

The most critical development today involves SkillDRE, which systematically tests how agent skills can be evolved through both pre-execution and runtime feedback. This matters because it shows a pathway for adversarial manipulation of an agent's capabilities before and during its actual operation.

SkillDRE explores this by using a dual-stage red-team evolution process to probe skill sets. It specifically examines how an agent's performance changes when it receives feedback both before starting a task and while the task is running, which reveals hidden vulnerabilities in the skill acquisition process.

CyberClear provides a benchmark for assessing LLM agent systems against advanced persistent threat attack chains by focusing on provenance tracking. Understanding where an attack comes from allows defenders to build better defenses against sophisticated threats.

Hearsay investigates the trustworthiness of records generated by deployed agents, asking whether an auditor can rely on what the agent actually writes when it performs a task. This directly impacts how we can verify the integrity of automated decision-making processes.

REFINE introduces a resilient framework for intelligent enterprise alert triage within security operations centers, aiming to improve how security teams handle incoming alerts. This work is significant because it focuses on making the triage process robust against unexpected or malicious inputs.

Trust the Brand, Lose Control looks at how identity hijacks the orchestration of LLM agents. This is a key concern because if an attacker gains control over an agent's identity, they can potentially redirect its intended actions.

Ask Without Telling examines a method where local small language models consult cloud large language models without exposing the actual task intent to the cloud model. This technique offers a way to leverage powerful external intelligence while maintaining some level of operational privacy.

Learning to Refer addresses privacy concerns by employing client-resolved generation for language models, which helps ensure that generated content respects user boundaries. This method is crucial for deploying these models in sensitive environments where data leakage is a major risk.

The most pressing work involves understanding how agents can leak sensitive information through their browser usage, which exposes user behavior in a way that traditional security measures might miss. A study on AgentTell explored this by measuring behavioral side-channel leakage in browser-use agents, showing they can reveal information about the underlying system or user actions. This suggests a new vector for covert data exfiltration and connects to checking leakage witnesses versus certifying bounded non-leakage.

Another significant piece of research addresses the reliability of detection methods when an agent's memory is being maliciously poisoned, specifically looking at retrieval observability bounds on provenance detection. Researchers measured how well these detectors cover different poisoning scenarios and found that a standalone detector often fails to provide accurate results. This means we need better ways to verify if an agent's memory is trustworthy, which relates to exploring LLMs for attack investigations.

Then there is the work on evasion attacks targeting cost-utility-based adversarial training for online AutoML in IoT networks. This demonstrates how attackers can bypass security measures designed to make these systems robust. This shows that even well-trained models can be tricked into making suboptimal or insecure decisions when facing targeted manipulation, contrasting with DegreeSpar which focuses on structured degree sparsity for efficient secure transformer inference.

The most significant development today involves TokenScanner, which aims to detect backdoors and discover triggers within text-to-image low-rank adaptations by performing a full vocabulary scan. This is important because it addresses growing security concerns surrounding generative AI models where hidden vulnerabilities could be exploited.

This work builds upon the concept of residual transferability in neural image watermarking, which explores how much information from one image can be transferred to another through a watermark. Understanding this transfer is key to measuring inference exposure. Furthermore, E3C presents a tool for evaluating communication and computation costs in authentication and key exchange protocols, offering practical metrics for assessing the efficiency of secure exchanges.

TokenScanner's deep vocabulary scan complements this by looking specifically at the textual prompts used to generate images, tying into how information might leak through different generative pathways. Similarly, Armadillo introduces robust single-server secure aggregation for federated learning. This is a vital step toward building more trustworthy decentralized machine learning systems and relies on input validation to maintain security while allowing model training across distributed data without centralizing sensitive information.

The most critical piece of work today involved developing a simulation study to attribute sensor deviations in oilfield digital twins to potential causes like degradation, weather, or direct attack. This matters because accurately diagnosing the source of a fault is essential for maintaining operational integrity and safety in these complex systems. The researchers explored probabilistic attribution methods within this framework to test how different failure modes influence the likelihood of a specific deviation being caused by wear versus an external event.

Following that, there was work on hardware-rooted physical unclonable functions for device-level traceability in knowledge distillation. This technique uses hardware randomness to create unique fingerprints for devices, which is important for ensuring that distilled models retain verifiable lineage back to their original physical components. This contrasts with the simulation work by focusing on device identity rather than environmental or operational fault diagnosis.

Another area of focus was a compact shielded CSV, a lightweight client-side validation blockchain designed to be post-quantum secure and private. This addresses security concerns by providing a decentralized ledger for validating data locally, which is significant given the increasing threat landscape. This contrasts with the traceability work by focusing on secure data handling rather than model provenance.

The research also touched upon VulContextBench, which serves as a benchmark for retrieving security context in coding agents. This tool helps evaluate how well these agents can understand and utilize necessary security information when performing tasks, linking conceptually to the simulation study's need to correctly interpret system states.

Finally, there was a neurophysiological framework examining how deepfakes exploit cognitive engagement and implicit visual evaluation. This work is significant because it moves beyond technical detection methods to look at the human vulnerability exploited by synthetic media, providing a different kind of context for understanding digital threats compared to infrastructure-focused studies.

The most significant development today involves understanding how AI orchestration at the expression layer can be exploited. Researchers explored weird machine compositors, which are systems that combine different computational elements to create novel outputs; they found ways these compositors can be manipulated to produce unintended results. This manipulation is important because it opens avenues for subtle control over complex AI behaviors.

A related piece of work looked at API secrets and how they interact with large language models, analyzing the threat of API credentials becoming part of the LLM's vocabulary. They empirically evaluated a vault-mediated execution boundary to see if this handling could be mitigated, suggesting a path toward safer agent systems. This contrasts with other security concerns, such as those in provenance-based intrusion detection where auditing evaluation is key to identifying intrusions based on data lineage.

Further down the line, there was work on verifiable credentials used for privacy-preserving federated analytics. This means developing methods to analyze shared data without exposing individual user information, building upon the need for secure handling discussed earlier when considering API secrets and LLM interaction. Separately, research into application agnostic side-channel emanations from FPGA clock distribution networks examined how hardware itself can leak information during computation.

Finally, there is a study on HESP, which separates what an alert triage agent should probe from when it should stop probing within a local LLM environment. This addresses the practical deployment of these AI systems by providing guardrails for their operation.

The most significant piece of work today involves COGNIT-Guard because it tackles the critical need for reliable decision making in autonomous systems by implementing calibrated standalone guardrails. This system uses a heterogeneous CPU and NPU confidence cascading mechanism to handle explicit latency and false-positive constraints, which is vital when these systems are operating in real-time environments.

SecProbe addresses agent security decisions by adaptively evaluating coding agents against known cybersecurity vulnerabilities. This work moves beyond simple checks by assessing how well these agents perform when encountering actual exploits. Following that, the research on Carpet-Bombing detection focuses on using per-packet uniformity testing to detect this type of attack. This is important because it offers a way to catch malicious flooding before it overwhelms the system.

Another area of focus is evaluating System One models for agent security decisions by examining their reliability and calibration when making selective automation choices. This helps us understand how trustworthy these models are when they are tasked with making high-stakes automated judgments, contrasting with the work on individual-level unlearning in vision-language models, which deals with forgetting specific personal data from these large models.

The paper on anytime-valid leakage detection on ML-KEM EM traces is also relevant because it provides a method for detecting subtle information leakage during cryptographic operations. This is closely related to optimizing and securing the modern watermarking channel for images, as both explore ways to ensure integrity or detect unauthorized access within data transmission pathways. Finally, the poster on ProofWeave aims for a privacy-minimised evidence plane anchored by continuous agentic assurance, suggesting a future direction for providing verifiable guarantees in complex agentic workflows.

The most pressing development is the traffic analysis attack against Introduction Protocol and Onion Services, which demonstrates a vulnerability where network traffic patterns reveal sensitive information about the underlying services. This finding is significant because it directly impacts the security of decentralized communication methods by showing how metadata can be exploited.

This work builds upon earlier concerns regarding agentic security auditing, specifically examining how to maintain continuous assurance for those auditors at software delivery decision gates. The research suggests that sustained participation in these channels is a key mechanism for ensuring robust security oversight within communities.

Another critical area involves the implementation of data diodes using commodity hardware and open source software. This provides a physical enforcement layer for data flow control, which is important because it offers a tangible way to isolate systems and prevent unauthorized outbound communication.

Furthermore, the study on SoK cryptocurrency mixing and anonymity details the architectures, threat models, and operational aspects related to cryptocurrency mixing services. This contributes to understanding the security implications of privacy-enhancing technologies in digital finance.

The research into dithered Gaussian mechanisms for randomness-efficient differential privacy addresses how to introduce controlled noise into systems while maintaining a level of privacy. This is relevant because it offers a practical way to balance utility and anonymity in data processing pipelines.

Finally, physics-attested federated learning focuses on securing collaborative anomaly detection within critical water infrastructure by leveraging physical laws for verification. This work is important because it moves beyond purely mathematical security assurances to incorporate verifiable physical constraints into machine learning models.

Today's papers

The papers

Important terms

Server-enforced watermarking
Embedding invisible markers directly into model updates during training to allow later verification of content source, treating watermarking as a monitoring primitive.
SkillDRE
A systematic test for evolving agent skills using pre-execution and runtime feedback. It reveals hidden vulnerabilities in how an agent acquires new capabilities.
TokenScanner
A tool that performs a full vocabulary scan to detect backdoors and triggers within text-to-image models, addressing security concerns in generative AI.
COGNIT-Guard
A system using heterogeneous hardware for calibrated guardrails. It handles latency and false positives in real-time autonomous systems, ensuring reliable decision making.