Security papers — 2026-10-10

Work was done on mAVE, a watermark method for joint audio-visual generation models to track the origin of generated media. This is important because knowing where outputs come from is crucial for safety and trust as these agents become more sophisticated.

Researchers also explored false claims in commercial image generators using a red-teaming benchmark to test how easily deceptive images can be created. This work helps us understand the limits of current generation technology when it comes to creating believable but untrue content.

There is a problem with certifying hidden paths in quantum key distribution networks through scalable topology assurance, which is important for securing future communication infrastructure. This connects to agent work because reliable communication channels are a prerequisite for secure agent operation.

Context-binding gaps in stateful zero-knowledge proximity proofs were also looked into, dealing with how context can leak or be misused in complex cryptographic checks. This is a technical hurdle that needs to be cleared before deploying agents that rely on these proofs for verification.

DCVD was also touched upon, which uses dual-channel cross-modal fusion for joint vulnerability detection and localization. This provides a way to pinpoint exactly where security flaws might exist within the system architecture being secured.

The most pressing issue today revolves around the security of large language models through various attack vectors. Work on Phantom Transfer explored how data poisoning attacks can evade existing data-level defenses, meaning malicious actors can still inject poisoned information that survives initial filtering.

A related concern is how these models are being attacked at the serving level. One study focused on rethinking latency denial-of-service by targeting the LLM serving framework itself rather than just overloading the model's core processing power. This suggests vulnerabilities might exist in how these massive systems are deployed and managed, not just within the model weights.

Another area of focus is resource hijacking when using LLM agents, which looks at how attackers can gain unauthorized access to system resources through these agents. This connects to a different line of research examining resource hijacking in LLM agents that goes beyond direct access methods.

On a more technical note, there was an attempt to improve tokenization security with OTRO, which introduced Oblivious Tokenization Path with Square-Root ORAM. This work aims to make the process of tokenizing data more secure by obscuring the path taken by the tokens. This is important because it addresses how information is broken down before it even enters the model's processing pipeline.

Finally, there is a piece on evaluating LLMs themselves, specifically designing a multi-perspective report evaluation for security operation centers. This work suggests that we need better ways to assess the security posture of these models through structured reporting mechanisms, which all feeds into ensuring robust and trustworthy AI systems in practice.

The most critical piece of work today involved understanding the inspection execution gap in agent skill scanners, which is vital because if we cannot trust how an agent actually performs a task after it has been scanned, our entire security posture built around these agents is flawed. The PyCache Trap was looked at to see where the scanner fails to match what it intends to inspect.

This failure point connects directly into MRCert, which aims for post-deployment patch robustness certification by using type-specific masking when samples are adversarially patched, ensuring that a patch holds up under attack. Furthermore, SoK was explored to create a taxonomy and design guidance for failure modes in common criteria product evaluations, providing the framework needed to identify these kinds of gaps systematically.

DITTO proposes a context-aware pickle-based pre-trained model scanner specifically for effective security audits, offering a different approach to pre-deployment checking. This contrasts with the work on when AI finds hidden messages and reports them, which examines the reporting mechanisms of models that uncover latent data.

Work was also done on aligning safety across recurrent depths in looped language models and BRANCH, which deals with bypassing multi-scanner AI guardrails using a different type of scanner altogether.

The most significant development today involves using diffusion models to guide adaptive purification in audio deepfake detection. This matters because it promises a more robust way to filter out synthetic speech by iteratively refining the signal based on learned noise characteristics. Researchers explored how these models can adjust purification steps dynamically, aiming for higher accuracy than static methods.

This work builds upon prior efforts where researchers investigated power side-channel membership inference attacks against embedded machine learning, which showed that attackers could infer membership in a model based on power consumption patterns. A related piece of research looked at CPU-Auth, which is a device fingerprinting technique using DVFS side-channels to authenticate devices, suggesting that hardware characteristics can be exploited for verification.

Another area of focus was understanding how flaws cascade within JavaScript engines, specifically looking at vulnerabilities and exploitation chains that arise from these engine weaknesses. This contrasts with work on speedbumps, which examined rejection attacks on speculative decoding mechanisms in large language models, highlighting another avenue where model inference security is being tested.

There was an empirical study examining the hint weight of ML-DSA signatures across three different FIPS 204 parameter sets, which suggests that the effectiveness of these digital signature schemes is key-dependent. This connects to NOMOS, which compiles written policies into statically verified tool-call gates for LLM agents, showing how policy enforcement can be made more reliable.

The most critical piece of work today involved GROB, which proposes a multi-agent architecture designed to investigate public traces of candidate agentic activity. This matters because it offers a systematic way to look into what agents are actually doing in public data streams.

This approach builds upon the foundational concepts explored in other areas, such as the survey of security research for operating systems, which provides necessary context for understanding system vulnerabilities. Furthermore, the work on MARC introduces multi-bit watermarking specifically targeting autoregressive audio generation to defend against codec attacks, showing how specific cryptographic techniques are being applied to protect data integrity.

The investigation into on-chain archaeology of Bitcoin oracles is also significant because it seeks evidence of actual use under limited observability, which speaks directly to the reliability of decentralized systems. This effort connects with the zero-knowledge signature framework for post-quantum message authentication in automated driving, as both deal with verifying information securely in complex environments.

Understanding where tokens go within LLM agents is important for reducing costs during vulnerability discovery efforts. This contrasts with LTBD, which focuses on learnable trust-boundary delimiters to defend against prompt injection attacks when these agents are being deployed.

The most pressing work this morning centers on Host Attack Graph for Botnet Propagation because understanding how these malicious networks spread is crucial for developing effective countermeasures against large-scale cyber threats. Researchers explored a framework that models the relationships between compromised hosts to map out propagation paths, which helps in identifying key nodes where intervention can stop the infection.

A separate line of inquiry looked at Anytime-valid detection of LLM weight exfiltration because protecting the intellectual property embedded in large language models is a major concern. They proposed a method for detecting when sensitive model weights are being stolen, which means we have a way to catch data theft as it happens rather than after the fact.

SemField introduces a simple semantic watermark designed to be linear and continuous while remaining robust against tampering. This technique essentially embeds an invisible signature into data so that its integrity can be checked later, linking it conceptually to how we might track the flow of information across different systems.

Work was also seen on Secure Aggregate Encryption with Identity-Based Authentication for Multi-Vendor FPGA Cloud Deployment, which addresses the security challenges of deploying hardware across different providers. This protocol offers a provably secure way to handle encryption when dealing with many different vendors in a cloud environment.

Moving toward practical network defense, there is Moving Target Defense in SDN-enabled EV Charging Network research, which focuses on making the network harder for attackers to target by constantly changing its configuration. This means the charging infrastructure becomes less predictable for hackers trying to cause disruption.

Finally, HPQ-AKE presents a provably secure sign-less hybrid authenticated key exchange protocol suitable for bandwidth-constrained IoT and edge networks. This protocol is important because it allows low-power devices to establish secure communication without needing heavy cryptographic signatures, which is vital for massive deployments.

The most significant piece of work today involved exploring how to protect CPU artificial intelligence on edge trusted execution environments by leveraging WebAssembly. This matters because it offers a pathway to secure computation outside traditional hardware boundaries. A preliminary study looked at LLM distillation inference, which essentially means taking a large language model and shrinking it down while still keeping its core abilities intact. This work suggests that distillation can be done in a way that maintains certain properties of the original model, though the specifics of how this manifests are still being mapped out.

Another important thread concerns characterizing statistical separability in TP-CRIV for probabilistic AI models, which is crucial because it helps us understand if different AI models can be distinguished based on their underlying statistical patterns. This research attempts to quantify this separability, providing a mathematical framework for assessing model differences. This connects to the work on BRACE, which uses differential privacy for dense associative memory with LSR energy; that latter project aims to build robust memory structures while ensuring privacy through noise injection.

ORCAGen orchestrates context-aware malware deception using RAG-guided generative AI, a system designed to create sophisticated traps for malicious software by using retrieval augmented generation. This deception method relies on generating plausible but ultimately misleading data based on retrieved context. Finally, there is the work on provable subexponential algorithms for NIST third-round lattice families, which deals with the theoretical limits of solving certain mathematical problems efficiently. This theoretical underpinning provides a benchmark against which practical implementations, like those involving WebAssembly security, can be measured.

The work on ProxyEraseAgent is particularly significant because it tackles the practical challenge of removing digital watermarks in real-world environments without alerting the underlying system. This agent was tested by attempting to blind watermark removal using a specific set of adversarial input patterns, which resulted in a successful erasure rate of seventy-two percent across varied datasets. This success builds upon prior work that explored similar obfuscation techniques, such as those detailed in the MORDOR paper, which focused on mitigating overhead from read disturbance preventive operations through elastic refresh scheduling.

The MORDOR approach aimed to reduce computational strain during read disturbance prevention by using a dynamic scheduling method, and it showed promise in reducing operational overhead. Moving down the list of importance, EIFL addressed protecting global model privacy and integrity when dealing with untrusted servers in federated learning settings; this involved developing methods to ensure that local model updates do not leak sensitive information to the central server.

One Node, Two Roles explored simultaneous contests for validation and attention within rollups, which suggests a way to improve the efficiency of validating data structures by assigning dual roles to nodes. This contrasts with ReSI, which focuses on recursive safety improvement toward creating more resistant and resilient artificial intelligence systems through iterative refinement processes. Finally, the lessons drawn from recent security incidents at OpenAI, Anthropic, and Google Agent Security Incidents highlight a necessary shift from reactive containment strategies toward proactive assurance in agent security protocols.

Today's papers

The papers

Important terms

mAVE
A watermark method used for joint audio-visual generation models to track where generated media originates. This is key for safety and trust in sophisticated AI agents.
Phantom Transfer
Research on how data poisoning attacks can bypass existing defenses by injecting poisoned information that survives initial filtering. This tests the limits of current data-level security.
Context-binding gaps
Issues in stateful zero-knowledge proximity proofs where context might leak or be misused during complex cryptographic checks. Clearing this is vital for deploying secure agents.
Host Attack Graph
A framework to model the relationships between compromised hosts in a botnet, helping to map out propagation paths and identify critical nodes for stopping infections.
GROB
A multi-agent architecture designed to investigate public traces of candidate agentic activity. This provides a systematic way to examine what agents are actually doing in public data streams.