Security papers — 2026-09-16
The work focusing on evaluating adversarial attacks against large language models centers on MarkSec, a framework designed to unify the analysis of stealing, scrubbing, and spoofing attacks against LLM watermarks. This is important because previous studies often looked at these attack types in isolation without shared calibration or metrics.
MarkSec introduces a common reporting protocol and a quality-constrained attack success metric to assess both effectiveness and text quality simultaneously. Experiments across various watermarks and attacks revealed that apparent winners depend heavily on text-quality constraints, the generality of the attack, and assumptions about model capability.
This connects to work on LLM agents where methods like Attacker Tool Filtering and Normal Tool Recalling were introduced to develop universal defenses against adversarial attacks targeting tool integration. These methods significantly reduced attack success rates across several models while maintaining task success, showing that simple, modular defenses can be very effective when layered correctly.
Another area of concern is the security of autonomous AI-penetration testing agents, where researchers are characterizing their trust boundaries to propose context-aware guardrails. This research highlights the need for specialized defenses against agent architecture attacks beyond standard conversational AI safeguards.
The work concerning the practical realizability of white-box backdoor constructions in machine learning models matters because it tests whether theoretical security guarantees hold up when implemented using standard computing tools. The implementation effort tried to realize a white-box attack against models trained with Random Fourier Features using only numpy and scipy, aiming to see if this threat requires specialized cryptographic infrastructure or can be achieved with commodity scientific-computing tools.
The team found no evidence of detectable difference between the backdoored and clean models across a range of sparsity ratios rho equal to d sparse over D. This means that despite implementing the construction end to end, they could not find any measurable distinction between the two types of models when tested using weight-space and functional black-box comparisons. They also reported which parts of the construction were relatively straightforward to realize, while noting that other components, such as the underlying lattice hardness reduction, required derivation not fully detailed in the original paper.
This practical testing effort contributes to understanding how feasible it is to execute the core concepts of white-box CLWE within real-world computational environments. This finding connects with research on homomorphic inference feasibility, which also deals with executing complex models under constraints, though that work focused on genomic foundation models and this implementation focused on feature extraction backdoors.
The most pressing work concerns understanding how to build robust security around the infrastructure that powers our daily digital lives, specifically in power grids where the separation between information technology and operational technology is so critical. This matters because a failure in one area can cascade into physical disruption, which is why analyzing standards like IEC 62351 and IEC 62443 alongside emerging AI-driven threat detection methods is so important for maintaining system integrity.
We also have foundational work addressing the inherent difficulty of securing complex hardware platforms, where the semantics of individual components are often poorly described. The Sockey project tackles this by creating a domain-specific language to formally describe hardware behavior from reference manuals, which allows them to prove security properties like memory confidentiality and integrity for eight different platforms. This formal description is crucial because it moves beyond guesswork when dealing with closed-source hardware and helps system integrators find counterexamples or confirm correct designs.
On the protocol side, the research into bidirectional fully encrypted protocols is significant because previous unidirectional attempts failed to capture the complexity of two-way communication, often leading to detectable issues like traffic imbalance. The team introduced new formal security definitions for bidirectional FEPs, resulting in provably secure BiFEPs for both datastream and datagram settings, which shows that existing deployed protocols do not meet the full set of required security properties.
Furthermore, there is important work on agentic detection frameworks designed to find hidden log file exposures within third-party software plugins for popular content management systems. This approach uses an LLM agent to analyze thousands of plugins, and they found that while some protective measures exist, multi-layered protection is often missing, leading to new best practices for developers.
Finally, the work on GPU privilege escalation using Rowhammer demonstrates that attacks previously thought limited to degrading model accuracy can actually lead to gaining root shell and system-wide control by exploiting page table management. This finding shows that hardware vulnerabilities are not just about subtle performance degradation but can lead to severe security compromises even in non-multi-tenant settings.
The most significant piece of work here is GPUHammer because it directly addresses a major gap in security research concerning machine learning hardware, specifically by demonstrating a practical Rowhammer attack on NVIDIA GPUs using GDDR6 memory. This matters because these attacks could allow attackers to tamper with trained machine learning models, causing substantial accuracy drops up to eighty percent.
The core of this attack involves novel techniques to reverse-engineer the physical memory row mappings within the GDDR DRAM, which is difficult due to proprietary hardware and high latency challenges. This mapping discovery work is foundational because it unlocks the ability to target specific memory locations for bit-flips. Following that, GPUHammer employs GPU-specific memory access optimizations designed to amplify the hammering intensity while simultaneously bypassing existing security mitigations within the GDDR chips.
The demonstration showed a successful attack injecting up to eight bit-flips across four DRAM banks on an NVIDIA A6000 card with GDDR6 memory. This result is critical because it proves that these vulnerabilities are not purely theoretical but can be exploited in real-world discrete GPU setups. Furthermore, the authors showed how this capability allows an attacker to directly tamper with ML models, leading to those significant accuracy reductions mentioned earlier.
This success builds upon the initial mapping discovery, which required reverse-engineering proprietary memory layouts using FPGA-based test platforms. This process is necessary because understanding where the rows are located is the prerequisite for effective hammering. The overall impact suggests that current security assumptions about GPU memory isolation are insufficient against this type of physical fault injection.
The most pressing concern right now is how easily attackers can manipulate the perceived distance of objects in autonomous systems, which directly impacts safety. This vulnerability arises from the way stereo cameras sample pixels and calibrate them, allowing simple repeating patterns to control estimated depths without needing complex machine learning tricks. This manipulation affects both traditional stereo matching algorithms like BM and SGBM, as well as deep learning models such as PSMNet, MoCha-Stereo, and UniMatch.
The impact is significant because in a real driving scenario evaluated in CARLA at speeds up to forty kilometers per hour, a brief half-second attack can cause an autonomous vehicle to initiate emergency braking. This shows that the issue isn't just theoretical; it affects deployed systems like the ZED2 and Intel RealSense D435 cameras, where obstacles can be shifted up to twenty meters farther away or twelve meters closer.
While this attack is feasible, our work also confirms that current state-of-the-art defenses are ineffective against it. We propose a new strategy that uses similarity scores to dynamically spot and suppress these depth errors. This approach builds upon the understanding of how these vulnerabilities exist in both stereo matching and deep learning depth estimation models.
The most crucial finding relates to how mobile agents can be tricked into performing malicious actions through UI desynchronization threats, which matters because it shows a fundamental breakdown in the assumption that users and agents see the same thing. This mismatch occurs because humans perceive interfaces visually, subject to limitations like occlusion and contrast, while agents consume digital screenshots that retain metadata inaccessible to humans.
The research demonstrated that a repackaged application clone could exploit this desynchronization to steer an agent toward attacker-designated actions while still appearing normal to the human user. This threat is feasible because perturbations embedded before deployment can cause these deviations without needing runtime user instructions or agent detection.
This feasibility is supported by automated framework development that constructs instruction-agnostic UI desynchronization attacks and realizes them in deployable APKs. These attacks were tested across five mobile-agent frameworks and three backbone models on five hundred forty-six tasks, yielding average misleading rates of seventy seven point nine percent and sixty six point nine percent, respectively.
This capability is further contextualized by a questionnaire study involving one hundred eighty-six participants, which indicated that the visual perturbations used in the attacks were difficult for human users to notice. This suggests that even when agents are compromised via this visual discrepancy, the human element remains relatively secure against detection.
Today's papers
- MarkSec: Capability-Aware Evaluation of Adversarial Attacks Against LLM Watermarks LLM watermarking helps trace text origin but faces stealing, scrubbing, and spoofing attacks that MarkSec unifies and evaluates. [paper]
- Can We Stop The Ads? Taxonomy and Characterization of Smartphone Splash Ads and Existing Countermeasures This paper analyzes defenses against full-screen ads on smartphones, finding that many require difficult modifications like rooting or jailbreaking to be effective. [paper]
- Toward Secure AI-Powered Penetration Testing Agents: Security Threats, Guardrails, and Architectural Perspectives This work proposes a threat taxonomy for autonomous AI penetration testing agents to better secure them against various lifecycle attacks. [paper]
- gr-PHYSEC: Real-time Channel-based Key Generation for Physical Layer Secure Wireless Communications This paper presents a new method using neural networks in GNU Radio to generate physical layer keys directly from wireless channel randomness in real time. [paper]
- Permutation-Based Stegomalware in Large Language Models: Threats and Countermeasures This paper explores how permutation symmetries can be used both to defend against and attack large language models by embedding malicious payloads into their weights. [paper]
- Universal Defenses for Tool-Integrated LLM Agents Against Adversarial Attacks This work introduces universal tool-based and prompt-based defenses to significantly reduce the attack success rates of LLM agents integrated with external tools. [paper]
- Not All Relations Are Equal: Relation-Balanced and Calibrated Graph Learning for Provenance-Based Intrusion Detection This paper proposes a method to detect intrusions by better capturing rare interaction patterns in system graphs while calibrating reconstruction errors against each relation's benign distribution. [paper]
- GPUThor: Amplifying Rowhammer Attacks via Non-Uniform Patterns to Exploit ECC-Protected GPUs This paper introduces GPUThor, an attack that uses non-uniform hammering patterns to achieve significantly higher bit flip rates and exploit privilege escalation on NVIDIA GPUs. [paper]
- Implementing a White-Box Undetectable Backdoor for Random Fourier Features This paper implements a white-box backdoor construction in models trained with Random Fourier Features to test its practical realizability using commodity scientific computing tools. [paper]
- InceptionRAG: Stealthy Poisoning Attack Against Retrieval-Augmented Generation This paper introduces InceptionRAG, a stealthy attack that fragments malicious data into harmless passages to bypass existing mitigation mechanisms in retrieval-augmented generation systems. [paper]
- Feasibility of Homomorphic Inference for a Genomic Foundation Model This paper assesses the feasibility of running genomic foundation models using client-assisted approximate homomorphic encryption to process sensitive data without exposing plaintext. [paper]
- CBW: Towards Dataset Ownership Verification for Speaker Verification via Clustering-based Backdoor Watermarking This paper proposes a clustering-based backdoor watermark to verify dataset ownership in speaker verification models by partitioning speakers into clusters with distinct triggers. [paper]
- Risk-Calibrated Bayesian Streaming Intrusion Detection with SRE-Aligned Decisions This paper presents a streaming intrusion detection approach that aligns alert decisions with Site Reliability Engineering error budgets using Bayesian Online Changepoint Detection. [paper]
- RuleAutoPilot: Synthesizing Deployable Suricata Rules from Network Traffic This paper introduces RuleAutoPilot, an agentic framework that automatically generates deployable Suricata rules directly from network traffic without needing prior threat intelligence. [paper]
- Observational Indistinguishability and Integrity Blind Regions in Hybrid Quantum-Classical Workflows This paper presents a claim-relative evidence/reference framework to analyze the integrity of hybrid quantum-classical workflows by identifying structural blind regions. [paper]
- Evaluating the NIST Bugs Framework Against CWE as a Successor for Automated Vulnerability Classification This paper evaluates the NIST Bugs Framework against CWE, suggesting it is a more structured and automation-friendly framework for classifying software vulnerabilities. [paper]
- Cybersecurity in Power Grids: Standards and Research Challenges This paper reviews the critical distinctions between IT and OT environments in smart grid cybersecurity, analyzing standards like IEC 62351. [paper]
- Sockeye: Bug-finding and proofs for platform configurations and hardware based on reference manuals This paper introduces a domain-specific language to formally describe hardware semantics from reference manuals, enabling security proofs for complex platforms. [paper]
- Closing the Loop: Bidirectional Fully Encrypted Protocols This paper introduces formal security definitions and provably secure bidirectional fully encrypted protocols designed to prevent detection attacks in two-way communication. [paper]
- Plug 'n' Pray: Agentic LLM-based Detection of Potential Log File Exposures in Third-Party Content Management System Plugins This paper presents an agentic framework that uses LLMs to automatically detect potential log file exposures within third-party plugins of content management systems. [paper]
- ROSETTA: Efficient and Accurate Privacy-Preserving LLM Decoding via Hybrid CKKS/TFHE Evaluation This paper proposes ROSETTA, a hybrid CKKS/TFHE framework that improves the speed of decoding in large language models for private inference by handling nonlinear operations efficiently. [paper]
- GPUBreach: Privilege Escalation Attacks on GPUs using Rowhammer This paper demonstrates that GPU Rowhammer attacks can be used to achieve privilege escalation and tamper with model code on NVIDIA GPUs. [paper]
- Understanding the Usability of Cryptographic Verification Tools This paper conducts a human-centered study revealing usability barriers in cryptographic protocol verification tools, suggesting a need for better diagnostics and visualization. [paper]
- SCHERI: Provably Secure Speculation Under the Constant-Time Policy for CHERI (Extended Version) This paper introduces SCHERI, a processor design that provides end-to-end secure speculation guarantees while maintaining constant-time policies. [paper]
- GPUHammer: Rowhammer Attacks on GPU Memories are Practical This paper introduces GPUHammer, the first attack targeting GDDR6 memory on NVIDIA GPUs to cause significant accuracy drops in machine learning models. [paper]
- You Shall Not Pass into Ring-0! A User Privacy-Friendly Anti-Cheat Architecture for Personal Computers This paper proposes Tirith, an anti-cheat architecture that uses protected virtual machines to provide kernel-level protection without compromising user privacy. [paper]
- SEMA-GUARD: Semantic and Graph-Based Vulnerability Detection in Assembly Code This paper presents SEMA-GUARD, a framework using semantic analysis and graph neural networks to identify vulnerabilities in compiled assembly code. [paper]
- From Hypervisor to Container: Cloud Security Vulnerabilities, Defense Mechanisms, and Open Challenges This paper reviews cloud security vulnerabilities like VM escape and container breakouts while introducing a quantitative scoring framework for defense systems. [paper]
- A Cyber Range Evaluation of Autonomous Network Incident Response Agents This paper tests the performance of reinforcement learning agents in a cyber range designed to evaluate autonomous network intrusion response policies. [paper]
- Human Factors in Cybersecurity in Icelandic Small and Medium-sized Enterprises This study surveys human factors affecting cybersecurity in Icelandic SMEs, recommending targeted training and cultural changes to mitigate threats. [paper]
- The MAL Simulator: Cyber Operations Simulation based on Attack & Defense Graphs This paper develops the MAL Simulator, a cyber operation simulator that uses an attack modeling language to train and evaluate offensive and defensive agents. [paper]
- Cross-Domain Inference for Human Localization: Applying Wi-Fi RSSI Data to CSI-Trained Models This paper shows that existing channel state information models can be adapted to predict human locations using less privileged received signal strength indicator data. [paper]
- Illusion of Depth: Revealing Hidden Stereo Vision Vulnerabilities in Depth Estimation This paper reveals an intrinsic vulnerability in stereo cameras and deep learning depth estimation models that allows attackers to manipulate perceived obstacle depth. [paper]
- MOZAIK: A Privacy-Preserving Analytics Platform for IoT Data Using MPC and FHE This paper presents MOZAIK, a privacy-preserving architecture for IoT data analytics using secure multi-party computation and fully homomorphic encryption. [paper]
- Do LLMs Make Neural Distinguishers Wise? This paper investigates whether large language models can improve the performance of neural distinguishers used in symmetric-key cryptography. [paper]
- Stream Assembly Is an Uncontrolled Treatment in Streaming Intrusion-Detection Benchmarks This paper demonstrates that reordering evaluation streams in streaming intrusion detection benchmarks changes the measured results, showing they are not comparable across studies. [paper]
- Analyzing Multi-Factor Authentication Through Cryptographic Security Properties This paper examines how modern multi-factor authentication systems use cryptographic security properties to validate user identity against replay attacks. [paper]
- RobResilience: Implementing and Evaluating a Resilience Framework for Cyber-Physical Embodied Systems This paper presents RobResilience, a formal resilience framework that allows cyber-physical systems to determine if disruptions are tolerable at runtime. [paper]
- No Bit Left Behind: Using Brute-Force Lifting to Achieve Fully Static Binary Recompilation This paper introduces a fully static binary recompiler that lifts arbitrary binaries to LLVM IR without needing runtime translation support on the target machine. [paper]
- GAUGE: A Formal Framework for Measuring Cryptographic Security under Heterogeneous Adversary Cost Models This paper formalizes cryptographic security as a function over adversary cost models, providing an auditable framework for comparing different cryptographic schemes. [paper]
- Exploiting and Securing Docker containers and Kubernetes pods from a MitM attack This systematic review explores techniques for securing containerized systems against Man-in-the-Middle attacks by proposing a zero trust architecture. [paper]
- When Agents See Differently: Exposing UI Desynchronization Threats in Mobile Agents This paper demonstrates that mobile agents can be steered toward attacker actions by exploiting the mismatch between user and agent perceptions of the same mobile interface. [paper]
The papers
- CBW: Towards Dataset Ownership Verification for Speaker Verification via Clustering-based Backdoor Watermarking —
- GPUHammer: Rowhammer Attacks on GPU Memories are Practical —
- Risk-Calibrated Bayesian Streaming Intrusion Detection with SRE-Aligned Decisions —
- Sockeye: Bug-finding and proofs for platform configurations and hardware based on reference manuals —
- MOZAIK: A Privacy-Preserving Analytics Platform for IoT Data Using MPC and FHE —
- GPUBreach: Privilege Escalation Attacks on GPUs using Rowhammer —
- Stream Assembly Is an Uncontrolled Treatment in Streaming Intrusion-Detection Benchmarks —
- Human Factors in Cybersecurity in Icelandic Small and Medium-sized Enterprises —
- Do LLMs Make Neural Distinguishers Wise? —
- Universal Defenses for Tool-Integrated LLM Agents Against Adversarial Attacks —
- Permutation-Based Stegomalware in Large Language Models: Threats and Countermeasures —
- Feasibility of Homomorphic Inference for a Genomic Foundation Model —
- Analyzing Multi-Factor Authentication Through Cryptographic Security Properties —
- RuleAutoPilot: Synthesizing Deployable Suricata Rules from Network Traffic —
- Exploiting and Securing Docker containers and Kubernetes pods from a MitM attack —
- Understanding the Usability of Cryptographic Verification Tools —
- Illusion of Depth: Revealing Hidden Stereo Vision Vulnerabilities in Depth Estimation —
- gr-PHYSEC: Real-time Channel-based Key Generation for Physical Layer Secure Wireless Communications —
- Implementing a White-Box Undetectable Backdoor for Random Fourier Features —
- No Bit Left Behind: Using Brute-Force Lifting to Achieve Fully Static Binary Recompilation —
- Evaluating the NIST Bugs Framework Against CWE as a Successor for Automated Vulnerability Classification —
- Not All Relations Are Equal: Relation-Balanced and Calibrated Graph Learning for Provenance-Based Intrusion Detection —
- A Cyber Range Evaluation of Autonomous Network Incident Response Agents —
- GPUThor: Amplifying Rowhammer Attacks via Non-Uniform Patterns to Exploit ECC-Protected GPUs —
- The MAL Simulator: Cyber Operations Simulation based on Attack & Defense Graphs —
- From Hypervisor to Container: Cloud Security Vulnerabilities, Defense Mechanisms, and Open Challenges —
- MarkSec: Capability-Aware Evaluation of Adversarial Attacks Against LLM Watermarks —
- Toward Secure AI-Powered Penetration Testing Agents: Security Threats, Guardrails, and Architectural Perspectives —
- When Agents See Differently: Exposing UI Desynchronization Threats in Mobile Agents —
- InceptionRAG: Stealthy Poisoning Attack Against Retrieval-Augmented Generation —
- ROSETTA: Efficient and Accurate Privacy-Preserving LLM Decoding via Hybrid CKKS/TFHE Evaluation —
- Cybersecurity in Power Grids: Standards and Research Challenges —
- Observational Indistinguishability and Integrity Blind Regions in Hybrid Quantum-Classical Workflows —
- Plug 'n' Pray: Agentic LLM-based Detection of Potential Log File Exposures in Third-Party Content Management System Plugins —
- Cross-Domain Inference for Human Localization: Applying Wi-Fi RSSI Data to CSI-Trained Models —
- SEMA-GUARD: Semantic and Graph-Based Vulnerability Detection in Assembly Code —
- GAUGE: A Formal Framework for Measuring Cryptographic Security under Heterogeneous Adversary Cost Models —
- Can We Stop The Ads? Taxonomy and Characterization of Smartphone Splash Ads and Existing Countermeasures —
- RobResilience: Implementing and Evaluating a Resilience Framework for Cyber-Physical Embodied Systems —
- Closing the Loop: Bidirectional Fully Encrypted Protocols —
- SCHERI: Provably Secure Speculation Under the Constant-Time Policy for CHERI (Extended Version) —
- You Shall Not Pass into Ring-0! A User Privacy-Friendly Anti-Cheat Architecture for Personal Computers —
Important terms
- MarkSec
- A framework designed to unify the analysis of stealing, scrubbing, and spoofing attacks against LLM watermarks by introducing a common reporting protocol and a quality-constrained success metric.
- Attacker Tool Filtering
- A method used in LLM agent defenses that helps reduce attack success rates by filtering out malicious tools an attacker might try to use.
- GPUHammer
- A practical Rowhammer attack demonstrating that physical memory vulnerabilities in NVIDIA GPUs can be exploited to flip bits and tamper with trained machine learning models.
- UI Desynchronization Threats
- Attacks where a repackaged application clone exploits the difference between how humans and agents perceive a user interface, tricking agents into performing malicious actions.