Security papers — 2026-09-16

The work focusing on evaluating adversarial attacks against large language models centers on MarkSec, a framework designed to unify the analysis of stealing, scrubbing, and spoofing attacks against LLM watermarks. This is important because previous studies often looked at these attack types in isolation without shared calibration or metrics.

MarkSec introduces a common reporting protocol and a quality-constrained attack success metric to assess both effectiveness and text quality simultaneously. Experiments across various watermarks and attacks revealed that apparent winners depend heavily on text-quality constraints, the generality of the attack, and assumptions about model capability.

This connects to work on LLM agents where methods like Attacker Tool Filtering and Normal Tool Recalling were introduced to develop universal defenses against adversarial attacks targeting tool integration. These methods significantly reduced attack success rates across several models while maintaining task success, showing that simple, modular defenses can be very effective when layered correctly.

Another area of concern is the security of autonomous AI-penetration testing agents, where researchers are characterizing their trust boundaries to propose context-aware guardrails. This research highlights the need for specialized defenses against agent architecture attacks beyond standard conversational AI safeguards.

The work concerning the practical realizability of white-box backdoor constructions in machine learning models matters because it tests whether theoretical security guarantees hold up when implemented using standard computing tools. The implementation effort tried to realize a white-box attack against models trained with Random Fourier Features using only numpy and scipy, aiming to see if this threat requires specialized cryptographic infrastructure or can be achieved with commodity scientific-computing tools.

The team found no evidence of detectable difference between the backdoored and clean models across a range of sparsity ratios rho equal to d sparse over D. This means that despite implementing the construction end to end, they could not find any measurable distinction between the two types of models when tested using weight-space and functional black-box comparisons. They also reported which parts of the construction were relatively straightforward to realize, while noting that other components, such as the underlying lattice hardness reduction, required derivation not fully detailed in the original paper.

This practical testing effort contributes to understanding how feasible it is to execute the core concepts of white-box CLWE within real-world computational environments. This finding connects with research on homomorphic inference feasibility, which also deals with executing complex models under constraints, though that work focused on genomic foundation models and this implementation focused on feature extraction backdoors.

The most pressing work concerns understanding how to build robust security around the infrastructure that powers our daily digital lives, specifically in power grids where the separation between information technology and operational technology is so critical. This matters because a failure in one area can cascade into physical disruption, which is why analyzing standards like IEC 62351 and IEC 62443 alongside emerging AI-driven threat detection methods is so important for maintaining system integrity.

We also have foundational work addressing the inherent difficulty of securing complex hardware platforms, where the semantics of individual components are often poorly described. The Sockey project tackles this by creating a domain-specific language to formally describe hardware behavior from reference manuals, which allows them to prove security properties like memory confidentiality and integrity for eight different platforms. This formal description is crucial because it moves beyond guesswork when dealing with closed-source hardware and helps system integrators find counterexamples or confirm correct designs.

On the protocol side, the research into bidirectional fully encrypted protocols is significant because previous unidirectional attempts failed to capture the complexity of two-way communication, often leading to detectable issues like traffic imbalance. The team introduced new formal security definitions for bidirectional FEPs, resulting in provably secure BiFEPs for both datastream and datagram settings, which shows that existing deployed protocols do not meet the full set of required security properties.

Furthermore, there is important work on agentic detection frameworks designed to find hidden log file exposures within third-party software plugins for popular content management systems. This approach uses an LLM agent to analyze thousands of plugins, and they found that while some protective measures exist, multi-layered protection is often missing, leading to new best practices for developers.

Finally, the work on GPU privilege escalation using Rowhammer demonstrates that attacks previously thought limited to degrading model accuracy can actually lead to gaining root shell and system-wide control by exploiting page table management. This finding shows that hardware vulnerabilities are not just about subtle performance degradation but can lead to severe security compromises even in non-multi-tenant settings.

The most significant piece of work here is GPUHammer because it directly addresses a major gap in security research concerning machine learning hardware, specifically by demonstrating a practical Rowhammer attack on NVIDIA GPUs using GDDR6 memory. This matters because these attacks could allow attackers to tamper with trained machine learning models, causing substantial accuracy drops up to eighty percent.

The core of this attack involves novel techniques to reverse-engineer the physical memory row mappings within the GDDR DRAM, which is difficult due to proprietary hardware and high latency challenges. This mapping discovery work is foundational because it unlocks the ability to target specific memory locations for bit-flips. Following that, GPUHammer employs GPU-specific memory access optimizations designed to amplify the hammering intensity while simultaneously bypassing existing security mitigations within the GDDR chips.

The demonstration showed a successful attack injecting up to eight bit-flips across four DRAM banks on an NVIDIA A6000 card with GDDR6 memory. This result is critical because it proves that these vulnerabilities are not purely theoretical but can be exploited in real-world discrete GPU setups. Furthermore, the authors showed how this capability allows an attacker to directly tamper with ML models, leading to those significant accuracy reductions mentioned earlier.

This success builds upon the initial mapping discovery, which required reverse-engineering proprietary memory layouts using FPGA-based test platforms. This process is necessary because understanding where the rows are located is the prerequisite for effective hammering. The overall impact suggests that current security assumptions about GPU memory isolation are insufficient against this type of physical fault injection.

The most pressing concern right now is how easily attackers can manipulate the perceived distance of objects in autonomous systems, which directly impacts safety. This vulnerability arises from the way stereo cameras sample pixels and calibrate them, allowing simple repeating patterns to control estimated depths without needing complex machine learning tricks. This manipulation affects both traditional stereo matching algorithms like BM and SGBM, as well as deep learning models such as PSMNet, MoCha-Stereo, and UniMatch.

The impact is significant because in a real driving scenario evaluated in CARLA at speeds up to forty kilometers per hour, a brief half-second attack can cause an autonomous vehicle to initiate emergency braking. This shows that the issue isn't just theoretical; it affects deployed systems like the ZED2 and Intel RealSense D435 cameras, where obstacles can be shifted up to twenty meters farther away or twelve meters closer.

While this attack is feasible, our work also confirms that current state-of-the-art defenses are ineffective against it. We propose a new strategy that uses similarity scores to dynamically spot and suppress these depth errors. This approach builds upon the understanding of how these vulnerabilities exist in both stereo matching and deep learning depth estimation models.

The most crucial finding relates to how mobile agents can be tricked into performing malicious actions through UI desynchronization threats, which matters because it shows a fundamental breakdown in the assumption that users and agents see the same thing. This mismatch occurs because humans perceive interfaces visually, subject to limitations like occlusion and contrast, while agents consume digital screenshots that retain metadata inaccessible to humans.

The research demonstrated that a repackaged application clone could exploit this desynchronization to steer an agent toward attacker-designated actions while still appearing normal to the human user. This threat is feasible because perturbations embedded before deployment can cause these deviations without needing runtime user instructions or agent detection.

This feasibility is supported by automated framework development that constructs instruction-agnostic UI desynchronization attacks and realizes them in deployable APKs. These attacks were tested across five mobile-agent frameworks and three backbone models on five hundred forty-six tasks, yielding average misleading rates of seventy seven point nine percent and sixty six point nine percent, respectively.

This capability is further contextualized by a questionnaire study involving one hundred eighty-six participants, which indicated that the visual perturbations used in the attacks were difficult for human users to notice. This suggests that even when agents are compromised via this visual discrepancy, the human element remains relatively secure against detection.

Today's papers

The papers

Important terms

MarkSec
A framework designed to unify the analysis of stealing, scrubbing, and spoofing attacks against LLM watermarks by introducing a common reporting protocol and a quality-constrained success metric.
Attacker Tool Filtering
A method used in LLM agent defenses that helps reduce attack success rates by filtering out malicious tools an attacker might try to use.
GPUHammer
A practical Rowhammer attack demonstrating that physical memory vulnerabilities in NVIDIA GPUs can be exploited to flip bits and tamper with trained machine learning models.
UI Desynchronization Threats
Attacks where a repackaged application clone exploits the difference between how humans and agents perceive a user interface, tricking agents into performing malicious actions.