Security papers — 2026-10-08

Today's focus is on how adversarial images can hijack web agents, which shows a new way to compromise automated systems from visual input to actual browser execution. We looked at methods like constitution-guided watermarking and visual memory attacks that persist through the key-value cache. This persistence means an attacker's influence can linger in the agent's short-term memory, making detection harder.

The research also explored constrained action AI remediation for SIEM and XDR systems using a NeMo Guardrails Proxy. This attempts to stop harmful actions by limiting what an agent can do based on its input. This is important because it moves beyond just flagging threats to actively controlling the system's behavior in real-time. Following that, we investigated sensitive topic leakage through LLM routing metadata and mitigation strategies for that risk.

Finally, there was a look at the cost of delay for post-quantum migration, comparing classical and harvest-now decrypt-later risks on a single ordered list. This helps frame the urgency of adopting new cryptographic standards versus waiting for future threats to materialize.

The most critical piece of work today involved investigating how to stop large language models from being tricked into revealing sensitive information through prompt engineering. If we cannot control what these models disclose, their security and trustworthiness are fundamentally compromised.

One line of research focused on the Trojan knowledge problem, which explored bypassing commercial LLM guardrails by weaving harmless prompts and using adaptive tree search to find loopholes in the model's safety mechanisms. This means they were trying to figure out clever ways to get the model to ignore its built-in rules and spit out restricted data.

Another significant effort looked at making sure agents are safe when they are attacked by decomposition attacks, using a new benchmark called DECOMPBENCH to test this resilience. This testing helps determine if an agent can be tricked into revealing hidden vulnerabilities even when it is being broken down into smaller steps.

Then there was the work on Curvature-Guided Module Localization for low-rank detoxification of backdoored large language models. This attempts to find and remove malicious code hidden within the model's parameters by looking at how the model's structure curves. This is a direct attempt to clean up compromised models before they are deployed.

Finally, there was COD-ssi, which deals with enforcing mutual privacy for credential oblivious disclosure in self-sovereign identity systems. This work is important because it tackles the issue of protecting personal credentials when using decentralized identity methods.

The most significant development today involves the work on MARS, which attempts to analyze malware by using rule-based scoring for claims made by large language models. This addresses the growing risk of relying on potentially flawed outputs from AI in security analysis. The authors found that while these models can generate plausible-sounding reports, they often make unsupported findings when reconstructing agent logs.

This is connected to the research on Multi-Aspect Runtime Verification for Simulation-Based V&V of LLM-Enabled Autonomous Agents. This work tries to check these autonomous agents while they are running by looking at multiple aspects simultaneously during simulations. This helps ensure their behavior remains compliant with expected security protocols.

Another piece of work focuses on Adversarial RL for Port-Scan Evasion in edge-deployed intrusion detection systems. This research investigates how attackers can evade these systems by learning the best ways to perform port scans. It tries to make the attacker's features visible so we can defend against them better.

Finally, there is the study on Visible-Spectrum Optical Covert Channels in Commodity Smart Lighting. This explores how hidden communication channels might exist within common lighting technology. This opens up new avenues for covert data transmission that might bypass traditional network monitoring tools.

The most critical piece of work on the day involved exploring how black box adversarial patch attacks can compromise Vision Language Models through ancestor vision language model exploitation. This matters because it shows a direct pathway for injecting malicious visual information into these complex models. Researchers tested their efficacy against various Vision Language Models using methods derived from ancestor VLM exploitation techniques. The findings indicated that the attacks were successful in manipulating the model's understanding of the input data, suggesting a vulnerability in how these models process visual context when subjected to targeted perturbations.

Following this, there was work on why defenses against malicious finetuning erode as training continues. This research investigated how repeated exposure to adversarial fine-tuning degrades the robustness of existing security measures within large language models. It suggests that simply adding defenses is not enough; the continuous training process itself can weaken those safeguards over time.

Another significant area explored was using LLM-guided reinforcement learning to create autonomous cyber defense systems. This involved setting up an agent to learn how to defend against threats by being guided by an LLM. This is a key step toward automated security responses.

The study on trust highlighted the weakest assumptions protocols need when operating in real-world scenarios. It examined the fundamental vulnerabilities inherent in established communication or operational protocols, pointing out where external actors can exploit weak points for malicious gain.

Finally, there was work on formal runtime verification for tool-using LLM agents, specifically comparing AgentDojo and STAC on an offline same-benchmark study. This aimed to formally prove the safety of agents that use tools by checking their execution paths in real time. This contrasts with the hybrid hierarchical runtime verification approach developed for edge-IoT security, which combines MonPoly and RTLola to secure devices at the network level.

The most significant development today concerns TwinGuard-Lite, which introduces a rule-based state admission gateway designed for generative patient digital twins. This is important because it directly addresses the security of creating personalized medical models by controlling what information flows into them. The work involved developing this gateway to manage the inputs for these digital twins.

A related piece of research looked at package hallucination attacks on coding agents, specifically focusing on prompt injection within rule files. Researchers tested how easily malicious instructions could trick automated code generation tools into producing flawed outputs based on their internal rules. This finding suggests a vulnerability in how these agents process structured directives.

Further security work explored Secure-CUA, which aims to control untrusted influence within computer-use agents. This effort builds upon the previous findings by focusing on controlling the agent's behavior when it interacts with external, potentially malicious inputs. The goal here is to establish boundaries for how these agents operate in real-world scenarios.

Then there was work on hierarchical security monitoring for edge-IoT using a formal methods approach. This study established a way to formally verify security properties across different layers of an Internet of Things system. This method provides a rigorous mathematical proof that certain security requirements are met at the hardware and software levels.

Another contribution involved defining purpose-limited secrets, which is about establishing clear boundaries for sensitive information within a system. This work seeks to define precisely what secrets should be accessible to specific parts of an application, aiming to reduce the surface area for potential misuse.

Finally, there is research on betweenCut, which deals with private heavy-node classification using doubly logarithmic error in tree height. This method provides a way to classify nodes in a large network while maintaining privacy guarantees and achieving a very efficient classification structure.

The most critical development today concerns the reliability of benchmarks used to test Large Language Model vulnerability patching. This matters because it directly impacts how we trust automated security fixes for these powerful systems. We looked at CredLeakBench, which evaluates credential leakage and recovery within LLM agents; this work suggests that current methods for assessing agent security are insufficient when dealing with sensitive information handling.

Another key piece of research focused on SLDR, which proposes a defense against malicious fine-tuning through selective layers recovery and dynamic routing. This technique is significant because it offers a way to actively defend models from adversarial fine-tuning attacks. CredLeakBench tells us what the current weaknesses are in agent credential management.

We also saw work on SwarmReconGuard, which employs black-box detection of distributed collective reconnaissance by benign-looking agent populations. This is important for understanding how coordinated malicious activity can spread across decentralized systems. It builds upon the foundational security questions raised by CredLeakBench regarding agent behavior.

Finally, there was research into ASPIRE, an agentic safety and prompt injection red-teaming engine designed to stress test these agents. This directly relates to the deployment-aware feasibility framework for machine learning-based IoT intrusion detection across edge, fog, and cloud architectures. Understanding how agents are exploited via prompt injection is crucial context when designing detection systems that operate across those different architectural layers.

The development of CYBERFORT shows how to build a compliance chain platform that operationalizes the Cyber Resilience Act for small and medium enterprises. This matters because it provides a concrete framework for SMEs to meet new cybersecurity requirements, moving beyond abstract legislation. The team focused on designing the core architecture of CYBERFORT, which involves creating a platform capable of tracking and managing the entire lifecycle of cyber resilience documentation. They tested several data models to ensure they could handle the varied technical specifications required by different types of products.

One key finding was that a modular approach to compliance checking significantly reduced implementation complexity for smaller firms. This means that instead of needing one massive system, SMEs can adopt pieces as their needs evolve, which is a practical takeaway from the initial design phase. Furthermore, the platform successfully integrated automated reporting features based on predefined regulatory checkpoints. This automation streamlines the tedious process of manual compliance checks for users who are not cybersecurity experts.

The research also explored user experience during the data input and verification stages of the compliance chain. Feedback indicated that intuitive interfaces were crucial for adoption, suggesting that usability is just as important as technical accuracy in this kind of system. Finally, while the platform demonstrates strong foundational capabilities, open questions remain regarding its scalability across vastly different industry verticals and how it will interface with existing legacy enterprise systems.

Today's papers

The papers

Important terms

Adversarial Images
These are manipulated visual inputs designed to hijack web agents by tricking them into executing harmful actions, showing a new way to compromise automated systems from visual data.
Constrained Action AI Remediation
This involves using tools like NeMo Guardrails Proxy to limit what an agent can do based on its input, actively controlling system behavior in real-time instead of just flagging threats.
Trojan Knowledge Problem
This research focuses on bypassing LLM guardrails by weaving harmless prompts and using adaptive search to find loopholes, allowing models to reveal restricted data they shouldn't.
Curvature-Guided Module Localization
This technique attempts to clean up backdoored large language models by finding and removing malicious code hidden within the model's structure based on how its parameters curve.
MARS Analysis
MARS uses rule-based scoring to analyze malware claims made by LLMs, addressing the risk that AI might generate plausible but unsupported findings when reconstructing agent logs.