Security papers — 2026-10-08
Today's focus is on how adversarial images can hijack web agents, which shows a new way to compromise automated systems from visual input to actual browser execution. We looked at methods like constitution-guided watermarking and visual memory attacks that persist through the key-value cache. This persistence means an attacker's influence can linger in the agent's short-term memory, making detection harder.
The research also explored constrained action AI remediation for SIEM and XDR systems using a NeMo Guardrails Proxy. This attempts to stop harmful actions by limiting what an agent can do based on its input. This is important because it moves beyond just flagging threats to actively controlling the system's behavior in real-time. Following that, we investigated sensitive topic leakage through LLM routing metadata and mitigation strategies for that risk.
Finally, there was a look at the cost of delay for post-quantum migration, comparing classical and harvest-now decrypt-later risks on a single ordered list. This helps frame the urgency of adopting new cryptographic standards versus waiting for future threats to materialize.
The most critical piece of work today involved investigating how to stop large language models from being tricked into revealing sensitive information through prompt engineering. If we cannot control what these models disclose, their security and trustworthiness are fundamentally compromised.
One line of research focused on the Trojan knowledge problem, which explored bypassing commercial LLM guardrails by weaving harmless prompts and using adaptive tree search to find loopholes in the model's safety mechanisms. This means they were trying to figure out clever ways to get the model to ignore its built-in rules and spit out restricted data.
Another significant effort looked at making sure agents are safe when they are attacked by decomposition attacks, using a new benchmark called DECOMPBENCH to test this resilience. This testing helps determine if an agent can be tricked into revealing hidden vulnerabilities even when it is being broken down into smaller steps.
Then there was the work on Curvature-Guided Module Localization for low-rank detoxification of backdoored large language models. This attempts to find and remove malicious code hidden within the model's parameters by looking at how the model's structure curves. This is a direct attempt to clean up compromised models before they are deployed.
Finally, there was COD-ssi, which deals with enforcing mutual privacy for credential oblivious disclosure in self-sovereign identity systems. This work is important because it tackles the issue of protecting personal credentials when using decentralized identity methods.
The most significant development today involves the work on MARS, which attempts to analyze malware by using rule-based scoring for claims made by large language models. This addresses the growing risk of relying on potentially flawed outputs from AI in security analysis. The authors found that while these models can generate plausible-sounding reports, they often make unsupported findings when reconstructing agent logs.
This is connected to the research on Multi-Aspect Runtime Verification for Simulation-Based V&V of LLM-Enabled Autonomous Agents. This work tries to check these autonomous agents while they are running by looking at multiple aspects simultaneously during simulations. This helps ensure their behavior remains compliant with expected security protocols.
Another piece of work focuses on Adversarial RL for Port-Scan Evasion in edge-deployed intrusion detection systems. This research investigates how attackers can evade these systems by learning the best ways to perform port scans. It tries to make the attacker's features visible so we can defend against them better.
Finally, there is the study on Visible-Spectrum Optical Covert Channels in Commodity Smart Lighting. This explores how hidden communication channels might exist within common lighting technology. This opens up new avenues for covert data transmission that might bypass traditional network monitoring tools.
The most critical piece of work on the day involved exploring how black box adversarial patch attacks can compromise Vision Language Models through ancestor vision language model exploitation. This matters because it shows a direct pathway for injecting malicious visual information into these complex models. Researchers tested their efficacy against various Vision Language Models using methods derived from ancestor VLM exploitation techniques. The findings indicated that the attacks were successful in manipulating the model's understanding of the input data, suggesting a vulnerability in how these models process visual context when subjected to targeted perturbations.
Following this, there was work on why defenses against malicious finetuning erode as training continues. This research investigated how repeated exposure to adversarial fine-tuning degrades the robustness of existing security measures within large language models. It suggests that simply adding defenses is not enough; the continuous training process itself can weaken those safeguards over time.
Another significant area explored was using LLM-guided reinforcement learning to create autonomous cyber defense systems. This involved setting up an agent to learn how to defend against threats by being guided by an LLM. This is a key step toward automated security responses.
The study on trust highlighted the weakest assumptions protocols need when operating in real-world scenarios. It examined the fundamental vulnerabilities inherent in established communication or operational protocols, pointing out where external actors can exploit weak points for malicious gain.
Finally, there was work on formal runtime verification for tool-using LLM agents, specifically comparing AgentDojo and STAC on an offline same-benchmark study. This aimed to formally prove the safety of agents that use tools by checking their execution paths in real time. This contrasts with the hybrid hierarchical runtime verification approach developed for edge-IoT security, which combines MonPoly and RTLola to secure devices at the network level.
The most significant development today concerns TwinGuard-Lite, which introduces a rule-based state admission gateway designed for generative patient digital twins. This is important because it directly addresses the security of creating personalized medical models by controlling what information flows into them. The work involved developing this gateway to manage the inputs for these digital twins.
A related piece of research looked at package hallucination attacks on coding agents, specifically focusing on prompt injection within rule files. Researchers tested how easily malicious instructions could trick automated code generation tools into producing flawed outputs based on their internal rules. This finding suggests a vulnerability in how these agents process structured directives.
Further security work explored Secure-CUA, which aims to control untrusted influence within computer-use agents. This effort builds upon the previous findings by focusing on controlling the agent's behavior when it interacts with external, potentially malicious inputs. The goal here is to establish boundaries for how these agents operate in real-world scenarios.
Then there was work on hierarchical security monitoring for edge-IoT using a formal methods approach. This study established a way to formally verify security properties across different layers of an Internet of Things system. This method provides a rigorous mathematical proof that certain security requirements are met at the hardware and software levels.
Another contribution involved defining purpose-limited secrets, which is about establishing clear boundaries for sensitive information within a system. This work seeks to define precisely what secrets should be accessible to specific parts of an application, aiming to reduce the surface area for potential misuse.
Finally, there is research on betweenCut, which deals with private heavy-node classification using doubly logarithmic error in tree height. This method provides a way to classify nodes in a large network while maintaining privacy guarantees and achieving a very efficient classification structure.
The most critical development today concerns the reliability of benchmarks used to test Large Language Model vulnerability patching. This matters because it directly impacts how we trust automated security fixes for these powerful systems. We looked at CredLeakBench, which evaluates credential leakage and recovery within LLM agents; this work suggests that current methods for assessing agent security are insufficient when dealing with sensitive information handling.
Another key piece of research focused on SLDR, which proposes a defense against malicious fine-tuning through selective layers recovery and dynamic routing. This technique is significant because it offers a way to actively defend models from adversarial fine-tuning attacks. CredLeakBench tells us what the current weaknesses are in agent credential management.
We also saw work on SwarmReconGuard, which employs black-box detection of distributed collective reconnaissance by benign-looking agent populations. This is important for understanding how coordinated malicious activity can spread across decentralized systems. It builds upon the foundational security questions raised by CredLeakBench regarding agent behavior.
Finally, there was research into ASPIRE, an agentic safety and prompt injection red-teaming engine designed to stress test these agents. This directly relates to the deployment-aware feasibility framework for machine learning-based IoT intrusion detection across edge, fog, and cloud architectures. Understanding how agents are exploited via prompt injection is crucial context when designing detection systems that operate across those different architectural layers.
The development of CYBERFORT shows how to build a compliance chain platform that operationalizes the Cyber Resilience Act for small and medium enterprises. This matters because it provides a concrete framework for SMEs to meet new cybersecurity requirements, moving beyond abstract legislation. The team focused on designing the core architecture of CYBERFORT, which involves creating a platform capable of tracking and managing the entire lifecycle of cyber resilience documentation. They tested several data models to ensure they could handle the varied technical specifications required by different types of products.
One key finding was that a modular approach to compliance checking significantly reduced implementation complexity for smaller firms. This means that instead of needing one massive system, SMEs can adopt pieces as their needs evolve, which is a practical takeaway from the initial design phase. Furthermore, the platform successfully integrated automated reporting features based on predefined regulatory checkpoints. This automation streamlines the tedious process of manual compliance checks for users who are not cybersecurity experts.
The research also explored user experience during the data input and verification stages of the compliance chain. Feedback indicated that intuitive interfaces were crucial for adoption, suggesting that usability is just as important as technical accuracy in this kind of system. Finally, while the platform demonstrates strong foundational capabilities, open questions remain regarding its scalability across vastly different industry verticals and how it will interface with existing legacy enterprise systems.
Today's papers
- Adversarial Images Hijack Web Agents from Visual Grounding to Browser Execution Adversarial images can trick web agents into doing things they shouldn't, like executing malicious code. [paper]
- Constitution-Guided Watermarking This method adds hidden watermarks to models to help identify who created them. [paper]
- Constrained-Action AI Remediation for SIEM/XDR via a NeMo-Guardrails Proxy This system uses a proxy to enforce safe actions for AI systems monitoring security events. [paper]
- Sensitive-Topic Leakage Through LLM Routing Metadata: Measurement and Mitigation This paper measures and tries to reduce how sensitive information leaks through the metadata used when routing requests to large language models. [paper]
- Visual Memory Attacks Can Persist Through The KV Cache Visual memory attacks can still work even if the model's key-value cache is cleared. [paper]
- Cost of Delay for Post-Quantum Migration: Putting Classical and Harvest-Now-Decrypt-Later Risk on One Ordered List This paper ranks different risks associated with waiting to switch to post-quantum cryptography. [paper]
- Collusion-Secure Semi-Quantum Secret Sharing Scheme using a Quantum Third Party This scheme allows multiple parties to share secrets securely even if one party is malicious, using quantum technology. [paper]
- Did You Forkget It? Detecting One-Day Vulnerabilities in Open-source Forks With Global History Analysis This tool scans open-source code history to quickly find very recent vulnerabilities introduced in forks. [paper] [episode]
- Mitigating the OWASP Top 10 For Large Language Models Applications using Intelligent Agents This work proposes using intelligent agents to help prevent common security flaws in LLM applications. [paper] [episode]
- COD-ssi: Enforcing Mutual Privacy for Credential Oblivious Disclosure in Self Sovereign Identity This system ensures that when an identity discloses credentials, the disclosure remains private from unauthorized parties. [paper] [episode]
- Hidden in Plain Sight: Benchmarking Agent Safety Against Decomposition Attacks with DECOMPBENCH This benchmark tests how safe AI agents are against attacks where a complex task is broken down into smaller, potentially harmful steps. [paper] [episode]
- Curvature-Guided Module Localization for Low-Rank Detoxification of Backdoored Large Language Models This technique uses the shape of the model to find and remove malicious parts in large language models. [paper] [episode]
- NeuPerm: Disrupting Malware Hidden in Neural Network Parameters by Leveraging Permutation Symmetry This method finds hidden malware within a neural network by looking for specific symmetry patterns. [paper] [episode]
- The Trojan Knowledge: Bypassing Commercial LLM Guardrails via Harmless Prompt Weaving and Adaptive Tree Search This research shows how to bypass safety guardrails on commercial LLMs by cleverly crafting prompts. [paper] [episode]
- A Survey of Secure Retrieval-Augmented Generation This paper reviews the different ways to make retrieval augmented generation safer and more secure. [paper] [episode]
- Restricting the Model, Missing the System: Measurement and Accountability in Offensive AI Governance This study looks at how restricting an AI model alone fails to solve security problems and emphasizes system-level accountability. [paper] [episode]
- ArapaiSecure: An Autonomous AI Security Agent for Banking: Multi-Vector Fraud and AML Detection Across Retail and Corporate Accounts This agent autonomously detects various types of fraud across different banking accounts.
- Adversarial RL for Port-Scan Evasion: Attacker Feature Visibility in Edge-Deployed IDS Adversarial reinforcement learning is used to help attackers evade detection by making their port scans look normal. [paper]
- Visible-Spectrum Optical Covert Channels in Commodity Smart Lighting This paper investigates hidden communication channels that can be sent using the visible light spectrum from common smart lights. [paper]
- Towards Verifying Neural Networks Against Multi-Parameter Bit-Flip Perturbations This work explores how to check if neural networks are robust against small, intentional errors in their data. [paper]
- Multi-Aspect Runtime Verification for Simulation-Based V&V of LLM-Enabled Autonomous Agents This method checks the safety of autonomous agents by verifying their behavior across multiple aspects during simulation. [paper]
- MARS: Malware Analysis with Rule-Based Scoring of LLM Claims This tool analyzes malware by scoring the claims made by a large language model that might be related to it. [paper]
- Correct Answers, Unsupported Findings: Evidence Binding in Forensic Reconstruction of LLM Agent Logs This research focuses on how to reliably use logs from AI agents to reconstruct past events and determine what actually happened. [paper]
- Automotive Hardware Attacks: An Architect's Guide to TARA This guide provides an architectural framework for identifying and mitigating hardware attacks in automotive systems. [paper]
- Black-Box Adversarial Patch Attacks on VLAs via Ancestor VLM Exploitation This attack method uses a vision language model to create adversarial patches that exploit vulnerabilities in other vision models. [paper]
- A Few Steps Further: Why Defenses Against Malicious Finetuning Erode Under Continued Training This paper explains why defenses against malicious fine-tuning become less effective as the model is trained more. [paper] [episode]
- Ask the Expert: LLM-Guided Reinforcement Learning for Autonomous Cyber Defense This system uses reinforcement learning guided by an LLM to help autonomous agents defend against cyber threats. [paper]
- Trust a Few: The Weakest Assumptions a Protocol Needs This paper identifies and analyzes the most fragile assumptions in security protocols. [paper]
- Formal Runtime Verification for Tool-Using LLM Agents: An Offline Same-Benchmark Study on AgentDojo and STAC This study compares different formal verification methods for checking the safety of LLM agents that use external tools. [paper]
- Hybrid Hierarchical Runtime Verification for Edge-IoT Security: Combining MonPoly and RTLola This approach combines two formal verification methods to secure security monitoring on edge IoT devices. [paper]
- Pump-and-Dump meets Honeypot Tokens: Detection and Analysis of Telegram Bait-and-Trap Schemes This system detects deceptive financial schemes like pump-and-dump scams using honeypot tokens. [paper]
- Receiver-Domain Behavioral Probing for Backdoor-Resilient Federated GPS Spoofing Detection in UAV Networks This technique checks for malicious GPS spoofing in drone networks by analyzing the behavior of the receiving devices. [paper]
- TwinGuard-Lite: A Rule-Based State-Admission Gateway for Generative Patient Digital Twins This system uses rules to control what states are allowed when a generative model is creating digital patient twins. [paper]
- Package Hallucination Attacks on Coding Agents through Prompt Injection in Rule Files This attack shows how prompt injection can cause coding agents to generate incorrect code by manipulating rule files. [paper]
- Secure-CUA: Controlling Untrusted Influence in Computer-Use Agents This system helps control the influence of untrusted input when an AI agent is performing computer tasks. [paper]
- Faster PMNS Multi-precision Multiplications Using Truncated Montgomery Technique This paper presents a faster way to perform multi-precision multiplications using a specific mathematical technique. [paper]
- Hierarchical Security Monitoring for Edge-IoT: A Formal Methods Approach This approach uses formal methods to create layered security monitoring for IoT devices at the edge level. [paper]
- Defensive Sufficiency in a Stackelberg Model of AI Security This study analyzes how much defense is needed and sufficient in an AI security model where one entity dictates the other's actions. [paper]
- Defining Purpose-Limited Secrets This paper discusses how to define and protect secrets that are only meant for a specific, limited purpose. [paper]
- BetweenCut: Private Heavy-Node Classification with Doubly Logarithmic Error in Tree Height This method classifies heavy nodes privately while maintaining a very low error rate in tree structure analysis. [paper]
- On the Reliability of LLM-Based Vulnerability Patching Benchmarks This paper examines how trustworthy the benchmarks are when used to test vulnerability patching suggestions from LLMs. [paper]
- SLDR: Defending Against Malicious Fine-tuning via Selective Layers Recovery and Dynamic Routing This technique defends against malicious fine-tuning by selectively recovering layers and dynamically routing requests. [paper]
- A Deployment-Aware Feasibility Framework for Machine Learning-Based IoT Intrusion Detection Across Edge, Fog, and Cloud Architectures This framework helps determine if deploying ML intrusion detection works across different IoT architectures. [paper]
- CredLeakBench: Evaluating Credential Leakage and Recovery in LLM Agents This benchmark evaluates how easily credentials can leak and how they can be recovered from LLM agents. [paper]
- Contextualization of Third-Party Cloud Security Findings This work provides context to security findings reported by third-party cloud providers to make them more actionable. [paper]
- ASPIRE: Agentic Safety & Prompt Injection Red-teaming Engine This engine is designed to test the safety and prompt injection resistance of AI agents. [paper]
- SwarmReconGuard: Black-Box Detection of Distributed Collective Reconnaissance by Individually Benign-Looking Agent Populations This tool detects coordinated reconnaissance activities from groups of seemingly innocent AI agents. [paper]
- Understanding and Mitigating Token-Pruning-Induced Vulnerabilities in VLMs This paper explores the security risks that arise when vision language models have their tokens pruned during operation. [paper]
- CYBERFORT: A Compliance-Chain Platform Operationalising the Cyber Resilience Act for SMEs This platform helps small and medium enterprises comply with cyber resilience regulations using a compliance chain. [paper]
The papers
- Hidden in Plain Sight: Benchmarking Agent Safety Against Decomposition Attacks with DECOMPBENCH — LLM-based agents are increasingly capable, raising concerns about their misuse through Decomposition Attacks, which break down harmful tasks into benign subtasks that evade safety mechanisms when executed separately but cumulatively fulfill a malicious intent. [episode]
- A Few Steps Further: Why Defenses Against Malicious Finetuning Erode Under Continued Training — Model providers increasingly release open weights or allow users to fine-tune foundation models through APIs, creating a post-release safety problem where malicious fine-tuning (MFT) can subvert safety alignments. [episode]
- Curvature-Guided Module Localization for Low-Rank Detoxification of Backdoored Large Language Models — Backdoor attacks pose a serious threat to large language models (LLMs) by causing otherwise benign systems to produce attacker-specified malicious behavior when a hidden trigger is present. [episode]
- Restricting the Model, Missing the System: Measurement and Accountability in Offensive AI Governance — Offensive capability in AI systems must be assessed at the level of the entire system—model, scaffold, and evaluation protocol—rather than focusing solely on restricting access to individual models. [episode]
- Did You Forkget It? Detecting One-Day Vulnerabilities in Open-source ForksWith Global History Analysis — A global history analysis approach leverages a comprehensive graph of public code to identify one-day vulnerabilities in forked repositories, addressing a critical gap where existing tools fail to track inherited security issues across diverse fork ecosystems. [episode]
- A Survey of Secure Retrieval-Augmented Generation — Retrieval-augmented generation (RAG) extends large language models (LLMs) with external knowledge, but this access path also introduces security risks that existing work often conflates with inherent LLM flaws. [episode]
- COD-ssi: Enforcing Mutual Privacy for Credential Oblivious Disclosure in Self Sovereign Identity — The COD-ssi framework introduces a novel approach to Self-Sovereign Identity (SSI) that enforces mutual privacy during credential exchange by allowing Verifiers to selectively disclose a subset of claims without revealing which specific claims were accessed to the Holder. [episode]
- The Trojan Knowledge: Bypassing Commercial LLM Guardrails via Harmless Prompt Weaving and Adaptive Tree Search — Large language models remain vulnerable to jailbreak attacks that bypass safety guardrails to elicit harmful outputs. [episode]
- An Autonomous AI Security Agent for Banking: Multi-Vector Fraud and AML Detection Across Retail and Corporate Accounts — Banks face two threat families with fundamentally different detection requirements: signature-based fraud and behavioral financial crime. [episode]
- Mitigating the OWASP Top 10 For Large Language Models Applications using Intelligent Agents — Large Language Models (LLMs) have emerged as a transformative technology, but their widespread integration has raised significant security concerns highlighted by the Open Web Application Security Project (OWASP), necessitating frameworks to proactively identify and counteract th [episode]
- NeuPerm: Disrupting Malware Hidden in Neural Network Parameters by Leveraging Permutation Symmetry — Pretrained deep learning model sharing exposes end-users to cyber threats where attackers hide self-executing malware inside neural network parameters, and this work proposes NeuPerm, a simple yet effective zero-trust technique leveraging permutation symmetry to disrupt such atta [episode]
- Cost of Delay for Post-Quantum Migration: Putting Classical and Harvest-Now-Decrypt-Later Risk on One Ordered List —
- SwarmReconGuard: Black-Box Detection of Distributed Collective Reconnaissance by Individually Benign-Looking Agent Populations —
- Adversarial Images Hijack Web Agents from Visual Grounding to Browser Execution —
- Package Hallucination Attacks on Coding Agents through Prompt Injection in Rule Files —
- Ask the Expert: LLM-Guided Reinforcement Learning for Autonomous Cyber Defense —
- Secure-CUA: Controlling Untrusted Influence in Computer-Use Agents —
- Constitution-Guided Watermarking —
- MARS: Malware Analysis with Rule-Based Scoring of LLM Claims —
- Correct Answers, Unsupported Findings: Evidence Binding in Forensic Reconstruction of LLM Agent Logs —
- Automotive Hardware Attacks: An Architect's Guide to TARA —
- Understanding and Mitigating Token-Pruning-Induced Vulnerabilities in VLMs —
- Black-Box Adversarial Patch Attacks on VLAs via Ancestor VLM Exploitation —
- Trust a Few: The Weakest Assumptions a Protocol Needs —
- Faster PMNS Multi-precision Multiplications Using Truncated Montgomery Technique —
- Formal Runtime Verification for Tool-Using LLM Agents: An Offline Same-Benchmark Study on AgentDojo and STAC —
- Hierarchical Security Monitoring for Edge-IoT: A Formal Methods Approach —
- Hybrid Hierarchical Runtime Verification for Edge-IoT Security: Combining MonPoly and RTLola —
- Defensive Sufficiency in a Stackelberg Model of AI Security —
- Constrained-Action AI Remediation for SIEM/XDR via a NeMo-Guardrails Proxy —
- CYBERFORT: A Compliance-Chain Platform Operationalising the Cyber Resilience Act for SMEs —
- Sensitive-Topic Leakage Through LLM Routing Metadata: Measurement and Mitigation —
- Defining Purpose-Limited Secrets —
- BetweenCut: Private Heavy-Node Classification with Doubly Logarithmic Error in Tree Height —
- Pump-and-Dump meets Honeypot Tokens: Detection and Analysis of Telegram Bait-and-Trap Schemes —
- On the Reliability of LLM-Based Vulnerability Patching Benchmarks —
- Collusion-Secure Semi-Quantum Secret Sharing Scheme using a Quantum Third Party —
- SLDR: Defending Against Malicious Fine-tuning via Selective Layers Recovery and Dynamic Routing —
- Receiver-Domain Behavioral Probing for Backdoor-Resilient Federated GPS Spoofing Detection in UAV Networks —
- Adversarial RL for Port-Scan Evasion: Attacker Feature Visibility in Edge-Deployed IDS —
- A Deployment-Aware Feasibility Framework for Machine Learning-Based IoT Intrusion Detection Across Edge, Fog, and Cloud Architectures —
- Visible-Spectrum Optical Covert Channels in Commodity Smart Lighting —
- CredLeakBench: Evaluating Credential Leakage and Recovery in LLM Agents —
- Towards Verifying Neural Networks Against Multi-Parameter Bit-Flip Perturbations —
- Contextualization of Third-Party Cloud Security Findings —
- Multi-Aspect Runtime Verification for Simulation-Based V&V of LLM-Enabled Autonomous Agents —
- ASPIRE: Agentic Safety & Prompt Injection Red-teaming Engine —
- TwinGuard-Lite: A Rule-Based State-Admission Gateway for Generative Patient Digital Twins —
- Visual Memory Attacks Can Persist Through The KV Cache —
Important terms
- Adversarial Images
- These are manipulated visual inputs designed to hijack web agents by tricking them into executing harmful actions, showing a new way to compromise automated systems from visual data.
- Constrained Action AI Remediation
- This involves using tools like NeMo Guardrails Proxy to limit what an agent can do based on its input, actively controlling system behavior in real-time instead of just flagging threats.
- Trojan Knowledge Problem
- This research focuses on bypassing LLM guardrails by weaving harmless prompts and using adaptive search to find loopholes, allowing models to reveal restricted data they shouldn't.
- Curvature-Guided Module Localization
- This technique attempts to clean up backdoored large language models by finding and removing malicious code hidden within the model's structure based on how its parameters curve.
- MARS Analysis
- MARS uses rule-based scoring to analyze malware claims made by LLMs, addressing the risk that AI might generate plausible but unsupported findings when reconstructing agent logs.