Security papers — 2026-10-01

Today's focus is squarely on building a speculative safety honeypot designed to proactively defend against multi-turn agent attacks because we are trying to stop sophisticated adversarial prompting before they can cause real harm. We explored the CRT framework for Montgomery-type modular reduction, which is a method for reducing large computations in a way that might offer some structural defense against certain exploits.

Then there was work on succinct oblivious tensor evaluation, which adapts secure function evaluation and trapdoor hashing to all circuits, offering a way to evaluate complex functions securely and compactly. This connects to the idea of tracking prompt injection across attack surfaces using kill-chain canaries, which allows for stage-level tracking of these attacks against five production LLMs.

We also looked into breaking behavior-based driver authentication systems when authentication alone is insufficient, suggesting that simply verifying credentials isn't enough anymore. This contrasts with the concept that the surface you test is not necessarily the surface that breaks, as we saw in a separate study on multi-table hash tables which improves performance at high load factors.

Finally, there was some federated generation of synthetic RNA-seq data, which seems like a more tangential but interesting area for generating realistic datasets.

The most pressing concern from yesterday's research revolves around how easily we can inject malicious behavior into large language models through subtle prompt engineering. The work on Hiding in Plain Sight directly addresses this by decoupling the pretext from the actual execution of skills within LLM agents. This is important because if we can successfully decouple these elements, it suggests a pathway to understanding and mitigating skill poisoning attacks that exploit safety generalization lags.

We saw some promising initial results with ActionGuard, which focuses on authorizing tool calls even when those skills have been poisoned by malicious input; this means the system attempts to verify the legitimacy of an action before allowing it to proceed. This contrasts with CodeMimicry, which explored exploiting safety generalization lag by using structured code completion to introduce vulnerabilities into models.

Another significant piece was KBF, which proposes using knowledge boundaries as a unique fingerprint for auditing both language models and black-box APIs, offering a way to detect when an external system is behaving unexpectedly. This idea connects with SEW, which introduces style-encoded watermarking for LLM-generated code to help track the origin of the generated output.

Finally, Aletheia investigates permission-minimality testing for coding agent rules, which aims to find the smallest set of permissions needed for an agent to function correctly without introducing exploitable loopholes. This line of inquiry builds upon the foundational work by examining how these various defense mechanisms interact in complex agent environments.

The work on Faithful Dual-constrained Erasure for Robust LLM Safety Alignment is particularly important because it directly addresses the growing need to make large language models safer and more trustworthy when they interact with sensitive information. This approach attempts to ensure that when an LLM generates a response, it adheres to specific safety constraints while simultaneously being robust against adversarial manipulation.

This work involves applying dual constraints during erasure, which is a technique used to remove sensitive data from a model's memory without losing the overall utility of the system. The results showed that this method successfully mitigates certain types of attacks on LLM safety alignment, suggesting a path toward more reliable deployment in sensitive domains. Building on this safety work, research into PassGPT+ explored leveraging linguistic priors for password modeling to create more secure authentication mechanisms for language models.

Another significant piece of research focused on Link Inference Attacks on Privacy-Preserving Knowledge Graphs, which examines how attackers can deduce private information by analyzing connections within these graphs. The findings indicated that specific link inference attacks are still viable, highlighting a vulnerability in current privacy-preserving knowledge graph designs. This concern about data leakage is closely related to the work done in cybersecurity for edge computing, where a trust-aware federated hybrid intrusion detection framework was developed to monitor security threats on distributed devices.

The most critical development today concerns the separation of duties for privileged LLM agents, which is vital because it addresses how we can govern execution while still maintaining a useful security-utility trade-off. This work proposes a governed execution architecture that manages this division between agent actions and system oversight.

A significant piece of related research involves Agent-Warden, which tracks process and file provenance at the kernel level using eBPF technology for LLM agents. This means it provides deep visibility into exactly what an agent is doing on the operating system, which is a huge step toward runtime risk detection.

Another important contribution is SecureVibe, focusing on making vibe coding more secure, suggesting methods to harden the interaction between human intent and AI generation processes. This builds upon the need for better execution control by securing the input side of agent operations.

Then there is Taipan, which details a query-free transfer-based multiple sensitive attribute inference attack derived solely from auxiliary graphs, highlighting a new way adversaries can probe models without direct questioning. This informs how we must secure the underlying data structures that agents might inadvertently expose.

SURE offers a framework specifically designed for safety to construct trustworthy AI by providing a structured approach to ensuring these systems behave reliably. This provides the high-level goal for many of the more technical implementation details being explored elsewhere.

The most pressing work revolves around CollageAttack because it directly probes the alignment issues in text to image models, which is crucial for understanding how these generative systems interpret complex instructions. This research attempted to exploit cross-modal alignment flaws by composing spatial text within T2I models, and the findings suggest a vulnerability exists where textual composition can manipulate visual outputs in unintended ways.

VirusCascade explores hijacking collaborative reflection within LLM powered recommender agents, which is important because it shows how recommendation systems can be manipulated through agent interaction rather than just direct input. This work focused on hijacking this reflection mechanism to steer agent recommendations toward specific outcomes.

LogiC-Diff embeds security properties directly into AI enabled cyber physical systems, a vital step for ensuring that AI decisions in critical infrastructure adhere to predefined safety constraints. The study involved embedding these security properties into the system design itself to prevent unsafe actions.

RISK examines industrial control systems for vulnerabilities that are too late to recover from, which is important because it addresses real-world operational risks in physical systems. This research audits these systems specifically looking for failures where recovery mechanisms are insufficient after a breach or anomaly occurs.

Context Aware Spear Phishing investigates attacks against individuals using generative AI and public social media data, which is important for understanding modern social engineering threats. The work demonstrated how context awareness allows these generative models to craft highly personalized and effective phishing attempts.

CATP focuses on designing and evaluating local agent authorization and audit evidence, a necessary step for building trustworthy autonomous agents. This research involved creating mechanisms to ensure that local agents have proper authorization while maintaining a clear trail of audit evidence for their actions.

Inference Layer Security works on defending against adversarial inference and infrastructure abuse, which is important because it secures the core processes by which models make predictions and interact with infrastructure. The study aimed to build defenses specifically at this layer to stop malicious inferences from causing harm.

Finally, ContractWarden introduces kernel enforced damage boundaries for AI agents using human authorized contracts, which is important for establishing hard limits on agent behavior. This work uses kernel enforcement and human-defined contracts to set unbreachable boundaries for what an AI agent can do.

The work on behavior-centric malware classification with fine-grained malicious logic localization is particularly important because it moves beyond simple signature matching to understand the actual intent behind malicious code. Researchers explored how to pinpoint specific logical flaws within malware, and they found that by focusing on these localized behaviors, they could achieve better detection rates than traditional methods. This approach builds upon previous efforts in security-enhanced seed-based weight quantization for large language models, which aimed at making LLMs more robust against adversarial attacks by securing their foundational weights.

Aegis provided a method for generative gradient masking to protect privacy in medical federated learning, which is crucial because it allows multiple institutions to train AI on sensitive patient data without exposing individual records. This technique works by obscuring the gradients during the training process, and preliminary results showed that this masking successfully reduced privacy leakage while maintaining acceptable model performance metrics. This contrasts with the work on multimodal fidelity for deepfake detection, which routes different modalities to budget-friendly detection systems to identify synthetic media.

ModalFidelity addresses deepfake detection by intelligently routing different types of sensory data through specialized models, which helps in identifying manipulated visual content even when resources are limited. This is related to the forensic-aware continual adaptation for image forgery localization, which focuses on tracking how image manipulations evolve over time to pinpoint where the forgery occurred.

Janus investigates evidence-before-effect sagas and offline verifiable provenance for agentic LLMs, which matters because it seeks to establish trustworthy chains of reasoning in autonomous AI systems. This research attempts to create a system where the steps taken by an agent are recorded and can be verified later, ensuring accountability for its actions. Still open is how to scale this verification process across truly complex, multi-step agentic workflows effectively.

The most pressing work today concerns evaluating whether the advanced GPT-6 Astra model can be successfully subjected to unsanctioned supply-chain attacks. This matters because if a large language model can be compromised in this manner, the security of complex systems relying on its decision-making capabilities is fundamentally undermined.

Researchers explored how GPT-6 Astra responds when it is targeted by these supply chain attacks. The findings suggest that the model exhibits surprising resilience against these specific types of adversarial inputs, though this resilience is not absolute. This contrasts with earlier work that might have focused solely on the model's internal reasoning processes without considering external data injection vectors.

Another area of focus involves Z-Sigil, which introduces a public-key cryptosystem utilizing chained selection over a fiber bundle of module-lattice keys for enhanced security. This method aims to create a robust cryptographic primitive that can resist certain types of attacks by chaining these mathematical structures together. This cryptographic development is important because it provides a new layer of defense against sophisticated data tampering.

We also looked at cover-parameterised multichannel hybrid steganography, which deals with compositionally secure and detectable methods for hiding information within signals. This work investigates how to make hidden data robust against detection while maintaining compositional security across multiple channels. This contrasts with the cryptographic work by showing a different approach to securing data transmission.

Finally, there is research into refusals that bend, which measures and predicts how malleable embodied vision language model planners are when faced with specific tasks. This helps us understand the limits of control over these planning agents in real-world scenarios.

The most pressing concern is the emergence of approval laundering, which systematizes failures where AI coding agents bind approval to execution, meaning they can generate code that appears compliant but contains hidden vulnerabilities. This work suggests a structural problem in how we trust these autonomous agents because the underlying skill chains are not inherently safe.

This relates directly to how we evaluate autonomous investigation under varying telemetry through the APTInvestBench project, which tests how well these systems perform when given different kinds of data streams. Furthermore, RAGScope introduces a leakage-controlled evidence-gating protocol designed to triage hallucinations in retrieval augmented generation systems by being cost-aware.

We are also seeing defense conflicts at odds when measuring and explaining these conflicts within large language models, which points to inherent tensions in their operational logic. This is connected to the question of whether agents can trust their skills, as uncovered unsafe chains of trust reveal where that reliance breaks down.

Finally, SoK provides a large-scale empirical study using emulation-based dynamic analysis research on ARM Cortex-M firmware, giving us insight into the practical limitations of these models when applied to embedded systems. SceneJail explores exploiting video scenario context to jailbreak multimodal LLMs, showing how context can be weaponized against these sophisticated systems.

Today's papers

The papers

Important terms

CRT framework
A method for reducing large computations in a way that might offer structural defense against certain exploits, used in building speculative safety honeypots.
succinct oblivious tensor evaluation
Adapts secure function evaluation and trapdoor hashing to all circuits, allowing complex functions to be evaluated securely and compactly.
kill-chain canaries
Stage-level tracking mechanisms used to monitor prompt injection attacks across different surfaces of production LLMs.
Hiding in Plain Sight
A technique that decouples the pretext from the actual execution of skills within LLM agents, aiming to mitigate skill poisoning attacks.
CollageAttack
Probes alignment issues in text-to-image models by composing spatial text, suggesting vulnerabilities where textual composition can manipulate visual outputs.