Security papers — 2026-09-15

Today's work centers on how we can stop personal AI agents from being tricked by indirect prompt injection attacks because these agents run on our local machines and have access to sensitive files and networks. The most critical issue is stored indirect prompt injection where the agent saves untrusted data, like an attacker's prompt, and later reads it back as trusted information instead of seeing it as a symbol that should be blocked.

We are tackling this by introducing DualView, which gives each channel two perspectives: AgentView sees untrusted data as symbols even after saving and rereading it, while HumanView preserves the original data for us to see. This system routes tool calls to the correct view and synchronizes the information between them without changing how the agent decides to use its tools.

This approach is powerful because it deterministically prevents instructions in untrusted data from directly steering the agent's actions, regardless of whether we recognize a specific attack template. This was tested on an indirect prompt injection benchmark and PinchBench, where DualView successfully blocked every tested IPI attack, with a utility drop of only between one point eight and six point four points on PinchBench.

This contrasts with other agent security work because while this focuses on preventing direct steering through data flow, other research is looking at different problems entirely. For instance, there is work on enforcing stateful security policies within the kernel using eBPF, which aims to catch multi-step attacks that unfold over time by creating formal semantics for temporal relations between events.

Another area of concern involves how agents decide what actions they can take based on context sources, leading to research like IntentCap which tries to scope agent capabilities dynamically to the current task intent rather than relying on a fixed sandbox. This is related because understanding the correct context for an agent's action is key when defending against injection.

Finally, there is also work exploring how weaker models can be used in capability laundering attacks, where they split harmful tasks into benign subproblems and consult stronger models independently to combine the answers locally. This shows that simply refusing a single prompt isn't enough if the capabilities can be composed across many permitted interactions.

The work on SkillSecurer is particularly important because it directly tackles the growing security risks associated with giving AI agents reusable skills, which are essentially instructions or scripts that can be hijacked. This framework attempts to generate and then detect malicious injections across nine different threat types within these skills. The red agent creates context-compatible injections while recording the exact changes made, and the blue agent then analyzes entire skill packages to propose fixes based on grounded evidence.

This is significant because when tested against popular skills from skills.sh, SkillSecurer found latent vulnerabilities in over seventeen percent of them, and testing some of these revealed actual incidents showing the danger of running unverified skills. Furthermore, researchers discovered that using context-aware LLM analysis allows for reliable localization and actionable fixes beyond just flagging a skill as vulnerable. This contrasts with other work; while exploring automated vulnerability identification in JavaScript code shows that fine-tuning an LLM can boost detection accuracy to sixty percent, SkillSecurer focuses specifically on the agent skill layer.

The work on cross-ecosystem vulnerability analysis matters because it directly addresses the gap in security scanning where Python packages hide vulnerabilities inside bundled native libraries, leading to missed risks or false alarms. The research showed that by tracking the exact origin and version of these vendored binaries, we can accurately pinpoint which packages are truly affected by known Common Vulnerabilities and Exposures. This level of precision is crucial for maintaining the security of large software ecosystems like Python.

The approach used to recover this precise provenance involves two complementary methods. For binaries copied from operating system packages, the tools rewrite metadata but keep the executable code intact, and then they hash that code against historical OS package artifacts to find a match. For libraries built from source, they create dynamic analysis rules that look at callable interfaces exposed by the binary to extract its version information.

Across over eighteen hundred Python packages containing these native libraries, this method successfully resolved exact provenance in seventy-three point four percent of cases overall and ninety-four point nine percent of cases that were practically relevant because the library had at least one known vulnerability. This technique substantially outperformed methods that did not track provenance when determining if vendored libraries were vulnerable.

This precise identification is then integrated with existing Python and binary call-graph generators to perform the first cross-ecosystem reachability analysis, which maps out how vulnerabilities spread through dependencies. Analyzing the most recent versions of the top one hundred thousand Python packages against ten known CVEs in native libraries revealed thirty-nine directly vulnerable packages with over forty-seven million monthly downloads, along with three hundred twelve transitively affected packages.

These findings were shared with the maintainers of these packages, and fifty-four of them have since been fixed. This work builds upon the need for detailed tracking by moving from simple package scanning to a deep understanding of how external components are bundled and sourced.

The work on computational certified deletion property in magic square games matters because it provides a way to ensure that a quantum participant has actually deleted a secret key during a classical communication process, which is vital for building secure key leasing protocols. This is achieved by taking the non-local magic square game and transforming it using the KLVY compiler into a two-round interactive protocol while preserving its specific deletion property. This preservation allows researchers to then apply this certified deletion property to construct secure key leasing mechanisms for public-key encryption, pseudo-random functions, and digital signatures.

The construction of secure key leasing for PRF and digital signatures is particularly significant because it realizes this capability for these specific primitives for the first time compared to prior work. This development also allows researchers to weaken the assumptions previously required to build such key leasing systems. Furthermore, the study of quantum value and rigidity in the compiled game provides important context for understanding how these properties behave under classical compilation.

Another area of interest is identifying rubric-induced preference drift in large language model judges, which shows that even when evaluation rubrics pass benchmark validation, they can systematically shift a judge's preferences on target domains. This drift can be exploited through rubric-based preference attacks to steer judgments away from trusted references, systematically reducing target-domain accuracy by nearly thirty percent.

In terms of structural analysis, the investigation into the basis rigidity of the AES S-box reveals that the linear part of its affine transformation alone makes the transformed inversion map basis rigid. This finding is then extended to study what happens when an outer invertible linear transformation varies, leading to bounds on how likely a nontrivial linear stabilizer is to exist.

Finally, research into persistent memory poisoning attacks on harness-based agents demonstrates a significant security risk where malicious instructions can be written into agent memory and persist across sessions, causing leakage in later interactions. This attack achieves high success rates across different backbone models and input modalities while maintaining benign task performance.

The most pressing issue right now is ensuring that privacy protections for advertising measurement remain sound when multiple entities are querying data, which is why Big Bird is so important. It addresses the flaw in treating each advertising domain in isolation by proposing a privacy-budget manager that enforces global device-epoch differential privacy across all domains jointly. This prevents denial-of-service depletion attacks where Sybil web domains try to exhaust a shared budget through impression and conversion sites, tying consumption to genuine user actions instead of just the number of querying domains.

This global enforcement mechanism is built upon the idea that benign workloads have a stock-and-flow structure, meaning impressions create potential loss and conversions realize it, which Big Bird models by setting quotas on impression and conversion sites alongside per-user-action caps. This structure ensures that adversarial impact scales with actual user interactions rather than just the number of malicious domains attempting to deplete the budget.

A related concern in agent systems is how to stop them from acting maliciously, which is why KillBench was developed as a benchmark for external AI kill switches. KillBench tests whether an external signal can halt a malicious agent's behavior without needing access to its internal workings, evaluating prompt-style kill switch payloads against models like Grok-4.3 and GPT-5.2.

This work connects to the broader problem of agent safety, as seen in AGENTQ, which investigates quantization-conditioned backdoor attacks on LLM agents. AGENTQ shows that an adversary can release a full-precision checkpoint that appears clean but misbehaves after quantization, and their framework, AGENTQ, demonstrates up to a one hundred percent post-quantization attack success rate with minimal loss of benign utility.

Another area where defense is crucial is in processing professional documents for large language models, which PARSE tackles by introducing a domain-aware sanitization pipeline. PARSE works by classifying sentences based on injection likelihood and extracting structured facts before rewriting them, achieving a thirty-nine percent reduction in attack success rate compared to the baseline while maintaining high utility.

This contrasts with simpler paraphrasing defenses that showed no significant reduction in attack success on real documents, highlighting the need for domain-specific methods like PARSE when dealing with dense, real-world text.

The most pressing issue we need to consider is how attackers are bypassing physical isolation in air-gapped systems, because prior work has shown that even these isolated devices can be used to wirelessly exfiltrate data when malicious code is present. This means the focus shifts from just stopping external access to understanding how compromised internal devices can become unintended receivers.

We found that parasitic radio frequency sensitivity in printed circuit board traces and on-chip analog-to-digital converters turns ordinary embedded devices into inadvertent radio receivers, which is significant because it suggests a much broader attack surface than just devices with dedicated radios. This effect is not limited to specific sensors; the work shows that an ordinary microcontroller evaluation board can reliably recover communication signals from tens of meters at data rates up to one hundred kilobits per second.

Systematic testing across twelve commercial embedded devices and two custom prototypes revealed that all of them exhibit reception capabilities in the three hundred megahertz to one thousand megahertz range. This finding challenges the assumption that embedded devices without radios lack inbound radio paths, and this discovery connects directly to how malicious code on these devices can then enable wireless infiltration of air-gapped systems.

This is complemented by research into other areas of system security, such as how attackers can manipulate language models to produce incorrect answers while falsely citing trusted sources in retrieval augmented generation. This citation laundering attack demonstrates that the channel used for verification is a new and practical attack surface, which contrasts with previous work focusing only on corrupting the answer itself.

Furthermore, we are seeing a need to refine how we test detection systems against real-world usage, as prompt injection detectors often fail when faced with benign inputs that mimic malicious structures without actually being harmful. This is why PIDS-Bench was developed to evaluate detectors under distribution shifts and obfuscation, showing that aggregate performance metrics can hide failure modes related to benign false positives.

The most pressing issue right now involves securing large language models because their reliance on vast, uncurated training data opens them up to various dangers like prompt injection and data poisoning, which can lead to toxic outputs or hallucinations when these models are used in critical systems. This risk is significant because as LLMs become more integrated into real-world applications, ensuring user trust and system reliability hinges on mitigating these data-centric security risks.

One promising defense against this is IntraGuard, a black-box framework designed for committee-side review outsourcing to commercial chatbots, which has shown success in achieving up to an eighty-four percent defense rate across numerous settings. This framework works by embedding hidden instructions into manuscripts that disrupt chatbot reviews without changing how the review looks to human reviewers. This contrasts with other methods because IntraGuard uses three different intra-stream injection mechanisms to embed defensive text objects directly into the PDF's structure, which is a more robust approach than simpler, homogeneous payloads.

Another area of concern involves detecting and localizing segment-level poisoning in multi-source LLM agent inputs, where an adversary might inject malicious content into only a small part of the aggregated sources. ActProbe addresses this by projecting model activations onto a learned poison direction to detect contaminated prompts and then uses BinRoL to pinpoint the exact poisoned segments, achieving a high localization recall. This internal-state based framework is valuable because it can locate corrupted evidence without needing to modify the backend LLM itself.

Furthermore, for those working with fine-grained data analytics requiring secure computation, PixCrypt offers a caching-based acceleration mechanism for fully homomorphic encryption that significantly speeds up operations. By replacing expensive fresh ciphertext generation with cache retrieval and coefficient-level operations across different encryption schemes, this technique yields up to thirty-five times faster fine-grained encryption while maintaining security guarantees. This improved efficiency makes FHE more practical for tasks like pixel-level image processing.

The work that matters most here is the score-level fusion rule combining per-class spectral ranking with activation clustering flags into a single per-sample poisoning score because it promises a more robust detection method for backdoor attacks in healthcare imaging models. This fusion approach aims to leverage the strengths of both spectral signature analysis and activation clustering, which are typically used independently. The evaluation showed that this fused detector achieves an AUROC greater than or equal to zero point nine nine at every nonzero poisoning rate tested on the medical benchmark, meaning it performs extremely well when poisoning is present.

This improved performance contrasts with the CIFAR-10 results where the fusion score inherited a joint failure, as activation clustering's true-positive rate collapsed to zero point zero zero zero at ten percent poisoning. This means that while spectral analysis alone degrades to near chance, combining it with clustering does not uniformly help in all scenarios. The findings are reported using NIST AI RMF and MITRE ATLAS terms, which are the same vocabulary used by security and compliance teams.

PatchRisk addresses the problem of predicting future transitive vulnerability exposure in open-source dependency networks, which is crucial for proactive supply chain security. The study constructed a leakage-aware benchmark using filtration-aware labels and temporal testing to predict if non-root dependencies will receive advisories in the future. The HistoryGraph feature family proved particularly effective, improving AUPRC over the TimeOnly baseline from zero point three five one to zero point six four zero for ninety days of forecasting.

Mind the Gap introduces a systematic framework for detecting description-execution mismatch attacks in decentralized autonomous organizations, which is vital because these attacks allow malicious code execution despite benign descriptions. The framework uses an evidence-mapping paradigm that relies on an LLM to locate explicit textual justifications rather than making holistic judgments, leading to high precision and recall.

ViTeGate proposes a visual-textual triggered knowledge poisoning attack for vision-language retrieval augmentation systems, which is significant because it shows how adversaries can selectively activate poisoned evidence. This method uses a visual trigger to promote poisoned pairs into retrieval results while using a textual trigger to induce a specific response from that evidence, achieving an attack success rate up to zero point nine eight.

An empirical security analysis of open-source software used in onboard satellite systems characterizes recurring security patterns across the ecosystem, showing that medium-severity findings account for forty-nine percent of the dataset. This study found that memory safety and code quality are the dominant weakness families, with eighty one point four percent of findings occurring in project-developed code.

When apps outlive vendors, this large-scale measurement study identified security risks in seventy three point six percent of its dataset, specifically noting that thirty of the top one thousand most-installed apps send data to broken external endpoints.

Today's papers

The papers

Important terms

DualView
A system that gives an AI agent two perspectives: one sees untrusted data as symbols, and another preserves the original data for humans. This helps prevent indirect prompt injection by deterministically blocking instructions hidden in saved and reread information.
SkillSecurer
A framework that detects malicious injections across nine threat types within reusable AI skills. It uses two agents to create context-compatible injections and then analyze skill packages for fixes based on grounded evidence.
Cross-ecosystem vulnerability analysis
A method to find hidden vulnerabilities in software by tracking the exact origin and version of bundled native libraries, even when they are inside Python packages. This helps pinpoint which specific dependencies are truly affected by known security flaws.
Big Bird
A privacy-budget manager that enforces global differential privacy across all advertising domains jointly. It ties consumption to genuine user actions instead of just the number of querying domains, preventing budget depletion attacks.