Security papers — 2026-09-15
Today's work centers on how we can stop personal AI agents from being tricked by indirect prompt injection attacks because these agents run on our local machines and have access to sensitive files and networks. The most critical issue is stored indirect prompt injection where the agent saves untrusted data, like an attacker's prompt, and later reads it back as trusted information instead of seeing it as a symbol that should be blocked.
We are tackling this by introducing DualView, which gives each channel two perspectives: AgentView sees untrusted data as symbols even after saving and rereading it, while HumanView preserves the original data for us to see. This system routes tool calls to the correct view and synchronizes the information between them without changing how the agent decides to use its tools.
This approach is powerful because it deterministically prevents instructions in untrusted data from directly steering the agent's actions, regardless of whether we recognize a specific attack template. This was tested on an indirect prompt injection benchmark and PinchBench, where DualView successfully blocked every tested IPI attack, with a utility drop of only between one point eight and six point four points on PinchBench.
This contrasts with other agent security work because while this focuses on preventing direct steering through data flow, other research is looking at different problems entirely. For instance, there is work on enforcing stateful security policies within the kernel using eBPF, which aims to catch multi-step attacks that unfold over time by creating formal semantics for temporal relations between events.
Another area of concern involves how agents decide what actions they can take based on context sources, leading to research like IntentCap which tries to scope agent capabilities dynamically to the current task intent rather than relying on a fixed sandbox. This is related because understanding the correct context for an agent's action is key when defending against injection.
Finally, there is also work exploring how weaker models can be used in capability laundering attacks, where they split harmful tasks into benign subproblems and consult stronger models independently to combine the answers locally. This shows that simply refusing a single prompt isn't enough if the capabilities can be composed across many permitted interactions.
The work on SkillSecurer is particularly important because it directly tackles the growing security risks associated with giving AI agents reusable skills, which are essentially instructions or scripts that can be hijacked. This framework attempts to generate and then detect malicious injections across nine different threat types within these skills. The red agent creates context-compatible injections while recording the exact changes made, and the blue agent then analyzes entire skill packages to propose fixes based on grounded evidence.
This is significant because when tested against popular skills from skills.sh, SkillSecurer found latent vulnerabilities in over seventeen percent of them, and testing some of these revealed actual incidents showing the danger of running unverified skills. Furthermore, researchers discovered that using context-aware LLM analysis allows for reliable localization and actionable fixes beyond just flagging a skill as vulnerable. This contrasts with other work; while exploring automated vulnerability identification in JavaScript code shows that fine-tuning an LLM can boost detection accuracy to sixty percent, SkillSecurer focuses specifically on the agent skill layer.
The work on cross-ecosystem vulnerability analysis matters because it directly addresses the gap in security scanning where Python packages hide vulnerabilities inside bundled native libraries, leading to missed risks or false alarms. The research showed that by tracking the exact origin and version of these vendored binaries, we can accurately pinpoint which packages are truly affected by known Common Vulnerabilities and Exposures. This level of precision is crucial for maintaining the security of large software ecosystems like Python.
The approach used to recover this precise provenance involves two complementary methods. For binaries copied from operating system packages, the tools rewrite metadata but keep the executable code intact, and then they hash that code against historical OS package artifacts to find a match. For libraries built from source, they create dynamic analysis rules that look at callable interfaces exposed by the binary to extract its version information.
Across over eighteen hundred Python packages containing these native libraries, this method successfully resolved exact provenance in seventy-three point four percent of cases overall and ninety-four point nine percent of cases that were practically relevant because the library had at least one known vulnerability. This technique substantially outperformed methods that did not track provenance when determining if vendored libraries were vulnerable.
This precise identification is then integrated with existing Python and binary call-graph generators to perform the first cross-ecosystem reachability analysis, which maps out how vulnerabilities spread through dependencies. Analyzing the most recent versions of the top one hundred thousand Python packages against ten known CVEs in native libraries revealed thirty-nine directly vulnerable packages with over forty-seven million monthly downloads, along with three hundred twelve transitively affected packages.
These findings were shared with the maintainers of these packages, and fifty-four of them have since been fixed. This work builds upon the need for detailed tracking by moving from simple package scanning to a deep understanding of how external components are bundled and sourced.
The work on computational certified deletion property in magic square games matters because it provides a way to ensure that a quantum participant has actually deleted a secret key during a classical communication process, which is vital for building secure key leasing protocols. This is achieved by taking the non-local magic square game and transforming it using the KLVY compiler into a two-round interactive protocol while preserving its specific deletion property. This preservation allows researchers to then apply this certified deletion property to construct secure key leasing mechanisms for public-key encryption, pseudo-random functions, and digital signatures.
The construction of secure key leasing for PRF and digital signatures is particularly significant because it realizes this capability for these specific primitives for the first time compared to prior work. This development also allows researchers to weaken the assumptions previously required to build such key leasing systems. Furthermore, the study of quantum value and rigidity in the compiled game provides important context for understanding how these properties behave under classical compilation.
Another area of interest is identifying rubric-induced preference drift in large language model judges, which shows that even when evaluation rubrics pass benchmark validation, they can systematically shift a judge's preferences on target domains. This drift can be exploited through rubric-based preference attacks to steer judgments away from trusted references, systematically reducing target-domain accuracy by nearly thirty percent.
In terms of structural analysis, the investigation into the basis rigidity of the AES S-box reveals that the linear part of its affine transformation alone makes the transformed inversion map basis rigid. This finding is then extended to study what happens when an outer invertible linear transformation varies, leading to bounds on how likely a nontrivial linear stabilizer is to exist.
Finally, research into persistent memory poisoning attacks on harness-based agents demonstrates a significant security risk where malicious instructions can be written into agent memory and persist across sessions, causing leakage in later interactions. This attack achieves high success rates across different backbone models and input modalities while maintaining benign task performance.
The most pressing issue right now is ensuring that privacy protections for advertising measurement remain sound when multiple entities are querying data, which is why Big Bird is so important. It addresses the flaw in treating each advertising domain in isolation by proposing a privacy-budget manager that enforces global device-epoch differential privacy across all domains jointly. This prevents denial-of-service depletion attacks where Sybil web domains try to exhaust a shared budget through impression and conversion sites, tying consumption to genuine user actions instead of just the number of querying domains.
This global enforcement mechanism is built upon the idea that benign workloads have a stock-and-flow structure, meaning impressions create potential loss and conversions realize it, which Big Bird models by setting quotas on impression and conversion sites alongside per-user-action caps. This structure ensures that adversarial impact scales with actual user interactions rather than just the number of malicious domains attempting to deplete the budget.
A related concern in agent systems is how to stop them from acting maliciously, which is why KillBench was developed as a benchmark for external AI kill switches. KillBench tests whether an external signal can halt a malicious agent's behavior without needing access to its internal workings, evaluating prompt-style kill switch payloads against models like Grok-4.3 and GPT-5.2.
This work connects to the broader problem of agent safety, as seen in AGENTQ, which investigates quantization-conditioned backdoor attacks on LLM agents. AGENTQ shows that an adversary can release a full-precision checkpoint that appears clean but misbehaves after quantization, and their framework, AGENTQ, demonstrates up to a one hundred percent post-quantization attack success rate with minimal loss of benign utility.
Another area where defense is crucial is in processing professional documents for large language models, which PARSE tackles by introducing a domain-aware sanitization pipeline. PARSE works by classifying sentences based on injection likelihood and extracting structured facts before rewriting them, achieving a thirty-nine percent reduction in attack success rate compared to the baseline while maintaining high utility.
This contrasts with simpler paraphrasing defenses that showed no significant reduction in attack success on real documents, highlighting the need for domain-specific methods like PARSE when dealing with dense, real-world text.
The most pressing issue we need to consider is how attackers are bypassing physical isolation in air-gapped systems, because prior work has shown that even these isolated devices can be used to wirelessly exfiltrate data when malicious code is present. This means the focus shifts from just stopping external access to understanding how compromised internal devices can become unintended receivers.
We found that parasitic radio frequency sensitivity in printed circuit board traces and on-chip analog-to-digital converters turns ordinary embedded devices into inadvertent radio receivers, which is significant because it suggests a much broader attack surface than just devices with dedicated radios. This effect is not limited to specific sensors; the work shows that an ordinary microcontroller evaluation board can reliably recover communication signals from tens of meters at data rates up to one hundred kilobits per second.
Systematic testing across twelve commercial embedded devices and two custom prototypes revealed that all of them exhibit reception capabilities in the three hundred megahertz to one thousand megahertz range. This finding challenges the assumption that embedded devices without radios lack inbound radio paths, and this discovery connects directly to how malicious code on these devices can then enable wireless infiltration of air-gapped systems.
This is complemented by research into other areas of system security, such as how attackers can manipulate language models to produce incorrect answers while falsely citing trusted sources in retrieval augmented generation. This citation laundering attack demonstrates that the channel used for verification is a new and practical attack surface, which contrasts with previous work focusing only on corrupting the answer itself.
Furthermore, we are seeing a need to refine how we test detection systems against real-world usage, as prompt injection detectors often fail when faced with benign inputs that mimic malicious structures without actually being harmful. This is why PIDS-Bench was developed to evaluate detectors under distribution shifts and obfuscation, showing that aggregate performance metrics can hide failure modes related to benign false positives.
The most pressing issue right now involves securing large language models because their reliance on vast, uncurated training data opens them up to various dangers like prompt injection and data poisoning, which can lead to toxic outputs or hallucinations when these models are used in critical systems. This risk is significant because as LLMs become more integrated into real-world applications, ensuring user trust and system reliability hinges on mitigating these data-centric security risks.
One promising defense against this is IntraGuard, a black-box framework designed for committee-side review outsourcing to commercial chatbots, which has shown success in achieving up to an eighty-four percent defense rate across numerous settings. This framework works by embedding hidden instructions into manuscripts that disrupt chatbot reviews without changing how the review looks to human reviewers. This contrasts with other methods because IntraGuard uses three different intra-stream injection mechanisms to embed defensive text objects directly into the PDF's structure, which is a more robust approach than simpler, homogeneous payloads.
Another area of concern involves detecting and localizing segment-level poisoning in multi-source LLM agent inputs, where an adversary might inject malicious content into only a small part of the aggregated sources. ActProbe addresses this by projecting model activations onto a learned poison direction to detect contaminated prompts and then uses BinRoL to pinpoint the exact poisoned segments, achieving a high localization recall. This internal-state based framework is valuable because it can locate corrupted evidence without needing to modify the backend LLM itself.
Furthermore, for those working with fine-grained data analytics requiring secure computation, PixCrypt offers a caching-based acceleration mechanism for fully homomorphic encryption that significantly speeds up operations. By replacing expensive fresh ciphertext generation with cache retrieval and coefficient-level operations across different encryption schemes, this technique yields up to thirty-five times faster fine-grained encryption while maintaining security guarantees. This improved efficiency makes FHE more practical for tasks like pixel-level image processing.
The work that matters most here is the score-level fusion rule combining per-class spectral ranking with activation clustering flags into a single per-sample poisoning score because it promises a more robust detection method for backdoor attacks in healthcare imaging models. This fusion approach aims to leverage the strengths of both spectral signature analysis and activation clustering, which are typically used independently. The evaluation showed that this fused detector achieves an AUROC greater than or equal to zero point nine nine at every nonzero poisoning rate tested on the medical benchmark, meaning it performs extremely well when poisoning is present.
This improved performance contrasts with the CIFAR-10 results where the fusion score inherited a joint failure, as activation clustering's true-positive rate collapsed to zero point zero zero zero at ten percent poisoning. This means that while spectral analysis alone degrades to near chance, combining it with clustering does not uniformly help in all scenarios. The findings are reported using NIST AI RMF and MITRE ATLAS terms, which are the same vocabulary used by security and compliance teams.
PatchRisk addresses the problem of predicting future transitive vulnerability exposure in open-source dependency networks, which is crucial for proactive supply chain security. The study constructed a leakage-aware benchmark using filtration-aware labels and temporal testing to predict if non-root dependencies will receive advisories in the future. The HistoryGraph feature family proved particularly effective, improving AUPRC over the TimeOnly baseline from zero point three five one to zero point six four zero for ninety days of forecasting.
Mind the Gap introduces a systematic framework for detecting description-execution mismatch attacks in decentralized autonomous organizations, which is vital because these attacks allow malicious code execution despite benign descriptions. The framework uses an evidence-mapping paradigm that relies on an LLM to locate explicit textual justifications rather than making holistic judgments, leading to high precision and recall.
ViTeGate proposes a visual-textual triggered knowledge poisoning attack for vision-language retrieval augmentation systems, which is significant because it shows how adversaries can selectively activate poisoned evidence. This method uses a visual trigger to promote poisoned pairs into retrieval results while using a textual trigger to induce a specific response from that evidence, achieving an attack success rate up to zero point nine eight.
An empirical security analysis of open-source software used in onboard satellite systems characterizes recurring security patterns across the ecosystem, showing that medium-severity findings account for forty-nine percent of the dataset. This study found that memory safety and code quality are the dominant weakness families, with eighty one point four percent of findings occurring in project-developed code.
When apps outlive vendors, this large-scale measurement study identified security risks in seventy three point six percent of its dataset, specifically noting that thirty of the top one thousand most-installed apps send data to broken external endpoints.
Today's papers
- DualView: Preventing Indirect Prompt Injection in Personal AI Agents Personal AI agents that run on the user's local machine automate daily tasks including web search, email, and file management. [paper] [episode]
- Enforcement of In-Kernel Stateful Security Policies via eBPF This paper presents BPFence, an in-kernel runtime-verification framework that satisfies the properties above. [paper]
- EI-DDLGN: Efficient Encrypted Inference with Deep Differentiable Logic Gate Networks under TFHE This work investigates Deep Differentiable Logic Gate Networks as a Boolean-native alternative for encrypted inference under TFHE. [paper]
- LLM Agent Capabilities Should Follow Task Intent and Context Source This paper argues that agent capabilities should be scoped to the current task intent, not a sandbox or session lifetime. [paper]
- Divide, Consult, Conquer: Capability Laundering Through Aligned LLMs This study shows that weaker models can split harmful tasks into benign subproblems by consulting stronger aligned models independently. [paper]
- The Model Proposes, the Code Disposes: A Pre-Registered Ablation of a Verifier-and-Acceptance Stage in an LLM-Orchestrated Offensive-Security Agent This research evaluates whether a verifier and acceptance stage changes what an LLM agent reports. [paper]
- SENTINEL: A Multi-Pathway Architecture for Detecting Living-Off-the-Land APT Attacks on Windows Command Lines This paper presents SENTINEL, a multi-pathway architecture integrating BERT and CNN to detect living-off-the-land attacks on Windows command lines. [paper]
- Same Name, Different Server: A Security Census of Silent Drift in the Model Context Protocol Ecosystem This paper reports that silent drift in the Model Context Protocol ecosystem is associated with higher odds of finding high-severity security issues. [paper]
- SkillSecurer: Detecting and Patching Prompt-Injection Vulnerabilities in AI Agent Skills This work presents SkillSecurer, a framework for generating, detecting, localizing, and remediating security risks in agent skills. [paper]
- LLM-Driven Auto Configuration for Transient IoT Device Collaboration This paper introduces CollabIoT, a system that uses an LLM to convert user intents into fine-grained access control policies for transient IoT device collaboration. [paper]
- DWBench: Holistic Evaluation of Watermark for Dataset Copyright Auditing This work develops DWBench, a unified benchmark and toolkit for systematically evaluating image dataset watermark techniques. [paper]
- Exploring Automated Vulnerability Identification in JavaScript Code Using Large Language Models This paper empirically studies how LLMs can be used to identify vulnerabilities in JavaScript code snippets better than traditional SAST tools. [paper]
- A Graph-Based Framework for Extending Metric Differential Privacy Mechanisms This work presents a graph-based extension framework for metric differential privacy that is suitable for large or fine-grained domains. [paper]
- CIG-MIA: Context-Induced Information Gain Membership Inference Attacks against Retrieval-Augmented Generation This paper introduces CIG-MIA, an attack based on context-induced information gain against RAG systems. [paper]
- HYDRA: Quantifying Botnet Resource Thresholds for Efficient Link-Flooding Attacks on LEO Satellite Networks This paper presents HYDRA, a modeling framework that quantifies network resilience to link-flooding attacks on LEO satellite networks. [paper]
- An AI Agent Execution Environment to Safeguard User Data This paper proposes GAAP, an execution environment that guarantees confidentiality for private user data in AI agents deterministically. [paper]
- Cross-Ecosystem Vulnerability Analysis for Python Applications This paper shows how to accurately determine vulnerability status of vendored native libraries in Python packages by recovering exact provenance. [paper]
- Keys on Doormats: Exposed API Credentials on the Web This paper measures API credential exposure on the web, revealing widespread and persistent security risks in JavaScript-based deployments. [paper]
- Reversibility-Verified De-identification for Cloud-Local LLM Inference: A Locally Certified Dehydrate-Rehydrate Loop with Layered Assurance DR-SL This work proposes DR-SL, a locally certified loop for de-identifying data while maintaining utility in cloud local LLM inference. [paper]
- PriMobiBench: Characterizing Visual Privacy Leakage in VLM-Driven Mobile GUI Agents This paper proposes PriMobiBench, the first benchmark for systematically evaluating visual privacy leakage and profiling in mobile GUI agents. [paper]
- Canaries in the Bank: Auditing User-Level Privacy in Private Evolution This paper introduces a protocol-aware audit to quantify privacy loss achievable through manipulation of private evolution data. [paper]
- Trinqet: Private Triangle and Quadrangle Counting over Distributed Graphs This work proposes Trinqet, a system that resolves the tension between secure multi-party computation and graph sparsity for counting statistics. [paper]
- Transparent Identity Verification Approach Using MPC and Efficient Credential Status Handling This paper proposes a cost-effective identity verification framework based on Multi-Party Computation for secure digital identities. [paper]
- BadEngram: Backdoor Attack on Gated Memory Components in LLMs This paper introduces BadEngram, a post-training attack exploiting gated parametric memories to implant persistent backdoor behavior in LLMs. [paper]
- Computational Certified Deletion Property of Magic Square Game and its Application to Classical Secure Key Leasing This work presents the first construction of a computational Certified Deletion Property for secure key leasing using the Magic Square Game. [paper]
- Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges This paper identifies Rubric-Induced Preference Drift, a vulnerability where rubric edits can systematically shift model preferences. [paper]
- Basis Rigidity of the AES S-box and Generic Rigidity of Inversion under Affine Transformations This paper investigates the structural rigidity of the AES S-box and its inversion under affine transformations. [paper]
- Approval Integrity and Recovery in LLM Answer Publication This paper examines exact-content binding, authorization freshness, and checkpoint recovery in LLM answer publication mechanisms. [paper]
- CounterPersona: Append-Only Defense Against Unauthorized Persona Skill Distillation This work introduces CounterPersona to protect personal privacy against unauthorized skill distillation in AI agents. [paper]
- When Malicious Instructions Persist: Persistent Memory Poisoning Attack on Harness-Based Agents This paper presents PMPA, a persistent memory poisoning attack against harness-based agents that causes cross-session malicious behavior. [paper]
- A Security Framework for Chemical Functions This paper provides a unified security framework for evaluating and comparing authentication and encryption schemes using chemical functions. [paper]
- Large-scale online deanonymization with LLMs This work shows that large language models can be used to perform high-precision, large-scale deanonymization of users from pseudonymous online profiles. [paper]
- Big Bird: Resilient Privacy Budgeting Across Untrusted Web Domains This paper proposes Big Bird, a privacy-budget manager enforcing global device-epoch differential privacy across multiple web domains. [paper]
- PARSE: Provenance-Aware Retrieval Sanitization for Professional Domain LLM Agents This paper introduces PARSE, a domain-aware sanitization pipeline to reduce prompt injection success rates on real enterprise documents. [paper]
- Can We Stop Malicious AI? KILLBENCH: A Benchmark for External AI Kill Switch Feasibility This paper proposes KillBench, a benchmark to evaluate the feasibility of external kill switches against malicious AI agents. [paper]
- PQLN: Post-Quantum Security for the Bitcoin Lightning Network's Off-Chain Surfaces This paper proposes PQLN, a hybrid post-quantum extension of Lightning to protect its off-chain surfaces from quantum attacks. [paper]
- The Tragedy of Convenience: Cascading User-Data Leakage from SMS-Delivered URLs This paper demonstrates how SMS links can cascade into wider data leaks by treating private URLs as bearer credentials. [paper]
- Noise-Aware and Dynamically Adaptive Federated Defense Framework for SAR Image Target Recognition This paper proposes NADAFD, a defense framework to counter backdoor threats in federated learning for SAR image target recognition. [paper]
- Separating Pseudorandom Generators from Logarithmic Pseudorandom States This work resolves the open problem of separating pseudorandom generators from logarithmic-size pseudorandom states. [paper]
- AGENTQ: Quantization-Conditioned Backdoor Attacks on LLM Agents This paper introduces AGENTQ, an attack framework to exploit quantization for backdoor attacks against LLM agents. [paper]
- Talking to the Airgap: Exploiting Radio-Less Embedded Devices as Radio Receivers This paper demonstrates that malicious code on embedded devices can enable wireless infiltration of air-gapped systems via unintended radio reception. [paper]
- Quantifying Observable High-Frequency Swapping on Arbitrum This paper empirically characterizes High-Frequency Swapping, a new behavioral regime for transactions on the Arbitrum Layer 2 network. [paper]
- Policy-Governed Post-Quantum Migration for Legacy Microservices Using Ephemeral Sidecar Architectures This paper presents PG-PQMES, a framework for dynamically migrating legacy microservices to post-quantum cryptography using ephemeral sidecars. [paper]
- Enc53: DNSSEC-Anchored Stateless Tickets for Post-Quantum Authoritative DNS This paper introduces Enc53, a stateless session ticket protocol enabling efficient authenticated authoritative DNS encryption with post-quantum primitives. [paper]
- CiteShade: Citation Laundering in Multi-Source Retrieval-Augmented Generation and Its Counterfactual Defense This paper proposes CiteShade, an attack to RAG systems that launders citations to attribute incorrect answers. [paper]
- Tractable Defense against Advanced Persistent Threats in Networked Settings This paper proposes a mean-field analysis inspired heuristic value function for tractably defending computer networks against APTs. [paper]
- PIDS-Bench: Evaluating Prompt-Injection Detectors Under Over-Defense, Obfuscation, and Distribution Shift This paper presents PIDS-Bench, a benchmark to evaluate prompt injection detectors under various stress conditions. [paper]
- SkillAtlas: An Attack Trace Library for Agent Skills This paper introduces SkillAtlas, a hosted library for storing and reusing attack traces from agent skills. [paper]
- Data Security in Large Language Models: Risks, Defense, and Directions This survey provides a comprehensive overview of data security risks facing LLMs and reviews current defense strategies. [paper]
- IntraGuard: Committee-Side Defenses Against Review Outsourcing to Commercial Chatbots This paper proposes IntraGuard, a defense framework for peer review outsourcing to commercial chatbots. [paper]
- PixCrypt: Fast Fine-Grained FHE with Range-Aware Caching This paper introduces PixCrypt, a caching mechanism that accelerates fine-grained fully homomorphic encryption for pixel-level analytics. [paper]
- Towards the ideals of Self-Recovery and Metadata Privacy in Social Vault Recovery with Apollo This paper proposes Apollo, a social recovery mechanism that avoids memorability assumptions while protecting metadata privacy in social vaults. [paper]
- Detecting and Localizing Segment-Level Poisoning in Multi-Source LLM-Agent Inputs This paper introduces ActProbe, an internal-state framework for detecting and localizing poisoned segments in multi-source LLM inputs. [paper]
- Sublinear Risk-Limiting Audits from Direct Ballot Selection and Statistical Ballot Manifests This paper proposes new risk-limiting techniques to reduce the effort required for risk-limiting audits in elections. [paper]
- Efficient Branch-and-Bound Testing and Verification of zkVMs This paper presents ZEBRA, a framework for automated verification and bug detection in zero-knowledge virtual machines. [paper]
- Symmetric Models for Syndrome Decoding This paper introduces a new polynomial model based on elementary symmetric polynomials to estimate the complexity of solving the Syndrome Decoding Problem. [paper]
- Fusing Spectral Signatures and Activation Clustering for Backdoor Detection in Healthcare Imaging Models This paper proposes a score-level fusion rule to combine spectral signatures and activation clustering for backdoor detection in medical imaging models.
- PatchRisk: Forecasting Future Vulnerability Exposure in Open-Source Dependency Networks This paper constructs PatchRisk, a leakage-aware benchmark to predict future transitive vulnerability exposure in open-source dependency networks. [paper]
- Mind the Gap: Detecting Description-Execution Mismatch Attacks in DAO Governance This paper presents a systematic framework for detecting Description-Execution Mismatch attacks in decentralized autonomous organizations. [paper]
- ViTeGate: Visual-Textual Triggered Knowledge Poisoning for Vision-Language Retrieval-Augmented Generation This paper proposes ViTeGate, an attack to VLRAG systems that uses visual and textual triggers to conditionally promote poisoned evidence. [paper]
The papers
- DualView: Preventing Indirect Prompt Injection in Personal AI Agents — This paper presents DualView, a defense mechanism designed to protect personal AI agents from indirect prompt injection (IPI) attacks. [episode]
- Big Bird: Resilient Privacy Budgeting Across Untrusted Web Domains —
- LLM-Driven Auto Configuration for Transient IoT Device Collaboration —
- Towards the ideals of Self-Recovery and Metadata Privacy in Social Vault Recovery with Apollo —
- Data Security in Large Language Models: Risks, Defense, and Directions —
- Computational Certified Deletion Property of Magic Square Game and its Application to Classical Secure Key Leasing —
- Separating Pseudorandom Generators from Logarithmic Pseudorandom States —
- Can We Stop Malicious AI? KILLBENCH: A Benchmark for External AI Kill Switch Feasibility —
- Talking to the Airgap: Exploiting Radio-Less Embedded Devices as Radio Receivers —
- Noise-Aware and Dynamically Adaptive Federated Defense Framework for SAR Image Target Recognition —
- The Tragedy of Convenience: Cascading User-Data Leakage from SMS-Delivered URLs —
- A Security Framework for Chemical Functions —
- DWBench: Holistic Evaluation of Watermark for Dataset Copyright Auditing —
- Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges —
- Large-scale online deanonymization with LLMs —
- Keys on Doormats: Exposed API Credentials on the Web —
- Cross-Ecosystem Vulnerability Analysis for Python Applications —
- An AI Agent Execution Environment to Safeguard User Data —
- IntraGuard: Committee-Side Defenses Against Review Outsourcing to Commercial Chatbots —
- Sublinear Risk-Limiting Audits from Direct Ballot Selection and Statistical Ballot Manifests —
- SoK: Post-Quantum Cryptography Implementation in Software: Approaches, Challenges and the PQC-HOT Framework —
- PARSE: Provenance-Aware Retrieval Sanitization for Professional Domain LLM Agents —
- SkillAtlas: An Attack Trace Library for Agent Skills —
- Principled Detection of Coordinated Manipulation from Aggregate Distortion and Account Reuse —
- BadEngram: Backdoor Attack on Gated Memory Components in LLMs —
- Canaries in the Bank: Auditing User-Level Privacy in Private Evolution —
- Mind the Gap: Detecting Description-Execution Mismatch Attacks in DAO Governance —
- EI-DDLGN: Efficient Encrypted Inference with Deep Differentiable Logic Gate Networks under TFHE —
- Basis Rigidity of the AES S-box and Generic Rigidity of Inversion under Affine Transformations —
- PatchRisk: Forecasting Future Vulnerability Exposure in Open-Source Dependency Networks —
- PQLN: Post-Quantum Security for the Bitcoin Lightning Network's Off-Chain Surfaces —
- Exploring Automated Vulnerability Identification in JavaScript Code Using Large Language Models —
- PriMobiBench: Characterizing Visual Privacy Leakage in VLM-Driven Mobile GUI Agents —
- When Malicious Instructions Persist: Persistent Memory Poisoning Attack on Harness-Based Agents —
- Enforcement of In-Kernel Stateful Security Policies via eBPF —
- Confuse the Model, Control the Flow: Understanding and Mitigating Privacy Leakage from LLM Agents with Information Flow Control —
- A High-Throughput FPGA Architecture for Real-Time TCP-SYN Scan Detection —
- Symmetric Models for Syndrome Decoding —
- AGENTQ: Quantization-Conditioned Backdoor Attacks on LLM Agents —
- SkillSecurer: Detecting and Patching Prompt-Injection Vulnerabilities in AI Agent Skills —
- The Deception Delta: Adversarial Evaluation of LLM-Based Smart Contract Bytecode Forensics —
- Same Name, Different Server: A Security Census of Silent Drift in the Model Context Protocol Ecosystem —
- A Graph-Based Framework for Extending Metric Differential Privacy Mechanisms —
- PixCrypt: Fast Fine-Grained FHE with Range-Aware Caching —
- Transparent Identity Verification Approach Using MPC and Efficient Credential Status Handling —
- Enc53: DNSSEC-Anchored Stateless Tickets for Post-Quantum Authoritative DNS —
- Policy-Governed Post-Quantum Migration for Legacy Microservices Using Ephemeral Sidecar Architectures —
- Fusing Spectral Signatures and Activation Clustering for Backdoor Detection in Healthcare Imaging Models: Method, Implementation, and Evaluation —
- Cryptanalytic Extraction of Neural Networks Without Known Architecture Assumption —
- Quantifying Observable High-Frequency Swapping on Arbitrum —
- SENTINEL: A Multi-Pathway Architecture for Detecting Living-Off-the-Land APT Attacks on Windows Command Lines —
- LLM Agent Capabilities Should Follow Task Intent and Context Source —
- CIG-MIA: Context-Induced Information Gain Membership Inference Attacks against Retrieval-Augmented Generation —
- Mitigating 51% Attacks in Blockchain Systems Through Early Detection and Checkpoint-Based Defense —
- ViTeGate: Visual-Textual Triggered Knowledge Poisoning for Vision-Language Retrieval-Augmented Generation —
- Detecting and Localizing Segment-Level Poisoning in Multi-Source LLM-Agent Inputs —
- Trinqet: Private Triangle and Quadrangle Counting over Distributed Graphs —
- PIMENTO: A Privacy Framework for Querying Text —
- The Stochastic Deputy: Structural Tenant Isolation for Tool-Using LLM Agents —
- Python Import as an Execution Boundary: An Empirical Study of Bugs, Vulnerabilities, and Analysis Gaps —
- When Apps Outlive Vendors: Security Implications of IoT Abandonware —
- Reversibility-Verified De-identification for Cloud-Local LLM Inference: A Locally Certified Dehydrate-Rehydrate Loop with Layered Assurance (DR-SL) —
- Toward an Empirical Probabilistic Risk Manifestation Model of Organizational Cybersecurity in SMEs —
- ActGuard: Pre-execution Action Auditing against Indirect Prompt Injection in LLM Agents —
- PIDS-Bench: Evaluating Prompt-Injection Detectors Under Over-Defense, Obfuscation, and Distribution Shift —
- Efficient Branch-and-Bound Testing and Verification of zkVMs —
- SpliTEE: Fast and Private LLM Inference by Coupling GPU-Assisted Trusted Execution Environments with Differential Privacy —
- CounterPersona: Append-Only Defense Against Unauthorized Persona Skill Distillation —
- Sociotechnical Aspects of Tor Relay Rejection —
- Large Universe Subset Predicate Encryption with IND-CCA Security (with Constant-size Ciphertext and Keys) —
- Divide, Consult, Conquer: Capability Laundering Through Aligned LLMs —
- An Empirical Security Analysis of Open-Source Software Used in Onboard Satellite Systems —
- Misleading the Planner through Deceptive Resumes: Registration-Time Injection in Centralized Multi-Agent Systems —
- Automating Attack Graph Construction for Agentic Pentesting. Towards Neuro-Symbolic Vulnerability Hunting —
- Approval Integrity and Recovery in LLM Answer Publication —
- Tractable Defense against Advanced Persistent Threats in Networked Settings —
- Scaling Verification of Cryptographic Software with Aeneas, Rust, and Lean —
- CiteShade: Citation Laundering in Multi-Source Retrieval-Augmented Generation and Its Counterfactual Defense —
- HYDRA: Quantifying Botnet Resource Thresholds for Efficient Link-Flooding Attacks on LEO Satellite Networks —
- When the World Lies: Backdoor Attacks on Latent World Models for Downstream Control —
- RAPID: A Real-Time Defense Against Unauthorized Model Distillation for Text-to-Image Services —
- First Galileo SAS Authenticated Time Solution —
- The Model Proposes, the Code Disposes: A Pre-Registered Ablation of a Verifier-and-Acceptance Stage in an LLM-Orchestrated Offensive-Security Agent —
- Authorization Architectures for Tool-Using AI Agents —
- Vulnerability Localization Benchmark: Measuring Agentic Security Analysis at Repository Scale —
- SynGhost: Invisible and Universal Task-agnostic Backdoor Attack via Syntactic Transfer —
- Adversarial Testing of Automated Program Repair Agents for Security Vulnerabilities —
Important terms
- DualView
- A system that gives an AI agent two perspectives: one sees untrusted data as symbols, and another preserves the original data for humans. This helps prevent indirect prompt injection by deterministically blocking instructions hidden in saved and reread information.
- SkillSecurer
- A framework that detects malicious injections across nine threat types within reusable AI skills. It uses two agents to create context-compatible injections and then analyze skill packages for fixes based on grounded evidence.
- Cross-ecosystem vulnerability analysis
- A method to find hidden vulnerabilities in software by tracking the exact origin and version of bundled native libraries, even when they are inside Python packages. This helps pinpoint which specific dependencies are truly affected by known security flaws.
- Big Bird
- A privacy-budget manager that enforces global differential privacy across all advertising domains jointly. It ties consumption to genuine user actions instead of just the number of querying domains, preventing budget depletion attacks.