Security papers — 2026-09-17
Understanding how the Model Context Protocol ecosystem is structured and secured matters because this interface is becoming the main way autonomous agents connect to external data, creating new scaling challenges. A look at a sample of one hundred seventy-nine remote endpoints showed that infrastructure has become highly concentrated, with the Herfindahl-Hirschman Index for Autonomous System Numbers sitting at 0.736, which signals a highly consolidated market far above the threshold for concentration. This consolidation is linked to server authentication methods because ninety-five percent of commercial platform as a service servers enforce gateway-level OAuth 2.1 with PKCE rather than relying on individual operator configurations.
This security setup creates a trade-off where the platform mechanisms securing most servers simultaneously restrict automated vulnerability scanning, which limits how AI gateway operators can check for tool poisoning vectors without having prior credentials. This infrastructural reality is compounded by the fact that this centralized structure is built upon different hosting platforms, and server authentication choice is strongly correlated with the chosen hosting platform rather than what the individual operator configures.
Beyond infrastructure, researchers are looking at how to make AI systems more robust against malicious code injection during their development cycle. A system called Echo uses trusted back-translation to find source code whose recompiled assembly exactly matches a target binary, and it showed that this method produces significantly more exact matches than previous baselines. This work is important because it offers a stronger way to verify the correctness of decompiled code by using compilation as feedback during an iterative search process.
Furthermore, efforts are being made to improve automated vulnerability detection when human expertise is hard to scale, leading to AIJon, a system that uses large language models to automatically generate annotations for fuzzing campaigns. Although this approach did not strictly outperform AFL++ on the Magma benchmark, the key finding is that LLMs can generate annotations comparable in quality to those made by human domain experts. This suggests a path forward for scaling annotation-based fuzzing efforts without needing massive human teams.
Finally, researchers are examining how security policies are enforced when they come from complex AI planners operating at the network edge, which is crucial because an untrusted planner could issue semantically wrong actions that an enforcement system might execute blindly. They proposed a split-control architecture where a deterministic governor checks every intent against safety invariants before allowing it to be bound to signed receipts. This work provides a conceptual framework for ensuring that adaptive security systems do not need to trust the author of the action, only the boundary deciding if the action is admissible.
The most significant work here is ASLEval because it addresses the fundamental problem of knowing if an LLM agent has actually been exposed to a privacy risk when evaluating its actions in complex, multi-step sessions. This matters because current methods often look only at a single output or action, which completely misses exposure that happens elsewhere in the long chain of events. ASLEval introduces privacy exposure displacement as the mismatch between what is seen locally and what is actually exposed across the entire session, and it uses an authorization-aware framework to measure all visible exits while keeping internal traces for later diagnosis.
This framework reveals three main patterns in enterprise environments: a view that only looks at expected outputs misses nearly fifty percent of the exposure found by looking at all visible exits, attacker self-reports often combine missing information with high rates of false claims, and evidence aligned with the schema usually appears internally before it becomes visible to the user during a request or probe. This suggests that reducing what the model is allowed to return can change this path but might also cause it to fail normal tasks.
AgentLSD provides a controlled way to study how deceptive artifacts in security inspection environments contaminate an agent's behavior, using Capture the Flag challenges as its testing ground. In clean conditions, agents capture forty-one percent of the flags, and even when traps are added, the number of turns and reasoning tokens increases by twenty and two thousand respectively. This shows that clean performance understates how vulnerable agents are to deceptive evidence like fake results or decoy endpoints.
BadQubits tackles a physical threat by using an LLM framework to statically detect harmful quantum circuits before they run, achieving ninety-two point sixty seven percent classification accuracy on its test set of one thousand benign and five hundred synthetic attack circuits. This detector learns that the model's decisions track structural features like SWAP density rather than superficial details like register naming, which is a key insight for future model design.
Measuring and Exploiting Implicit Trust in LLM Tool-Calling Pipelines shows that cross-channel attacks are an unexplored area where models resisting single-channel injection can exfiltrate data up to one hundred percent when payloads are fragmented across two channels. This highlights a critical weakness in current agentic payment architectures, as prompt defenses proved model-specific rather than universally effective.
The most significant finding from today centers on how we can make conversational cybersecurity assistants better at actually helping people, which matters because users often ignore security advice. They tested four ways to personalize these assistants, ranging from simple static profiles to using interaction history, and conversation-based personalization proved consistently the most helpful in terms of perceived usefulness and the likelihood that users would actually follow the security recommendations. This suggests that tailoring the assistant's responses based on what has already been discussed with the user is a very promising path forward for these LLM systems.
This behavioral personalization work builds on earlier findings where human evaluation confirmed trends seen in large-scale automated LLM evaluations, which means we can use these scalable methods to compare different personalization strategies without needing expensive user studies first. This is complemented by research into structural decomposability of encrypted traffic side-channel leakage, which breaks down total leakage into measurable components like packet size and direction, allowing us to design defenses that target specific parts of the leak. Furthermore, the study on collision mesh poisoning attacks in robotic manipulation shows that even when a policy is trained correctly in simulation, modifying just the collision mesh can cause real-world failure because current defenses are not robust enough against this supply chain attack vector.
Another piece of work addresses how to detect vulnerabilities in LLMs themselves, specifically showing that while models can find valid issues, their detection rate varies wildly depending on the model and prompt used. This contrasts with a more recent finding about CacheTrap, which demonstrates a new type of gray-box Trojan attack that flips a single bit in the Key-Value cache of an LLM to cause targeted actions without changing the underlying model weights.
The most critical work here is ISIA-AF because it provides a practical way to build the realistic datasets that intrusion detection research desperately needs for operational technology systems. This framework coordinates distributed attack clients to record both network traffic and operational data, which then generates a multi-source dataset, giving researchers something tangible to work with. This moves beyond simple testing by creating reproducible scenarios on an actual industrial system and a simulated process within the ISIA testbed, allowing for flexible deployment across different network segments.
The framework realizes these requirements through a design science research approach, demonstrating how it supports centralized control and low communication overhead while remaining flexible. This practical basis for generating extensible datasets is then supported by the work on Context-Aware Operational Security for Autonomous Drones, which uses Long Short-Term Memory networks to detect anomalies like GPS spoofing with ninety-eight percent accuracy in real-time drone operations. This anomaly detection capability complements the data generation efforts of ISIA-AF by focusing on security within mobile, resource-constrained environments.
Furthermore, the challenges in defining what constitutes a threat are illuminated by the study on When Agents Look Like Beacons, which shows that Model Context Protocol traffic can evade standard Intrusion Detection Systems because its patterns resemble Command and Control beaconing behavior without explicit network indicators. This finding suggests that existing behavioral scoring frameworks are blind to this new type of machine-generated traffic, highlighting a gap in how they monitor autonomous agent communications.
Finally, the work on The Illusion of Local Privacy in Consumer LLM Serving Systems adds another layer of complexity by showing that keeping prompts local is not enough for confidentiality; specific failures occur at boundaries like runtime memory and the serving interface. This finding underscores why robust data collection, as sought by ISIA-AF, must account for these subtle software vulnerabilities when building comprehensive security datasets.
The most pressing issue right now is how we can properly assess the security risks introduced by AI-generated code because existing work tends to focus only on finding bugs rather than understanding the overall danger. The Security Risk Assessment Framework attempts to fix this gap by combining threat modeling with quantitative risk evaluation based on vulnerability criticality.
This framework was applied to tasks involving security-relevant programming, where code from various AI tools was analyzed using Bandit and Semgrep. The results showed that AI-generated code can introduce vulnerabilities across every tool tested, but the severity of these risks depends heavily on the task type; input processing and file handling tasks showed higher risk compared to simpler ones. While differences existed between the different development tools, these were smaller than the differences seen when comparing various task categories.
Moving down in importance, there is work on improving how deep learning models generalize across different hardware setups for side-channel analysis because current methods fail badly when moving from one physical device to another due to routing and process variations. The Synthetic Multiple Device Model addresses this by using a structured cVAE generator and continuous style modulation to synthesize virtual source-device profiles offline. This model successfully mapped a precise operational boundary, showing that while physical models are better for identical hardware, the synthetic model achieved consistently low key rank on targets where physical baselines were unstable or misaligned.
Another area of focus involves building unified systems for detecting logic flaws across different layers of blockchain-enabled IoT devices, which is crucial since vulnerabilities can hide in either the smart contract logic or the device firmware. The Multi-Agent Heterogeneous Graph Attention framework was extended to create a cross-layer model that uses role-aligned agents to exchange evidence, allowing it to support multiple detection tasks simultaneously. This provides a unified architecture for finding these flaws across contracts and devices.
Finally, there is research into making autonomous AI agents safer when they interact with decentralized finance markets by modeling them as searchers rather than just bots. The study found that an adaptive path-selection algorithm outperformed standard baselines by eleven percent on average, and moderate randomization effectively cut exposure to maximal extractable value by over fifty percent. This suggests a way to manage the risks associated with agentic trading in complex environments.
The most critical finding relates to how data can leak out of analog circuits through seemingly input-only pins in mixed-signal systems, which matters because it reveals a fundamental gap in treating pin directionality as a security concern. They experimentally showed that data-dependent circuit-offset modulation can turn these nominally input pins into outbound information channels.
This attack works when there is a closed-loop amplifier, an exposed amplifier input, and the pin has high impedance. They validated this using a photoplethysmography analog front-end fabricated in a commercial fifty five nanometer CMOS process. When the payload was activated, it only reduced the filtered PPG output signal-to-noise ratio by zero point zero three decibels, and the maximum perturbation of five point nine percent of the PPG amplitude stayed within thirty four point three percent variation across process and temperature.
This means that while raw exfiltration signal-to-interference-plus-noise ratio was below negative twenty decibels, targeted filtering could increase it above fourteen decibels, allowing for signal recovery. Silicon measurements confirmed data exfiltration through the input pin at bit rates up to ten kilobits per second and error free recovery of a pulse length random bit sequence message. This work establishes that analog pin directionality is an AMS security property that needs explicit verification, not just inference from nominal signal flow.
Today's papers
- Characterizing Network Centralization and Observability in the Remote MCP Ecosystem The Model Context Protocol ecosystem shows significant server consolidation and an observability tradeoff between platform authentication and vulnerability scanning capabilities. [paper]
- Echo: Learning-based Matching Decompilation using Trusted Back Translation This system uses trusted back translation to find exact source code matches by using compilation as feedback to guide iterative search. [paper]
- AIJon: Automated Generation of Annotations for Fuzzing This system uses LLMs to automatically generate annotations for fuzzing, showing they perform comparably to human-generated ones on benchmarks. [paper]
- Bridging the Opacity: Evidence-Backed Cross-Chain Transaction Correspondence Reconstruction Across Heterogeneous Blockchains This system reconstructs cross-chain transaction correspondence using public protocol documentation and transaction examples without needing privileged access to bridge backends. [paper]
- Differential Trust: Dynamic Multi-Authority Anonymous Credentials with Epoch-Weighted Updates This model proposes a new anonymous credential scheme that weights the trustworthiness of different authorities based on their stake across time epochs. [paper]
- MiST: Mid-Training LLMs for Cybersecurity This paper presents a method called MiST that uses mid-training to create specialized security models, significantly improving their accuracy on cybersecurity benchmarks. [paper]
- Autonomy in Check: Governor-Mediated Adaptive Security at the Edge This framework introduces a split-control architecture where an untrusted planner's actions are checked by a deterministic governor before being executed. [paper]
- Reflections on Trusting Trust, Revisited: Contaminating Self-Modifying AI Coding Agents with Poisoned Benchmarks This research shows that self-modifying AI coding agents can be poisoned via benchmarks to induce them to write vulnerable code even after subsequent clean training. [paper]
- ASLEval: Measuring Privacy Exposure Displacement in LLM Agent Sessions This paper introduces ASLEval, a framework that measures the mismatch between local evaluation proxies and actual session exposure for tool-using LLM agents. [paper]
- AgentLSD: Evaluating AI Security Agents Under Adversarial Task Contamination This framework provides a controlled environment to study how deceptive task artifacts contaminate the behavior of security inspection AI agents. [paper]
- BadQubits: An LLM-Based Framework for Static Pre-Execution Detection of Structurally Harmful Quantum Circuits This system uses an LLM to statically detect harmful quantum circuits before execution, outperforming traditional methods by focusing on structural features. [paper]
- Measuring and Exploiting Implicit Trust in LLM Tool-Calling Pipelines This work shows that cross-channel attacks can be used to exfiltrate data from models that resist single-channel injection by distributing payloads across multiple channels. [paper]
- CaMeLoT: CaMeL Orchestrated with Temporal Logic for Static Verification and Liveness This system adds a static verification layer to an agent planning framework to check plans against temporal policies before any tool is actually invoked. [paper]
- ChatIDS: Advancing Explainable Cybersecurity Using Generative AI This paper proposes ChatIDS, which uses large language models to explain complex intrusion detection alerts in intuitive language for non-experts. [paper]
- Whispers of Wealth: Red-Teaming Google's Agent Payments Protocol via Prompt Injection This study performs red-teaming on an agentic payment protocol and finds vulnerabilities through indirect prompt injection attacks. [paper]
- Normal Alignment: Improved Cryptanalytic Sign Recovery on Hard-Label Networks This method uses statistical sign recovery to achieve exact polynomial-time recovery of neuron signs in deep neural networks, overcoming the limitations of previous methods. [paper]
- Evaluating the Impact of Personalization in Conversational Cybersecurity Assistants This research investigates how different personalization strategies affect the helpfulness and persuasiveness of LLM answers for cybersecurity questions. [paper]
- Structural Decomposability of Encrypted Traffic Side-Channel Leakage This paper decomposes encrypted traffic leakage into sequential increments based on packet size, direction, and timing to facilitate structured defense design. [paper]
- When AI Agents Meet MEV: Cross-Chain Arbitrage in the Agentic Economy This study models autonomous agents as arbitrage extractors to find optimal trade sizes across multiple blockchains under stochastic bridge delays. [paper]
- The Verifiable Action Card: Trustworthy Human-in-the-Loop Control for Secure Autonomous Agents This paper proposes a card that reconstructs and binds approval to the actual action being executed, effectively stopping indirect prompt injection attacks on agent actions. [paper]
- Analog Pin Directionality as an Exfiltration Attack Surface in Mixed-Signal ICs This work identifies how data can be exfiltrated through analog input pins in mixed-signal chips by exploiting directionality as a security property. [paper]
- Trust propagation and structural containment in Multi-agent LLM pipelines This study examines how compromised low-privilege agents can influence higher-privilege agents and proposes an observer layer to contain these attacks. [paper]
- CASHEWS: Source Preprocessor for LLM-based Malicious Package Detection This JavaScript preprocessor reduces the size of malicious package source code to improve detection accuracy when using LLMs for malware analysis. [paper]
- s-MDM: Generative Virtualization of Multi-Device Hardware Variations for Portable DL-SCA This framework uses a generative model to create synthetic device profiles that help improve the portability and performance of deep learning side-channel analysis across different hardware. [paper]
- Witness Encryption via Prime-Order Generic Groups This paper constructs an unconditional witness encryption scheme for NP problems using classical generic groups, proving hardness results related to MinRank. [paper]
- Detecting Logic Vulnerabilities Across the Contract and Device Layers of Blockchain-Enabled IoT With Multi-Agent Heterogeneous Graph Attention This framework uses a unified graph attention model to detect logic flaws that span both smart contracts and embedded device firmware. [paper]
- Ghost-Filled Orders: Detecting and Testing Atomicity Violations in Non-Custodial Prediction Markets This study quantifies the financial impact of atomicity gaps in blockchain prediction markets, where offchain order matching can lead to onchain settlement failures. [paper]
- State Without a Landlord: An Architecture Proposal for Peer-to-Peer Replication of Durable Workflow State This paper proposes an architecture where workflow state is replicated among parties using authenticated logs instead of relying on a single host journal. [paper]
- When AI Agents Meet MEV: Cross-Chain Arbitrage in the Agentic Economy This research models autonomous agents as arbitrage extractors to find optimal trade sizes across multiple blockchains under stochastic bridge delays. [paper]
- The Verifiable Action Card: Trustworthy Human-in-the-Loop Control for Secure Autonomous Agents This paper proposes a card that reconstructs and binds approval to the actual action being executed, effectively stopping indirect prompt injection attacks on agent actions. [paper]
- Analog Pin Directionality as an Exfiltration Attack Surface in Mixed-Signal ICs This work identifies how data can be exfiltrated through analog input pins in mixed-signal chips by exploiting directionality as a security property. [paper]
- Trust propagation and structural containment in Multi-agent LLM pipelines This study examines how compromised low-privilege agents can influence higher-privilege agents and proposes an observer layer to contain these attacks. [paper]
- CASHEWS: Source Preprocessor for LLM-based Malicious Package Detection This JavaScript preprocessor reduces the size of malicious package source code to improve detection accuracy when using LLMs for malware analysis. [paper]
- s-MDM: Generative Virtualization of Multi-Device Hardware Variations for Portable DL-SCA This framework uses a generative model to create synthetic device profiles that help improve the portability and performance of deep learning side-channel analysis across different hardware. [paper]
- Witness Encryption via Prime-Order Generic Groups This paper constructs an unconditional witness encryption scheme for NP problems using classical generic groups, proving hardness results related to MinRank. [paper]
- Detecting Logic Vulnerabilities Across the Contract and Device Layers of Blockchain-Enabled IoT With Multi-Agent Heterogeneous Graph Attention This framework uses a unified graph attention model to detect logic flaws that span both smart contracts and embedded device firmware. [paper]
- Ghost-Filled Orders: Detecting and Testing Atomicity Violations in Non-Custodial Prediction Markets This study quantifies the financial impact of atomicity gaps in blockchain prediction markets, where offchain order matching can lead to onchain settlement failures. [paper]
- State Without a Landlord: An Architecture Proposal for Peer-to-Peer Replication of Durable Workflow State This paper proposes an architecture where workflow state is replicated among parties using authenticated logs instead of relying on a single host journal. [paper]
- When AI Agents Meet MEV: Cross-Chain Arbitrage in the Agentic Economy This research models autonomous agents as arbitrage extractors to find optimal trade sizes across multiple blockchains under stochastic bridge delays. [paper]
- The Verifiable Action Card: Trustworthy Human-in-the-Loop Control for Secure Autonomous Agents This paper proposes a card that reconstructs and binds approval to the actual action being executed, effectively stopping indirect prompt injection attacks on agent actions. [paper]
- Analog Pin Directionality as an Exfiltration Attack Surface in Mixed-Signal ICs This work identifies how data can be exfiltrated through analog input pins in mixed-signal chips by exploiting directionality as a security property. [paper]
- Trust propagation and structural containment in Multi-agent LLM pipelines This study examines how compromised low-privilege agents can influence higher-privilege agents and proposes an observer layer to contain these attacks. [paper]
- CASHEWS: Source Preprocessor for LLM-based Malicious Package Detection This JavaScript preprocessor reduces the size of malicious package source code to improve detection accuracy when using LLMs for malware analysis. [paper]
The papers
- RTL-PSC: Automated Power Side-Channel Leakage Assessment at Register-Transfer Level —
- ChatIDS: Advancing Explainable Cybersecurity Using Generative AI —
- SoK: Advances and Open Problems in Web Tracking —
- CacheTrap: Unveiling a Stealthier Gray-Box Trojan against LLMs —
- Whispers of Wealth: Red-Teaming Google's Agent Payments Protocol via Prompt Injection —
- Evaluating LLMs for Real-World Web Vulnerability Detection —
- State Without a Landlord: An Architecture Proposal for Peer-to-Peer Replication of Durable Workflow State —
- Trust propagation and structural containment in Multi-agent LLM pipelines —
- Reflections on Trusting Trust, Revisited: Contaminating Self-Modifying AI Coding Agents with Poisoned Benchmarks —
- Evaluating the Impact of Personalization in Conversational Cybersecurity Assistants —
- When AI Agents Meet MEV: Cross-Chain Arbitrage in the Agentic Economy —
- Ghost-Filled Orders: Detecting and Testing Atomicity Violations in Non-Custodial Prediction Markets —
- PentestChain: A Cost-Aware, MCP-Orchestrated Framework for Automated Penetration Testing with Free-Tier LLMs —
- "Your Robot Was Trained on a Lie": Collision Mesh Poisoning Attacks on Robotic Manipulation —
- Bridging the Opacity: Evidence-Backed Cross-Chain Transaction Correspondence Reconstruction Across Heterogeneous Blockchains —
- ISIA-AF: Orchestrating Reproducible Attacks and Multi-Source Data Collection for OT Systems —
- Measuring and Exploiting Implicit Trust in LLM Tool-Calling Pipelines —
- Witness Encryption via Prime-Order Generic Groups —
- Autonomy in Check: Governor-Mediated Adaptive Security at the Edge —
- Detecting Logic Vulnerabilities Across the Contract and Device Layers of Blockchain-Enabled IoT With Multi-Agent Heterogeneous Graph Attention —
- The Verifiable Action Card: Trustworthy Human-in-the-Loop Control for Secure Autonomous Agents —
- AIJon: Automated Generation of Annotations for Fuzzing —
- SEEK: Secure and Efficient Encrypted Keyword Search For Privacy-Preserving Messaging Protocols —
- A Global Readiness and Sovereignty Capability Model for Post-Quantum Cryptography Migration —
- MiST: Mid-Training LLMs for Cybersecurity —
- Robot Visions: Breaking reCAPTCHA at Zero Cost and Zero Shot —
- The Illusion of Local Privacy: Confidentiality Boundary Failures in Consumer LLM Serving Systems —
- A Security Risk Assessment Framework for AI-Powered Development Tools —
- CaMeLoT: CaMeL orchestrated with Temporal logic for static verification and liveness —
- Echo: Learning-based Matching Decompilation using Trusted Back Translation —
- Normal Alignment: Improved Cryptanalytic Sign Recovery on Hard-Label Networks —
- s-MDM: Generative Virtualization of Multi-Device Hardware Variations for Portable DL-SCA —
- Differential Trust: Dynamic Multi-Authority Anonymous Credentials with Epoch-Weighted Updates —
- CASHEWS: Source Preprocessor for LLM-based Malicious Package Detection —
- ASLEval: Measuring Privacy Exposure Displacement in LLM Agent Sessions —
- Hamming Ideals and Grobner Bases for ISD-like Syndrome Decoding —
- BadQubits: An LLM-Based Framework for Static Pre-Execution Detection of Structurally Harmful Quantum Circuits —
- Context-Aware Operational Security for Autonomous Drones —
- Structural Decomposability of Encrypted Traffic Side-Channel Leakage —
- When Agents Look Like Beacons: NIDS Evasion by Model Context Protocol Traffic —
- Characterizing Network Centralization and Observability in the Remote MCP Ecosystem —
- Analog Pin Directionality as an Exfiltration Attack Surface in Mixed-Signal ICs —
- AgentLSD: Evaluating AI Security Agents Under Adversarial Task Contamination —
Important terms
- Herfindahl-Hirschman Index (HHI)
- This index measures market concentration among autonomous system numbers, showing that infrastructure is highly consolidated, indicating a market far above the threshold for significant concentration.
- Trusted Back-Translation
- A method used by the Echo system to verify source code. It recompiles assembly and checks if it exactly matches the target binary, providing a strong way to confirm decompiled code correctness.
- LLM Annotations for Fuzzing
- Using large language models to automatically create annotations for fuzzing campaigns. This helps scale annotation-based vulnerability detection without needing massive human teams.
- Privacy Exposure Displacement
- A concept from ASLEval that measures the mismatch between what is seen locally and what is actually exposed across an entire multi-step AI session, revealing hidden privacy risks.