Security papers — 2026-09-17

Understanding how the Model Context Protocol ecosystem is structured and secured matters because this interface is becoming the main way autonomous agents connect to external data, creating new scaling challenges. A look at a sample of one hundred seventy-nine remote endpoints showed that infrastructure has become highly concentrated, with the Herfindahl-Hirschman Index for Autonomous System Numbers sitting at 0.736, which signals a highly consolidated market far above the threshold for concentration. This consolidation is linked to server authentication methods because ninety-five percent of commercial platform as a service servers enforce gateway-level OAuth 2.1 with PKCE rather than relying on individual operator configurations.

This security setup creates a trade-off where the platform mechanisms securing most servers simultaneously restrict automated vulnerability scanning, which limits how AI gateway operators can check for tool poisoning vectors without having prior credentials. This infrastructural reality is compounded by the fact that this centralized structure is built upon different hosting platforms, and server authentication choice is strongly correlated with the chosen hosting platform rather than what the individual operator configures.

Beyond infrastructure, researchers are looking at how to make AI systems more robust against malicious code injection during their development cycle. A system called Echo uses trusted back-translation to find source code whose recompiled assembly exactly matches a target binary, and it showed that this method produces significantly more exact matches than previous baselines. This work is important because it offers a stronger way to verify the correctness of decompiled code by using compilation as feedback during an iterative search process.

Furthermore, efforts are being made to improve automated vulnerability detection when human expertise is hard to scale, leading to AIJon, a system that uses large language models to automatically generate annotations for fuzzing campaigns. Although this approach did not strictly outperform AFL++ on the Magma benchmark, the key finding is that LLMs can generate annotations comparable in quality to those made by human domain experts. This suggests a path forward for scaling annotation-based fuzzing efforts without needing massive human teams.

Finally, researchers are examining how security policies are enforced when they come from complex AI planners operating at the network edge, which is crucial because an untrusted planner could issue semantically wrong actions that an enforcement system might execute blindly. They proposed a split-control architecture where a deterministic governor checks every intent against safety invariants before allowing it to be bound to signed receipts. This work provides a conceptual framework for ensuring that adaptive security systems do not need to trust the author of the action, only the boundary deciding if the action is admissible.

The most significant work here is ASLEval because it addresses the fundamental problem of knowing if an LLM agent has actually been exposed to a privacy risk when evaluating its actions in complex, multi-step sessions. This matters because current methods often look only at a single output or action, which completely misses exposure that happens elsewhere in the long chain of events. ASLEval introduces privacy exposure displacement as the mismatch between what is seen locally and what is actually exposed across the entire session, and it uses an authorization-aware framework to measure all visible exits while keeping internal traces for later diagnosis.

This framework reveals three main patterns in enterprise environments: a view that only looks at expected outputs misses nearly fifty percent of the exposure found by looking at all visible exits, attacker self-reports often combine missing information with high rates of false claims, and evidence aligned with the schema usually appears internally before it becomes visible to the user during a request or probe. This suggests that reducing what the model is allowed to return can change this path but might also cause it to fail normal tasks.

AgentLSD provides a controlled way to study how deceptive artifacts in security inspection environments contaminate an agent's behavior, using Capture the Flag challenges as its testing ground. In clean conditions, agents capture forty-one percent of the flags, and even when traps are added, the number of turns and reasoning tokens increases by twenty and two thousand respectively. This shows that clean performance understates how vulnerable agents are to deceptive evidence like fake results or decoy endpoints.

BadQubits tackles a physical threat by using an LLM framework to statically detect harmful quantum circuits before they run, achieving ninety-two point sixty seven percent classification accuracy on its test set of one thousand benign and five hundred synthetic attack circuits. This detector learns that the model's decisions track structural features like SWAP density rather than superficial details like register naming, which is a key insight for future model design.

Measuring and Exploiting Implicit Trust in LLM Tool-Calling Pipelines shows that cross-channel attacks are an unexplored area where models resisting single-channel injection can exfiltrate data up to one hundred percent when payloads are fragmented across two channels. This highlights a critical weakness in current agentic payment architectures, as prompt defenses proved model-specific rather than universally effective.

The most significant finding from today centers on how we can make conversational cybersecurity assistants better at actually helping people, which matters because users often ignore security advice. They tested four ways to personalize these assistants, ranging from simple static profiles to using interaction history, and conversation-based personalization proved consistently the most helpful in terms of perceived usefulness and the likelihood that users would actually follow the security recommendations. This suggests that tailoring the assistant's responses based on what has already been discussed with the user is a very promising path forward for these LLM systems.

This behavioral personalization work builds on earlier findings where human evaluation confirmed trends seen in large-scale automated LLM evaluations, which means we can use these scalable methods to compare different personalization strategies without needing expensive user studies first. This is complemented by research into structural decomposability of encrypted traffic side-channel leakage, which breaks down total leakage into measurable components like packet size and direction, allowing us to design defenses that target specific parts of the leak. Furthermore, the study on collision mesh poisoning attacks in robotic manipulation shows that even when a policy is trained correctly in simulation, modifying just the collision mesh can cause real-world failure because current defenses are not robust enough against this supply chain attack vector.

Another piece of work addresses how to detect vulnerabilities in LLMs themselves, specifically showing that while models can find valid issues, their detection rate varies wildly depending on the model and prompt used. This contrasts with a more recent finding about CacheTrap, which demonstrates a new type of gray-box Trojan attack that flips a single bit in the Key-Value cache of an LLM to cause targeted actions without changing the underlying model weights.

The most critical work here is ISIA-AF because it provides a practical way to build the realistic datasets that intrusion detection research desperately needs for operational technology systems. This framework coordinates distributed attack clients to record both network traffic and operational data, which then generates a multi-source dataset, giving researchers something tangible to work with. This moves beyond simple testing by creating reproducible scenarios on an actual industrial system and a simulated process within the ISIA testbed, allowing for flexible deployment across different network segments.

The framework realizes these requirements through a design science research approach, demonstrating how it supports centralized control and low communication overhead while remaining flexible. This practical basis for generating extensible datasets is then supported by the work on Context-Aware Operational Security for Autonomous Drones, which uses Long Short-Term Memory networks to detect anomalies like GPS spoofing with ninety-eight percent accuracy in real-time drone operations. This anomaly detection capability complements the data generation efforts of ISIA-AF by focusing on security within mobile, resource-constrained environments.

Furthermore, the challenges in defining what constitutes a threat are illuminated by the study on When Agents Look Like Beacons, which shows that Model Context Protocol traffic can evade standard Intrusion Detection Systems because its patterns resemble Command and Control beaconing behavior without explicit network indicators. This finding suggests that existing behavioral scoring frameworks are blind to this new type of machine-generated traffic, highlighting a gap in how they monitor autonomous agent communications.

Finally, the work on The Illusion of Local Privacy in Consumer LLM Serving Systems adds another layer of complexity by showing that keeping prompts local is not enough for confidentiality; specific failures occur at boundaries like runtime memory and the serving interface. This finding underscores why robust data collection, as sought by ISIA-AF, must account for these subtle software vulnerabilities when building comprehensive security datasets.

The most pressing issue right now is how we can properly assess the security risks introduced by AI-generated code because existing work tends to focus only on finding bugs rather than understanding the overall danger. The Security Risk Assessment Framework attempts to fix this gap by combining threat modeling with quantitative risk evaluation based on vulnerability criticality.

This framework was applied to tasks involving security-relevant programming, where code from various AI tools was analyzed using Bandit and Semgrep. The results showed that AI-generated code can introduce vulnerabilities across every tool tested, but the severity of these risks depends heavily on the task type; input processing and file handling tasks showed higher risk compared to simpler ones. While differences existed between the different development tools, these were smaller than the differences seen when comparing various task categories.

Moving down in importance, there is work on improving how deep learning models generalize across different hardware setups for side-channel analysis because current methods fail badly when moving from one physical device to another due to routing and process variations. The Synthetic Multiple Device Model addresses this by using a structured cVAE generator and continuous style modulation to synthesize virtual source-device profiles offline. This model successfully mapped a precise operational boundary, showing that while physical models are better for identical hardware, the synthetic model achieved consistently low key rank on targets where physical baselines were unstable or misaligned.

Another area of focus involves building unified systems for detecting logic flaws across different layers of blockchain-enabled IoT devices, which is crucial since vulnerabilities can hide in either the smart contract logic or the device firmware. The Multi-Agent Heterogeneous Graph Attention framework was extended to create a cross-layer model that uses role-aligned agents to exchange evidence, allowing it to support multiple detection tasks simultaneously. This provides a unified architecture for finding these flaws across contracts and devices.

Finally, there is research into making autonomous AI agents safer when they interact with decentralized finance markets by modeling them as searchers rather than just bots. The study found that an adaptive path-selection algorithm outperformed standard baselines by eleven percent on average, and moderate randomization effectively cut exposure to maximal extractable value by over fifty percent. This suggests a way to manage the risks associated with agentic trading in complex environments.

The most critical finding relates to how data can leak out of analog circuits through seemingly input-only pins in mixed-signal systems, which matters because it reveals a fundamental gap in treating pin directionality as a security concern. They experimentally showed that data-dependent circuit-offset modulation can turn these nominally input pins into outbound information channels.

This attack works when there is a closed-loop amplifier, an exposed amplifier input, and the pin has high impedance. They validated this using a photoplethysmography analog front-end fabricated in a commercial fifty five nanometer CMOS process. When the payload was activated, it only reduced the filtered PPG output signal-to-noise ratio by zero point zero three decibels, and the maximum perturbation of five point nine percent of the PPG amplitude stayed within thirty four point three percent variation across process and temperature.

This means that while raw exfiltration signal-to-interference-plus-noise ratio was below negative twenty decibels, targeted filtering could increase it above fourteen decibels, allowing for signal recovery. Silicon measurements confirmed data exfiltration through the input pin at bit rates up to ten kilobits per second and error free recovery of a pulse length random bit sequence message. This work establishes that analog pin directionality is an AMS security property that needs explicit verification, not just inference from nominal signal flow.

Today's papers

The papers

Important terms

Herfindahl-Hirschman Index (HHI)
This index measures market concentration among autonomous system numbers, showing that infrastructure is highly consolidated, indicating a market far above the threshold for significant concentration.
Trusted Back-Translation
A method used by the Echo system to verify source code. It recompiles assembly and checks if it exactly matches the target binary, providing a strong way to confirm decompiled code correctness.
LLM Annotations for Fuzzing
Using large language models to automatically create annotations for fuzzing campaigns. This helps scale annotation-based vulnerability detection without needing massive human teams.
Privacy Exposure Displacement
A concept from ASLEval that measures the mismatch between what is seen locally and what is actually exposed across an entire multi-step AI session, revealing hidden privacy risks.