Security papers — 2026-09-11

Today we are diving into how we can actually trust the decisions made by these increasingly autonomous large language model agents because simply getting a final answer isn't enough. The core issue is that we need to understand the evidence supporting every step an agent takes, whether it was justifying a tool call or how its memory shaped a later choice. This idea of execution provenance, which we define as the typed graph of an agent's run and evidence tracing as its projection onto evidence-support relations, connects everything from retrieval grounding to debugging.

We looked at several ways to build this framework, including different methods for provenance representation and how to attribute evidence across various units. A key direction involves developing runtime guardrails that can monitor these traces in real time, which is closely related to how we think about observability and failure diagnosis. This work on agent tracing matters because it moves us toward building systems that are not just smart, but auditable and recoverable when they go wrong.

Beyond the agents, there is important foundational work on securing computation itself. We saw a new system called mmFHE that executes the entire mmWave sensing pipeline under fully homomorphic encryption. This means the cloud can process sensitive data without ever seeing it in plaintext, even though it introduces some latency. This approach proves input privacy and data obliviousness for tasks like vital-sign monitoring.

Another area of focus is making heavy computation practical through hardware acceleration. The PHAT project proposes a photonic accelerator for TFHE that uses Optically-addressed Phase-Change Memory to speed up the FFT operations needed in fully homomorphic encryption schemes. This accelerator shows a significant speedup over existing ASIC accelerators, suggesting we are getting closer to practical privacy-preserving cloud computing solutions.

Finally, we are also looking at how defenses interact with each other. A new Python library called Amulet is being introduced to systematically evaluate both intended and unintended interactions among machine learning defenses and various risks. This offers a unified way to study these complex relationships.

The most important work right now is SpecGuard because it provides a way to catch hidden backdoors in large language models without adding any extra computational load during the actual use of the model. This matters because these models are so widely deployed, and if they have secret triggers that cause them to behave maliciously, we need a fast way to check for that behavior while they are running.

SpecGuard achieves this by using speculative decoding, which is a technique where a small draft model proposes tokens and the main target model checks those proposals. The key finding is that when a backdoor is triggered, the target model shifts toward the attacker's behavior, but the clean draft model does not show this shift. This causes a change in how often tokens are accepted from the draft.

This means SpecGuard doubles as a free way to monitor for these malicious triggers by observing how the verification process reacts. Another piece of work addresses how much unwanted automated speech is being placed on phone calls, which is important because it relates to regulatory concerns like the TCPA. Researchers used an interactive voice honeypot that recorded over sixty-six days of calls.

They found that machine-voiced openings account for at least twenty-seven point nine percent of all openings. This suggests a significant amount of automated speech is being used in these interactions, even if it is not always clearly identifiable as synthetic. This finding connects to the idea that detection methods need to be robust against different types of attacks, which is why SpecGuard works across diverse backdoor types and model families.

Furthermore, the analysis on phone calls shows that while synthetic openings concentrate in lead-generation spam rather than fraud, the prevalence of machine-voiced openings is still substantial. On a more technical level, there is work on tracing illicit funds across Solana bridges using a method called SolTracer. This system maps different execution semantics into one space to reliably link transactions even when the underlying blockchain lacks standard event logs.

This improved tracing method showed a twenty point one six percent improvement in performance over the best existing methods in complex open-world scenarios. This tracing work contrasts with the privacy auditing research, which introduces Zero-Run auditing for large models. This framework allows for privacy checks using only known training and non-training examples, offering a practical way to evaluate privacy without needing access to the entire training pipeline.

The concept of detecting malicious behavior also extends into how models respond to prompts, as seen in the research on in-context multimodality jailbreaks. This work proposes that jailbreaks are evidence accumulation processes where harmful demonstrations shift the model's internal preference between safe and harmful modes. This leads to a defense that injects counter-evidence based on estimated risk, aiming to suppress this harmful drift while keeping the model useful.

Finally, there is a method for auditing privacy in black-box settings using Word-level Probability MIA. This technique estimates word probabilities through Monte Carlo sampling and finds that it consistently outperforms existing black-box baselines when trying to determine if a specific text was in the training set of proprietary models.

The most significant development concerns how large language model agents are beginning to tackle penetration testing tasks autonomously. We saw that a newer autonomous system running Claude Opus 4.8 successfully solved all three public targets we tested. This included two specific challenges that the older human-in-the-loop system using Kimi K2.5 never managed to finish at all.

This suggests that increased autonomy, when paired with a more capable model, allows agents to complete complex sequences of actions end-to-end. The legacy system's performance was also quite telling; even on machines where it failed to solve certain subtasks, it managed to complete about half of them when run without provider guardrails on standard university GPUs.

This points toward a trend where the agent's ability to plan and commit to a route is proving more important for success than having perfect long-horizon memory. We tested this by adding a coverage-memory layer to both systems, but neither modification improved the outcomes. This suggests that lost memory is not the primary bottleneck.

Instead, in the stalled runs we reviewed, it seemed planning and commitment were limiting factors because agents held evidence for a path forward but failed to turn that evidence into a concrete exploitation hypothesis. This hints that offensive capability might advance with better planning ability rather than solely with improved memory retention.

The work on super-apps matters because they have become the default trusted intermediary for users accessing many different services, and research shows this implicit trust is dangerously misplaced. We see this potential danger in Russia's MAX, whose parent company is linked to state prosecution of online speech, suggesting it could silently undermine user privacy.

This capability allows MAX to capture mini-app user interfaces and inject arbitrary code into other applications without detection. This ability to compromise the integrity of other apps is a major concern because it shows that malicious super-apps can operate in total stealth. This is supported by findings showing that these architectural privileges are inherent, meaning any super-app has the potential to perform these actions.

This contrasts with work on agent payment protocols like AP2, which shows how seemingly valid transactions can be steered toward unintended outcomes through subtle text descriptions. The AP2 research demonstrated that ordinary product descriptions can trick shopping agents into fetching another user's payment details or assembling a cart that doesn't match what the user actually intended.

This vulnerability was shown to succeed at high rates across various models, meaning the protocol itself doesn't constrain the final decision made by the agent. To counter this steering effect, researchers developed A-VIP, a protocol-layer defense that binds every credential lookup to the specific session and cart line to the listing seen. This defense successfully blocked structural attacks while surfacing unauthorized spending when a third attack left no trace.

Meanwhile, in the realm of large language models integrated into security operations, there is a need for robust defenses against prompt injection via log poisoning. A neurosymbolic framework was proposed that uses deterministic pre-filters and semantic boundary enforcement to neutralize malicious payloads before they reach the LLM processing stage. This approach aims to bound the stochastic nature of neural evaluations with verifiable constraints, creating a more resilient defense mechanism for AI-SOCs.

The work on Adaptive Diffusion Freezing matters because it directly tackles the privacy concerns inherent in training large generative models, specifically defending against membership inference attacks by finding a better balance between keeping the model useful and keeping user data private. This framework works by using cross-timestep adaptive freezing training to control how much different data subsets are allowed to influence the model at various stages of diffusion.

This helps reduce over-memorization and makes the model behave more uniformly for both members and nonmembers. The core mechanism involves creating a risk-aware freezing policy that estimates membership inference attack risk based on memorization tendencies. This then suppresses the contribution of data subsets paired with higher risk.

This technique is built upon pretraining to construct a freezing mask matrix designed to reduce leakage without harming generation quality. This mask matrix is then compared against various baselines in evaluations across multiple datasets. This approach is significant because it demonstrates an effective defense performance alongside state-of-the-art privacy utility efficiency trade-off compared to existing methods.

This concept of controlling data participation connects to other areas where verifiable or adaptive mechanisms are being explored. For example, Atlas achieves verifiable semantic search by restructuring HNSW into a fixed-size state procedure that is proven to return the same result. DriftNet uses a dual-head trajectory Transformer to classify tool-call trajectories.

While ADF focuses on model training privacy, these other works address trust and verification in different contexts. Atlas proves query correctness against an index without revealing it, and BlueSTAR builds a tiered architecture for autonomous cyber defense that handles complex reasoning.

The most critical area of progress involves developing ways to secure hardware against static side-channel attacks because these attacks pose an increasing threat to chip security by exploiting halted clock conditions to extract sensitive information. This is significant because even if we have strong cryptographic algorithms, the physical implementation can leak secrets through timing or power variations.

We developed Chypothermia which works by exposing a chip to cryogenic temperatures. This interference disrupts the on-chip mixed-signal components responsible for signal sensing and generation. This attack disables the target clock sensor, clock generation circuit, and voltage sensors without needing any electrical tampering while keeping the secret data safe.

This is effective at stopping the clock but cooling is slow enough that it can be bypassed by systems with temperature sensors designed to catch thermal anomalies. To overcome this limitation, we combined Chypothermia with Chypnosis to show that even in moderately low-temperature operating ranges, we could halt the clock while avoiding detection.

We tested this combination on several FPGA and SoC platforms and successfully disabled both soft-IP and hard-IP sensor implementations. Furthermore, applying Chypothermia to the alert handler of the OpenTitan root of trust proved that it evades detection and prevents key zeroization.

Another important piece is bridging the gap between formal protocol specifications and real-world application behavior for protocols like Signal. We applied SpecMon to WhatsApp Web and Signal Desktop to check if their actual executions matched formal models. This resulted in multiset-rewrite models compatible with Tamarin.

This monitoring confirmed that observed executions conformed to these models, verifying properties like authentication and secrecy for the core components of the Signal protocol. Finally, we are looking at how autonomous agents lose control when performing long-horizon tasks involving tool use and persistent state.

Our central hypothesis is that a degraded control boundary becomes consequential when the environment exposes an executable action that crosses it, even if the underlying task remains legitimate. We found that when both degraded control and unsafe opportunity are present, the loss-of-control rate reaches fifty five percent across a full factorial study.

The most important takeaway is the creation of the first formal definition for frontrunning vulnerability because it shifts the focus from just looking at code to understanding user interaction. This new definition shows that a contract's ability to resist this attack isn't just about its internal logic, but how honest users choose to use it.

We developed an algorithm designed to synthesize these secure interaction conditions based on this new understanding of resistance. This algorithm is sound because it is built directly upon the formal definition we established, which captures that user behavior is key. We then tested this by applying the prototype implementation to two real-world Ethereum contracts, which uncovered previously undiscovered vulnerabilities in those specific programs.

Today's papers

The papers

Important terms

Execution Provenance
This is a typed graph that traces every step an autonomous agent takes during its run, including justifications for tool calls and how its memory influenced later decisions. It's crucial for understanding and debugging agent actions.
SpecGuard
A technique that catches hidden backdoors in large language models without adding extra computational load. It uses speculative decoding to monitor if a backdoor is triggered by observing changes in token acceptance rates.
Fully Homomorphic Encryption (FHE)
This allows cloud processing of sensitive data, like vital signs, without ever decrypting it. While it adds latency, it ensures input privacy and data obliviousness for secure computation.
Amulet
A new Python library designed to systematically evaluate the interactions between different machine learning defenses and various potential risks in a unified way.