Security papers — 2026-09-11
Today we are diving into how we can actually trust the decisions made by these increasingly autonomous large language model agents because simply getting a final answer isn't enough. The core issue is that we need to understand the evidence supporting every step an agent takes, whether it was justifying a tool call or how its memory shaped a later choice. This idea of execution provenance, which we define as the typed graph of an agent's run and evidence tracing as its projection onto evidence-support relations, connects everything from retrieval grounding to debugging.
We looked at several ways to build this framework, including different methods for provenance representation and how to attribute evidence across various units. A key direction involves developing runtime guardrails that can monitor these traces in real time, which is closely related to how we think about observability and failure diagnosis. This work on agent tracing matters because it moves us toward building systems that are not just smart, but auditable and recoverable when they go wrong.
Beyond the agents, there is important foundational work on securing computation itself. We saw a new system called mmFHE that executes the entire mmWave sensing pipeline under fully homomorphic encryption. This means the cloud can process sensitive data without ever seeing it in plaintext, even though it introduces some latency. This approach proves input privacy and data obliviousness for tasks like vital-sign monitoring.
Another area of focus is making heavy computation practical through hardware acceleration. The PHAT project proposes a photonic accelerator for TFHE that uses Optically-addressed Phase-Change Memory to speed up the FFT operations needed in fully homomorphic encryption schemes. This accelerator shows a significant speedup over existing ASIC accelerators, suggesting we are getting closer to practical privacy-preserving cloud computing solutions.
Finally, we are also looking at how defenses interact with each other. A new Python library called Amulet is being introduced to systematically evaluate both intended and unintended interactions among machine learning defenses and various risks. This offers a unified way to study these complex relationships.
The most important work right now is SpecGuard because it provides a way to catch hidden backdoors in large language models without adding any extra computational load during the actual use of the model. This matters because these models are so widely deployed, and if they have secret triggers that cause them to behave maliciously, we need a fast way to check for that behavior while they are running.
SpecGuard achieves this by using speculative decoding, which is a technique where a small draft model proposes tokens and the main target model checks those proposals. The key finding is that when a backdoor is triggered, the target model shifts toward the attacker's behavior, but the clean draft model does not show this shift. This causes a change in how often tokens are accepted from the draft.
This means SpecGuard doubles as a free way to monitor for these malicious triggers by observing how the verification process reacts. Another piece of work addresses how much unwanted automated speech is being placed on phone calls, which is important because it relates to regulatory concerns like the TCPA. Researchers used an interactive voice honeypot that recorded over sixty-six days of calls.
They found that machine-voiced openings account for at least twenty-seven point nine percent of all openings. This suggests a significant amount of automated speech is being used in these interactions, even if it is not always clearly identifiable as synthetic. This finding connects to the idea that detection methods need to be robust against different types of attacks, which is why SpecGuard works across diverse backdoor types and model families.
Furthermore, the analysis on phone calls shows that while synthetic openings concentrate in lead-generation spam rather than fraud, the prevalence of machine-voiced openings is still substantial. On a more technical level, there is work on tracing illicit funds across Solana bridges using a method called SolTracer. This system maps different execution semantics into one space to reliably link transactions even when the underlying blockchain lacks standard event logs.
This improved tracing method showed a twenty point one six percent improvement in performance over the best existing methods in complex open-world scenarios. This tracing work contrasts with the privacy auditing research, which introduces Zero-Run auditing for large models. This framework allows for privacy checks using only known training and non-training examples, offering a practical way to evaluate privacy without needing access to the entire training pipeline.
The concept of detecting malicious behavior also extends into how models respond to prompts, as seen in the research on in-context multimodality jailbreaks. This work proposes that jailbreaks are evidence accumulation processes where harmful demonstrations shift the model's internal preference between safe and harmful modes. This leads to a defense that injects counter-evidence based on estimated risk, aiming to suppress this harmful drift while keeping the model useful.
Finally, there is a method for auditing privacy in black-box settings using Word-level Probability MIA. This technique estimates word probabilities through Monte Carlo sampling and finds that it consistently outperforms existing black-box baselines when trying to determine if a specific text was in the training set of proprietary models.
The most significant development concerns how large language model agents are beginning to tackle penetration testing tasks autonomously. We saw that a newer autonomous system running Claude Opus 4.8 successfully solved all three public targets we tested. This included two specific challenges that the older human-in-the-loop system using Kimi K2.5 never managed to finish at all.
This suggests that increased autonomy, when paired with a more capable model, allows agents to complete complex sequences of actions end-to-end. The legacy system's performance was also quite telling; even on machines where it failed to solve certain subtasks, it managed to complete about half of them when run without provider guardrails on standard university GPUs.
This points toward a trend where the agent's ability to plan and commit to a route is proving more important for success than having perfect long-horizon memory. We tested this by adding a coverage-memory layer to both systems, but neither modification improved the outcomes. This suggests that lost memory is not the primary bottleneck.
Instead, in the stalled runs we reviewed, it seemed planning and commitment were limiting factors because agents held evidence for a path forward but failed to turn that evidence into a concrete exploitation hypothesis. This hints that offensive capability might advance with better planning ability rather than solely with improved memory retention.
The work on super-apps matters because they have become the default trusted intermediary for users accessing many different services, and research shows this implicit trust is dangerously misplaced. We see this potential danger in Russia's MAX, whose parent company is linked to state prosecution of online speech, suggesting it could silently undermine user privacy.
This capability allows MAX to capture mini-app user interfaces and inject arbitrary code into other applications without detection. This ability to compromise the integrity of other apps is a major concern because it shows that malicious super-apps can operate in total stealth. This is supported by findings showing that these architectural privileges are inherent, meaning any super-app has the potential to perform these actions.
This contrasts with work on agent payment protocols like AP2, which shows how seemingly valid transactions can be steered toward unintended outcomes through subtle text descriptions. The AP2 research demonstrated that ordinary product descriptions can trick shopping agents into fetching another user's payment details or assembling a cart that doesn't match what the user actually intended.
This vulnerability was shown to succeed at high rates across various models, meaning the protocol itself doesn't constrain the final decision made by the agent. To counter this steering effect, researchers developed A-VIP, a protocol-layer defense that binds every credential lookup to the specific session and cart line to the listing seen. This defense successfully blocked structural attacks while surfacing unauthorized spending when a third attack left no trace.
Meanwhile, in the realm of large language models integrated into security operations, there is a need for robust defenses against prompt injection via log poisoning. A neurosymbolic framework was proposed that uses deterministic pre-filters and semantic boundary enforcement to neutralize malicious payloads before they reach the LLM processing stage. This approach aims to bound the stochastic nature of neural evaluations with verifiable constraints, creating a more resilient defense mechanism for AI-SOCs.
The work on Adaptive Diffusion Freezing matters because it directly tackles the privacy concerns inherent in training large generative models, specifically defending against membership inference attacks by finding a better balance between keeping the model useful and keeping user data private. This framework works by using cross-timestep adaptive freezing training to control how much different data subsets are allowed to influence the model at various stages of diffusion.
This helps reduce over-memorization and makes the model behave more uniformly for both members and nonmembers. The core mechanism involves creating a risk-aware freezing policy that estimates membership inference attack risk based on memorization tendencies. This then suppresses the contribution of data subsets paired with higher risk.
This technique is built upon pretraining to construct a freezing mask matrix designed to reduce leakage without harming generation quality. This mask matrix is then compared against various baselines in evaluations across multiple datasets. This approach is significant because it demonstrates an effective defense performance alongside state-of-the-art privacy utility efficiency trade-off compared to existing methods.
This concept of controlling data participation connects to other areas where verifiable or adaptive mechanisms are being explored. For example, Atlas achieves verifiable semantic search by restructuring HNSW into a fixed-size state procedure that is proven to return the same result. DriftNet uses a dual-head trajectory Transformer to classify tool-call trajectories.
While ADF focuses on model training privacy, these other works address trust and verification in different contexts. Atlas proves query correctness against an index without revealing it, and BlueSTAR builds a tiered architecture for autonomous cyber defense that handles complex reasoning.
The most critical area of progress involves developing ways to secure hardware against static side-channel attacks because these attacks pose an increasing threat to chip security by exploiting halted clock conditions to extract sensitive information. This is significant because even if we have strong cryptographic algorithms, the physical implementation can leak secrets through timing or power variations.
We developed Chypothermia which works by exposing a chip to cryogenic temperatures. This interference disrupts the on-chip mixed-signal components responsible for signal sensing and generation. This attack disables the target clock sensor, clock generation circuit, and voltage sensors without needing any electrical tampering while keeping the secret data safe.
This is effective at stopping the clock but cooling is slow enough that it can be bypassed by systems with temperature sensors designed to catch thermal anomalies. To overcome this limitation, we combined Chypothermia with Chypnosis to show that even in moderately low-temperature operating ranges, we could halt the clock while avoiding detection.
We tested this combination on several FPGA and SoC platforms and successfully disabled both soft-IP and hard-IP sensor implementations. Furthermore, applying Chypothermia to the alert handler of the OpenTitan root of trust proved that it evades detection and prevents key zeroization.
Another important piece is bridging the gap between formal protocol specifications and real-world application behavior for protocols like Signal. We applied SpecMon to WhatsApp Web and Signal Desktop to check if their actual executions matched formal models. This resulted in multiset-rewrite models compatible with Tamarin.
This monitoring confirmed that observed executions conformed to these models, verifying properties like authentication and secrecy for the core components of the Signal protocol. Finally, we are looking at how autonomous agents lose control when performing long-horizon tasks involving tool use and persistent state.
Our central hypothesis is that a degraded control boundary becomes consequential when the environment exposes an executable action that crosses it, even if the underlying task remains legitimate. We found that when both degraded control and unsafe opportunity are present, the loss-of-control rate reaches fifty five percent across a full factorial study.
The most important takeaway is the creation of the first formal definition for frontrunning vulnerability because it shifts the focus from just looking at code to understanding user interaction. This new definition shows that a contract's ability to resist this attack isn't just about its internal logic, but how honest users choose to use it.
We developed an algorithm designed to synthesize these secure interaction conditions based on this new understanding of resistance. This algorithm is sound because it is built directly upon the formal definition we established, which captures that user behavior is key. We then tested this by applying the prototype implementation to two real-world Ethereum contracts, which uncovered previously undiscovered vulnerabilities in those specific programs.
Today's papers
- From Agent Traces to Trust: A Survey of Evidence Tracing and Execution Provenance in LLM Agents Large language model agents are evolving into autonomous systems whose behavior needs to be verified through evidence tracing and execution provenance. [paper] [episode]
- Amulet: a Python Library for Assessing Interactions Among ML Defenses and Risks This library helps evaluate how different machine learning defenses interact with each other, both intended and unintended interactions. [paper] [episode]
- A Survey of Threats Against Voice Authentication and Anti-Spoofing Systems This survey reviews the modern threats against voice authentication systems, including deepfakes and adversarial attacks. [paper] [episode]
- mmFHE: mmWave Sensing with End-to-End Fully Homomorphic Encryption This system allows for the entire mmWave sensing pipeline to be executed securely on a cloud using fully homomorphic encryption. [paper] [episode]
- PHAT: PHotonic Accelerator for TFHE This paper proposes a photonic accelerator to speed up fully homomorphic encryption by optimizing FFT operations. [paper]
- ToxicRAG: Compromising Retrieval-Augmented Generation Systems via Single-Shot Knowledge Poisoning Attacks This study shows how poisoned documents can successfully manipulate retrieval augmented generation systems through narrative attacks. [paper]
- No-Box Vulnerability Analysis: Description-only Detection of Indirect Prompt Injection Vulnerabilities in MCP Servers This work introduces a method to find indirect prompt injection vulnerabilities by analyzing only the metadata of a system without direct access. [paper]
- Few-Shot Learning for Network Intrusion Detection: Methods, Datasets, and Performance This paper systematically reviews few-shot learning approaches for training anomaly-based network intrusion detection systems. [paper]
- SpecGuard: Inference-Time Backdoor Detection For Free This method uses speculative decoding to detect hidden backdoors in large language models without adding extra model computation cost. [paper]
- The Machines Are Calling: Measuring Automated and Synthetic Voices in Unwanted Inbound Calls This research measures the prevalence of machine-voiced calls and synthetic speech in unwanted inbound traffic. [paper]
- Heterogeneous Cross-Chain Transaction Tracing for Solana Bridges via Candidate-Set Selective Decision This paper proposes a method to reliably trace cross-chain transactions on Solana by mapping disparate execution semantics into a unified event space. [paper]
- Empirical Evaluation of Data Poisoning Attacks in Supervised Learning This study evaluates how different data poisoning attacks, like label flipping and backdoor poisoning, affect the performance of various supervised learning models. [paper]
- Privacy Auditing with Zero (0) Training Run This framework allows for privacy auditing of models post-hoc using only fixed datasets without requiring any intervention during training. [paper]
- Black-Box Membership Inference via Word-Level Probability Estimation This method estimates word-level probabilities to perform membership inference attacks on large language models in a black-box setting. [paper]
- Understanding In-Context Multimodal Jailbreaks via Posterior Reweighting This framework models harmful demonstrations as evidence that shifts the model's internal preference between safe and harmful behaviors. [paper]
- You've Got a BUD in Me: Authenticated Reads from Per-Block Write Logs This paper introduces a Block Update Digest to provide authenticated reads for historical membership proofs on blockchains. [paper]
- Big Enough to Break Out: Tracking the Rising Capability of LLM Penetration-Testing Agents This study compares different LLM agents used for penetration testing and tracks their increasing capability in solving complex tasks. [paper]
- An Empirical Measurement of Jailbreaking Evaluators This paper systematically compares six automated jailbreak evaluators to determine which one performs best on human-labeled data. [paper]
- BenchShield: Formal Model-Backed Instrumentation for Reward Integrity in LLM-Agent Evaluation Infrastructure This tool provides instrumentation to detect reward hacking during LLM agent evaluations by analyzing the lifecycle of reward-relevant events. [paper]
- Temporal and Multimodal Deep Learning for Cyberattack Detection in LEO Satellite Systems This work investigates using structured multimodal deep learning to detect cyberattacks across heterogeneous data sources from low-earth orbit satellites. [paper]
- Deep-Fake CAPTCHA: Mitigating Next-Generation Social Engineering Attacks This paper proposes an active defense against deepfake impersonation by requiring callers to complete simple challenge-response tasks. [paper]
- SoK: Privacy Attacks on Machine Learning via Explainable AI This study analyzes how explanations for machine learning models can be exploited to extract private information and data. [paper]
- A2ABreak: Systematic Security Analysis of the A2A Protocol This paper provides a rigorous security analysis of the Agent2Agent protocol, uncovering new vulnerabilities through formal modeling. [paper]
- From Intent to Execution Grant: An Execution-Boundary Conformance Profile for High-Risk AI Actions This paper specifies a profile for deciding when an AI action can receive execution authority based on a structured intent object. [paper]
- Don't Trust the Super-App: A Case Study of Russia's Max This paper demonstrates how malicious super-apps can silently compromise the security and privacy of mini-apps within a mobile architecture. [paper]
- Signing the Transaction but Not the Decision: Whisper Attacks and a Binding Defense for AP2 This paper introduces a defense protocol to prevent software agents from being steered into unintended purchases through transaction signing. [paper]
- From Cycle Space to Cycle Manifold: Limits and Achievability of Blind False Data Injection Attacks This work establishes that the weighted cycle space is necessary and sufficient for performing blind false data injection attacks on IEEE systems. [paper]
- Accountability in Certificate Transparency and Variants This paper analyzes how Certificate Transparency aims to reduce trust in certificate authorities by providing accountability for certificate issuance logs. [paper]
- Beyond Static Guarantees: Measuring the Static-Pass Dynamic-Fail Gap in Security-Sensitive and LLM-Generated Python Code This paper introduces a pipeline to detect vulnerabilities that pass static analysis but fail during dynamic execution. [paper]
- AspisAI: A Canonical, Machine-Interpretable Governance Framework for Automated Multi-Standard Compliance Monitoring This framework translates multiple cybersecurity standards into a machine-interpretable control model for automated compliance monitoring. [paper]
- Architecting the Secure AI-SOC: A Neurosymbolic Framework for Pipeline Integrity and Threat Mitigation This paper proposes a neurosymbolic defense architecture to secure LLM security operations centers against prompt injection attacks. [paper]
- PEARL: A Task-Aware Framework for Evaluating Differentially Private Synthetic Educational Data This framework approves synthetic educational datasets only when they meet strict criteria for privacy, usefulness, and suitability for personalized learning. [paper]
- Adaptive Diffusion Freezing: Privacy-preserving Diffusion Models Against Membership Inference Attacks This method uses cross-timestep adaptive freezing in diffusion models to defend against membership inference attacks while maintaining utility. [paper]
- Atlas: Efficient Verifiable Semantic Search This system proves that semantic search can be made verifiable at scale by using zero-knowledge proofs for graph-based retrieval algorithms. [paper]
- Empirical Evaluation of Membership Inference Attacks on NLP Text Classifiers: A Baseline Study on SST-2 This study compares the privacy leakage of different text classification models under membership inference attacks. [paper]
- BlueSTAR: Tiered Agentic Architecture for Autonomous Cyber Defense This paper presents a tiered agentic architecture to autonomously defend IT/OT networks against automated cyber attacks. [paper]
- DriftNet: A Dual-Head Trajectory Transformer for Detecting and Localizing Prompt Injection in LLM Agents This model uses a dual-head transformer to classify and localize prompt injection within an LLM agent's tool-call trajectory. [paper]
- Differentially Private EEG Feature Anonymization: A Privacy-Utility Case Study in Clinical Neurophysiology This study investigates the trade-off between privacy protection and utility when applying differential privacy to clinical EEG features. [paper]
- Lower Bounds for PIR with Preprocessing from Blackbox Cryptography This paper establishes computation lower bounds for private information retrieval schemes that use preprocessing to achieve sublinear query time. [paper]
- Threshold Choice, Not Sample Size, Bounds Trustless Verification of Nondeterministic Compound AI Workflows This protocol provides a trustless verification method for nondeterministic compound AI workflows by focusing on the threshold rather than sample size. [paper]
- Chypothermia: Clock Freezing for Static Side-channel Attacks This paper introduces an attack that disables clock sensors and generation circuits using cryogenic temperatures to bypass static side-channel defenses. [paper]
- From Specs to Apps: Verifying and Monitoring Models of Signal and WhatsApp This work applies a runtime monitor to verify that the actual execution of messaging applications conforms to their formal protocol specifications. [paper]
- The Missing Boundary: How Autonomous Agents Lose Control This study investigates how a loss of control can emerge when autonomous agents pursue legitimate tasks in complex environments with executable unsafe opportunities. [paper]
- HermiCache: Enclave-Aware Cache Replacement for Trusted Execution Environments This paper introduces HermiCache, a cache replacement policy designed to protect trusted execution environments from cache-based side-channel attacks. [paper]
- DeFiFusion: Combining Transaction Events with Smart Contracts to Detect Price Manipulation Attacks This framework fuses transaction events and smart contract semantics to detect price manipulation attacks in decentralized finance. [paper]
- A Systematic Study of TEE Build Reproducibility in the Wild This paper investigates the reproducibility of trusted execution environment builds across different platforms and reveals ecosystem-level challenges.
- You Get What You Sample: Evaluating Sampling Strategies for Web Security Measurements This work compares different sampling strategies used in web security measurements to determine which ones provide unbiased estimates of vulnerability prevalence. [paper]
- CARTS: Contextual Autoregressive Rank Transcoding Steganography for Full-Capacity Keyed Text Encoding This paper formally analyzes the security properties of using autoregressive language models to create steganographic payloads. [paper]
- On Identifying Sound Conditions for Frontrunning Resistance This paper proposes a formal definition of frontrunning vulnerability for smart contracts based on how honest users interact with them. [paper]
The papers
- From Agent Traces to Trust: A Survey of Evidence Tracing and Execution Provenance in LLM Agents — This survey paper addresses the "process-level accountability gap" emerging in Large Language Model (LLM)-based agents. [episode]
- Amulet: a Python Library for Assessing Interactions Among ML Defenses and Risks — I apologize, but the provided text appears to be an excerpt from a bibliography page (Page 13) containing citations for various machine learning security papers. [episode]
- A Survey of Threats Against Voice Authentication and Anti-Spoofing Systems — I apologize, but you have provided only a section of a bibliography and not the actual content of the paper, "A Survey of Threats Against Voice Authentication and Anti-Spoofing Systems." To fulfill your request—which requires extracting specific details, quoting key phrases, an [episode]
- mmFHE: mmWave Sensing with End-to-End Fully Homomorphic Encryption — The paper "mmFHE: mmWave Sensing with End-to-End Fully Homomorphic Encryption" presents a novel framework for performing complex radar signal processing entirely within an encrypted environment. [episode]
- Threshold Choice, Not Sample Size, Bounds Trustless Verification of Nondeterministic Compound AI Workflows —
- Adaptive Diffusion Freezing: Privacy-preserving Diffusion Models Against Membership Inference Attacks —
- Black-Box Membership Inference via Word-Level Probability Estimation —
- PEARL: A Task-Aware Framework for Evaluating Differentially Private Synthetic Educational Data —
- Understanding In-Context Multimodal Jailbreaks via Posterior Reweighting —
- SoK: Privacy Attacks on Machine Learning via Explainable AI —
- From Cycle Space to Cycle Manifold: Limits and Achievability of Blind False Data Injection Attacks —
- HermiCache: Enclave-Aware Cache Replacement for Trusted Execution Environments —
- Architecting the Secure AI-SOC: A Neurosymbolic Framework for Pipeline Integrity and Threat Mitigation —
- CARTS: Contextual Autoregressive Rank Transcoding Steganography for Full-Capacity Keyed Text Encoding —
- Temporal and Multimodal Deep Learning for Cyberattack Detection in LEO Satellite Systems —
- Beyond Static Guarantees: Measuring the Static-Pass Dynamic-Fail Gap in Security-Sensitive and LLM-Generated Python Code —
- Big Enough to Break Out: Tracking the Rising Capability of LLM Penetration-Testing Agents —
- No-Box Vulnerability Analysis: Description-only Detection of Indirect Prompt Injection Vulnerabilities in MCP Servers —
- A2ABreak: Systematic Security Analysis of the A2A Protocol —
- AspisAI: A Canonical, Machine-Interpretable Governance Framework for Automated Multi-Standard Compliance Monitoring —
- DriftNet: A Dual-Head Trajectory Transformer for Detecting and Localizing Prompt Injection in LLM Agents —
- Empirical Evaluation of Membership Inference Attacks on NLP Text Classifiers: A Baseline Study on SST-2 —
- Empirical Evaluation of Data Poisoning Attacks in Supervised Learning —
- DeFiFusion: Combining Transaction Events with Smart Contracts to Detect Price Manipulation Attacks —
- The Missing Boundary: How Autonomous Agents Lose Control —
- BenchShield: Formal Model-Backed Instrumentation for Reward Integrity in LLM-Agent Evaluation Infrastructure —
- ToxicRAG: Compromising Retrieval-Augmented Generation Systems via Single-Shot Knowledge Poisoning Attacks —
- The Machines Are Calling: Measuring Automated and Synthetic Voices in Unwanted Inbound Calls —
- You Get What You Sample: Evaluating Sampling Strategies for Web Security Measurements —
- You've Got a BUD in Me: Authenticated Reads from Per-Block Write Logs —
- Few-Shot Learning for Network Intrusion Detection: Methods, Datasets, and Performance —
- Deep-Fake CAPTCHA: Mitigating Next-Generation Social Engineering Attacks —
- "They don't care about this": A Systematic Study of TEE Build Reproducibility in the Wild —
- Heterogeneous Cross-Chain Transaction Tracing for Solana Bridges via Candidate-Set Selective Decision —
- Chypothermia: Clock Freezing for Static Side-channel Attacks —
- On Identifying Sound Conditions for Frontrunning Resistance —
- Accountability in Certificate Transparency and Variants —
- From Intent to Execution Grant: An Execution-Boundary Conformance Profile for High-Risk AI Actions —
- PHAT: PHotonic Accelerator for TFHE —
- Signing the Transaction but Not the Decision: Whisper Attacks and a Binding Defense for AP2 —
- Differentially Private EEG Feature Anonymization: A Privacy-Utility Case Study in Clinical Neurophysiology —
- SpecGuard: Inference-Time Backdoor Detection For Free —
- Don't Trust the Super-App: A Case Study of Russia's Max —
- Atlas: Efficient Verifiable Semantic Search —
- BlueSTAR: Tiered Agentic Architecture for Autonomous Cyber Defense —
- From Specs to Apps: Verifying and Monitoring Models of Signal and WhatsApp —
- Privacy Auditing with Zero (0) Training Run —
- Lower Bounds for PIR with Preprocessing from Blackbox Cryptography —
- An Empirical Measurement of Jailbreaking Evaluators —
Important terms
- Execution Provenance
- This is a typed graph that traces every step an autonomous agent takes during its run, including justifications for tool calls and how its memory influenced later decisions. It's crucial for understanding and debugging agent actions.
- SpecGuard
- A technique that catches hidden backdoors in large language models without adding extra computational load. It uses speculative decoding to monitor if a backdoor is triggered by observing changes in token acceptance rates.
- Fully Homomorphic Encryption (FHE)
- This allows cloud processing of sensitive data, like vital signs, without ever decrypting it. While it adds latency, it ensures input privacy and data obliviousness for secure computation.
- Amulet
- A new Python library designed to systematically evaluate the interactions between different machine learning defenses and various potential risks in a unified way.