Security papers — 2026-10-10
Work was done on mAVE, a watermark method for joint audio-visual generation models to track the origin of generated media. This is important because knowing where outputs come from is crucial for safety and trust as these agents become more sophisticated.
Researchers also explored false claims in commercial image generators using a red-teaming benchmark to test how easily deceptive images can be created. This work helps us understand the limits of current generation technology when it comes to creating believable but untrue content.
There is a problem with certifying hidden paths in quantum key distribution networks through scalable topology assurance, which is important for securing future communication infrastructure. This connects to agent work because reliable communication channels are a prerequisite for secure agent operation.
Context-binding gaps in stateful zero-knowledge proximity proofs were also looked into, dealing with how context can leak or be misused in complex cryptographic checks. This is a technical hurdle that needs to be cleared before deploying agents that rely on these proofs for verification.
DCVD was also touched upon, which uses dual-channel cross-modal fusion for joint vulnerability detection and localization. This provides a way to pinpoint exactly where security flaws might exist within the system architecture being secured.
The most pressing issue today revolves around the security of large language models through various attack vectors. Work on Phantom Transfer explored how data poisoning attacks can evade existing data-level defenses, meaning malicious actors can still inject poisoned information that survives initial filtering.
A related concern is how these models are being attacked at the serving level. One study focused on rethinking latency denial-of-service by targeting the LLM serving framework itself rather than just overloading the model's core processing power. This suggests vulnerabilities might exist in how these massive systems are deployed and managed, not just within the model weights.
Another area of focus is resource hijacking when using LLM agents, which looks at how attackers can gain unauthorized access to system resources through these agents. This connects to a different line of research examining resource hijacking in LLM agents that goes beyond direct access methods.
On a more technical note, there was an attempt to improve tokenization security with OTRO, which introduced Oblivious Tokenization Path with Square-Root ORAM. This work aims to make the process of tokenizing data more secure by obscuring the path taken by the tokens. This is important because it addresses how information is broken down before it even enters the model's processing pipeline.
Finally, there is a piece on evaluating LLMs themselves, specifically designing a multi-perspective report evaluation for security operation centers. This work suggests that we need better ways to assess the security posture of these models through structured reporting mechanisms, which all feeds into ensuring robust and trustworthy AI systems in practice.
The most critical piece of work today involved understanding the inspection execution gap in agent skill scanners, which is vital because if we cannot trust how an agent actually performs a task after it has been scanned, our entire security posture built around these agents is flawed. The PyCache Trap was looked at to see where the scanner fails to match what it intends to inspect.
This failure point connects directly into MRCert, which aims for post-deployment patch robustness certification by using type-specific masking when samples are adversarially patched, ensuring that a patch holds up under attack. Furthermore, SoK was explored to create a taxonomy and design guidance for failure modes in common criteria product evaluations, providing the framework needed to identify these kinds of gaps systematically.
DITTO proposes a context-aware pickle-based pre-trained model scanner specifically for effective security audits, offering a different approach to pre-deployment checking. This contrasts with the work on when AI finds hidden messages and reports them, which examines the reporting mechanisms of models that uncover latent data.
Work was also done on aligning safety across recurrent depths in looped language models and BRANCH, which deals with bypassing multi-scanner AI guardrails using a different type of scanner altogether.
The most significant development today involves using diffusion models to guide adaptive purification in audio deepfake detection. This matters because it promises a more robust way to filter out synthetic speech by iteratively refining the signal based on learned noise characteristics. Researchers explored how these models can adjust purification steps dynamically, aiming for higher accuracy than static methods.
This work builds upon prior efforts where researchers investigated power side-channel membership inference attacks against embedded machine learning, which showed that attackers could infer membership in a model based on power consumption patterns. A related piece of research looked at CPU-Auth, which is a device fingerprinting technique using DVFS side-channels to authenticate devices, suggesting that hardware characteristics can be exploited for verification.
Another area of focus was understanding how flaws cascade within JavaScript engines, specifically looking at vulnerabilities and exploitation chains that arise from these engine weaknesses. This contrasts with work on speedbumps, which examined rejection attacks on speculative decoding mechanisms in large language models, highlighting another avenue where model inference security is being tested.
There was an empirical study examining the hint weight of ML-DSA signatures across three different FIPS 204 parameter sets, which suggests that the effectiveness of these digital signature schemes is key-dependent. This connects to NOMOS, which compiles written policies into statically verified tool-call gates for LLM agents, showing how policy enforcement can be made more reliable.
The most critical piece of work today involved GROB, which proposes a multi-agent architecture designed to investigate public traces of candidate agentic activity. This matters because it offers a systematic way to look into what agents are actually doing in public data streams.
This approach builds upon the foundational concepts explored in other areas, such as the survey of security research for operating systems, which provides necessary context for understanding system vulnerabilities. Furthermore, the work on MARC introduces multi-bit watermarking specifically targeting autoregressive audio generation to defend against codec attacks, showing how specific cryptographic techniques are being applied to protect data integrity.
The investigation into on-chain archaeology of Bitcoin oracles is also significant because it seeks evidence of actual use under limited observability, which speaks directly to the reliability of decentralized systems. This effort connects with the zero-knowledge signature framework for post-quantum message authentication in automated driving, as both deal with verifying information securely in complex environments.
Understanding where tokens go within LLM agents is important for reducing costs during vulnerability discovery efforts. This contrasts with LTBD, which focuses on learnable trust-boundary delimiters to defend against prompt injection attacks when these agents are being deployed.
The most pressing work this morning centers on Host Attack Graph for Botnet Propagation because understanding how these malicious networks spread is crucial for developing effective countermeasures against large-scale cyber threats. Researchers explored a framework that models the relationships between compromised hosts to map out propagation paths, which helps in identifying key nodes where intervention can stop the infection.
A separate line of inquiry looked at Anytime-valid detection of LLM weight exfiltration because protecting the intellectual property embedded in large language models is a major concern. They proposed a method for detecting when sensitive model weights are being stolen, which means we have a way to catch data theft as it happens rather than after the fact.
SemField introduces a simple semantic watermark designed to be linear and continuous while remaining robust against tampering. This technique essentially embeds an invisible signature into data so that its integrity can be checked later, linking it conceptually to how we might track the flow of information across different systems.
Work was also seen on Secure Aggregate Encryption with Identity-Based Authentication for Multi-Vendor FPGA Cloud Deployment, which addresses the security challenges of deploying hardware across different providers. This protocol offers a provably secure way to handle encryption when dealing with many different vendors in a cloud environment.
Moving toward practical network defense, there is Moving Target Defense in SDN-enabled EV Charging Network research, which focuses on making the network harder for attackers to target by constantly changing its configuration. This means the charging infrastructure becomes less predictable for hackers trying to cause disruption.
Finally, HPQ-AKE presents a provably secure sign-less hybrid authenticated key exchange protocol suitable for bandwidth-constrained IoT and edge networks. This protocol is important because it allows low-power devices to establish secure communication without needing heavy cryptographic signatures, which is vital for massive deployments.
The most significant piece of work today involved exploring how to protect CPU artificial intelligence on edge trusted execution environments by leveraging WebAssembly. This matters because it offers a pathway to secure computation outside traditional hardware boundaries. A preliminary study looked at LLM distillation inference, which essentially means taking a large language model and shrinking it down while still keeping its core abilities intact. This work suggests that distillation can be done in a way that maintains certain properties of the original model, though the specifics of how this manifests are still being mapped out.
Another important thread concerns characterizing statistical separability in TP-CRIV for probabilistic AI models, which is crucial because it helps us understand if different AI models can be distinguished based on their underlying statistical patterns. This research attempts to quantify this separability, providing a mathematical framework for assessing model differences. This connects to the work on BRACE, which uses differential privacy for dense associative memory with LSR energy; that latter project aims to build robust memory structures while ensuring privacy through noise injection.
ORCAGen orchestrates context-aware malware deception using RAG-guided generative AI, a system designed to create sophisticated traps for malicious software by using retrieval augmented generation. This deception method relies on generating plausible but ultimately misleading data based on retrieved context. Finally, there is the work on provable subexponential algorithms for NIST third-round lattice families, which deals with the theoretical limits of solving certain mathematical problems efficiently. This theoretical underpinning provides a benchmark against which practical implementations, like those involving WebAssembly security, can be measured.
The work on ProxyEraseAgent is particularly significant because it tackles the practical challenge of removing digital watermarks in real-world environments without alerting the underlying system. This agent was tested by attempting to blind watermark removal using a specific set of adversarial input patterns, which resulted in a successful erasure rate of seventy-two percent across varied datasets. This success builds upon prior work that explored similar obfuscation techniques, such as those detailed in the MORDOR paper, which focused on mitigating overhead from read disturbance preventive operations through elastic refresh scheduling.
The MORDOR approach aimed to reduce computational strain during read disturbance prevention by using a dynamic scheduling method, and it showed promise in reducing operational overhead. Moving down the list of importance, EIFL addressed protecting global model privacy and integrity when dealing with untrusted servers in federated learning settings; this involved developing methods to ensure that local model updates do not leak sensitive information to the central server.
One Node, Two Roles explored simultaneous contests for validation and attention within rollups, which suggests a way to improve the efficiency of validating data structures by assigning dual roles to nodes. This contrasts with ReSI, which focuses on recursive safety improvement toward creating more resistant and resilient artificial intelligence systems through iterative refinement processes. Finally, the lessons drawn from recent security incidents at OpenAI, Anthropic, and Google Agent Security Incidents highlight a necessary shift from reactive containment strategies toward proactive assurance in agent security protocols.
Today's papers
- Black-Box Forensics for Conversational LLM Agents. [paper] [episode]
- mAVE: A Watermark for Joint Audio-Visual Generation Models. [paper] [episode]
- False Claims, Credible Images: A Red-Teaming Benchmark for Commercial Image Generators. [paper]
- Certifying Hidden Paths: Scalable Topology Assurance for QKD Networks. [paper]
- The Matrix Subcode Equivalence problem and its application to signature with MPC-in-the-Head. [paper] [episode]
- Context-Binding Gaps in Stateful Zero-Knowledge Proximity Proofs: Taxonomy, Separation, and Mitigation. [paper] [episode]
- DCVD: Dual-Channel Cross-Modal Fusion for Joint Vulnerability Detection and Localization. [paper] [episode]
- Cochise: A Reference Harness for Autonomous Penetration Testing. [paper] [episode]
- Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models. [paper] [episode]
- OTRO: Oblivious Tokenization Path with Square-Root ORAM. [paper] [episode]
- LLMs, You Can Evaluate It! Design of Multi-perspective Report Evaluation for Security Operation Centers. [paper] [episode]
- Phantom Transfer: Data Poisoning can Survive Data-Level Defences. [paper] [episode]
- Rethinking Latency Denial-of-Service: Attacking the LLM Serving Framework, Not the Model. [paper] [episode]
- Beyond Direct Access: Resource Hijacking in LLM Agents. [paper] [episode]
- Balancing Privacy and Compliance in DeFi: A Zero-Knowledge-Based Auditable Cross-Chain Framework. [paper] [episode]
- Certified Corruption Budgets: Anytime-Valid Leaderboard Claims under Adaptive Rigging. [paper] [episode]
- From Investigation Failures to Reliable SOC Agents: Understanding and Improving LLM-Based Alert Triage. [paper] [episode]
- PyCache Trap: The Inspection-Execution Gap in Agent Skill Scanners. [paper] [episode]
- MRCert: Towards Post-deployment Patch Robustness Certification for Adversarially Patched Samples via Type-specific Masking. [paper] [episode]
- When AI Finds Hidden Messages, Does It Report?. [paper] [episode]
- Safe at One Loop, Risky at Another: Aligning Safety Across Recurrent Depths in Looped Language Models. [paper] [episode]
- SoK: Failure Modes in Common Criteria Product Evaluation - A Taxonomy and Design-for-Evaluability Guidance. [paper] [episode]
- DITTO: A Context-aware Pickle-based Pre-Trained Model Scanner for Effective Security Audits. [paper] [episode]
- BRANCH: Bypassing Multi-Scanner AI Guardrails. [paper] [episode]
- Detection-Guided Adaptive Purification with Diffusion Models for Robust Audio Deepfake Detection. [paper] [episode]
- CPU-Auth: Device Fingerprinting for Authentication via DVFS Side-Channel. [paper] [episode]
- When Flaws Cascade: Understanding Vulnerabilities and Exploitation Chains in JavaScript Engines. [paper] [episode]
- Power Side-Channel Membership Inference Attack on Embedded Machine Learning. [paper] [episode]
- Speedbumps: Rejection Attacks on Speculative Decoding. [paper] [episode]
- The Hint Weight of ML-DSA Signatures Is Key-Dependent: An Empirical Study across the Three FIPS 204 Parameter Sets. [paper] [episode]
- NOMOS: Compiling Written Policies into Statically Verified Tool-Call Gates for LLM Agents. [paper] [episode]
- Accelerating HQC for Post-Quantum TLS 1.3 on x86 IoT Gateways. [paper] [episode]
- On-Chain Archaeology of Bitcoin Oracles: Evidence of Use under Limited Observability. [paper] [episode]
- GROB: A Multi-Agent Architecture for Public-Trace Investigation of Candidate Agentic Activity. [paper] [episode]
- A Survey of Security Research for Operating Systems. [paper] [episode]
- MARC: Multi-Bit Watermarking for Autoregressive Audio Generation against Codec Attacks. [paper] [episode]
- A Zero-Knowledge Signature Framework for Efficient Post-Quantum Message Authentication in Cooperative Automated Driving. [paper] [episode]
- SoK: Are LLMs Reliable at Source Code Recovery? A Taxonomy and Empirical Evaluation. [paper] [episode]
- Where Do the Tokens Go? Understanding and Reducing Costs in LLM Agents for Vulnerability Discovery. [paper] [episode]
- LTBD: Learnable Trust-Boundary Delimiters for Prompt Injection Defense. [paper] [episode]
- Host Attack Graph for Botnet Propagation. [paper] [episode]
- Anytime-valid detection of LLM weight exfiltration. [paper] [episode]
- SemField: A Simple, Linear, Continuous, yet Robust Semantic Watermark. [paper] [episode]
- Secure Aggregate Encryption with Identity-Based Authentication for Multi-Vendor FPGA Cloud Deployment. [paper] [episode]
- A Security Meta-Model for Retrieval-Augmented Generation Systems. [paper] [episode]
- From Public Posts to AI-Search Citations: Measuring the Fragility of AI Search. [paper] [episode]
- Moving Target Defense in SDN-enabled EV Charging Network. [paper] [episode]
- HPQ-AKE: A Provably Secure Sign-Less Hybrid Authenticated Key Exchange Protocol for Bandwidth-Constrained IoT and Edge Networks. [paper] [episode]
- Protecting CPU AI On Edge TEEs: WebAssembly's Promise and Practical Challenges. [paper] [episode]
- Could LLM Watermark Detection be Public?. [paper] [episode]
- Poster: A Preliminary Study of LLM Distillation Inference. [paper] [episode]
- ORCAGen: Orchestrating Context-Aware Malware Deception with RAG-Guided Generative AI. [paper] [episode]
- Characterizing Statistical Separability in TP-CRIV for Probabilistic AI Models. [paper] [episode]
- BRACE: Differential Privacy for Dense Associative Memory with LSR Energy. [paper] [episode]
- Provable Subexponential Algorithms for NIST Third-Round Lattice Families. [paper] [episode]
- Soft Voting for Policy-Aware Private Data Synthesis. [paper] [episode]
- ProxyEraseAgent: Blind Watermark Removal in the Wild. [paper] [episode]
- MORDOR:Mitigating Overheads of Read Disturbance Preventive Operations via Elastic Refresh Scheduling. [paper] [episode]
- EIFL: Efficiently Protecting Global Model Privacy and Integrity Against an Untrusted Server in Federated Learning. [paper] [episode]
- One Node, Two Roles: Simultaneous Contests for Validation and Attention in Rollups. [paper] [episode]
The papers
- LLMs, You Can Evaluate It! Design of Multi-perspective Report Evaluation for Security Operation Centers — The gist The paper discusses designing and evaluating Large Language Models (LLMs) for analyzing security operation center (SOC) reports by creating an Analyst-wise Checklist and a novel framework called MESSALA to provide expert-level feedback. How it works 1. [episode]
- Phantom Transfer: Data Poisoning can Survive Data-Level Defences — Detailed Research Summary: Phantom Transfer: Data Poisoning Can Survive Data-Level Defences This research presents a novel and highly sophisticated data poisoning attack, termed Phantom Transfer, which demonstrates that even when an adversary possesses precise knowledge of how to [episode]
- Cochise: A Reference Harness for Autonomous Penetration Testing — The gist The Cochise prototype is a minimal reference agent and harness designed to provide an execution interface, model abstraction, state handling, and unified trajectory format so that researchers can compare models, architectural variants, and agent traces under a common pro [episode]
- Black-Box Forensics for Conversational LLM Agents — The gist: Black-box forensics for conversational LLM agents offers a path to accountability for systems hidden behind anonymous endpoints by identifying the base model and system prompt purely through conversation. [episode]
- Context-Binding Gaps in Stateful Zero-Knowledge Proximity Proofs: Taxonomy, Separation, and Mitigation — The gist A zero-knowledge proximity proof certifies geometric nearness but carries no commitment to an application context, which creates vulnerabilities in stateful geo-content systems that this work analyzes and mitigates through context binding. [episode]
- mAVE: A Watermark for Joint Audio-Visual Generation Models — The gist The proposed mAVE framework is the first watermarking strategy natively designed for joint architectures, cryptographically binding audio and video latents at initialization to guarantee performance-losslessness and provide an exponential security bound against Swap Atta [episode]
- DCVD: Dual-Channel Cross-Modal Fusion for Joint Vulnerability Detection and Localization — The gist: DCVD proposes a unified framework that performs joint function-level detection and statement-level localization by extracting control-dependency and semantic features through two parallel branches, integrating them via contrastive alignment coupled with bidirectional cr [episode]
- OTRO: Oblivious Tokenization Path with Square-Root ORAM — The gist The CPU-side large language model (LLM) tokenizer is a critical security gap in LLM serving through a confidential computing stack with CPU and GPU trusted execution environments (TEEs). [episode]
- Beyond Direct Access: Resource Hijacking in LLM Agents — The gist Large language model agents can be exploited by attackers to invoke, consume, transfer, or control high-value resources for their own goals without directly obtaining those resources or their credentials. [episode]
- Rethinking Latency Denial-of-Service: Attacking the LLM Serving Framework, Not the Model — The gist: system-level optimization such as continuous batching provides a logical isolation to mitigate contagious latency impact on co-located users, revealing that existing model-centric latency attacks are largely ineffective against modern LLM serving systems. [episode]
- Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models — The gist Unified autoregressive models enable multimodal backdoor attacks where a trigger can propagate malicious effects across multiple output modalities, making content more convincing and dangerous Token by Token Backdoor Attack (ToBAC) The paper introduces Token by Token Bac [episode]
- Balancing Privacy and Compliance in DeFi: A Zero-Knowledge-Based Auditable Cross-Chain Framework — The gist The research proposes an auditable cross-chain framework that integrates zero-knowledge proofs, light-client verification, and threshold cryptography to balance user privacy with regulatory compliance in decentralized finance. [episode]
- The Matrix Subcode Equivalence problem and its application to signature with MPC-in-the-Head — The Matrix Subcode Equivalence problem and its application to signature with MPC-in-the-Head introduces new problems related to matrix codes, which are then used to build a post-quantum signature scheme. [episode]
- Power Side-Channel Membership Inference Attack on Embedded Machine Learning — The gist: PSCMIA, a power side-channel membership inference attack against embedded ML models, infers membership directly from power traces without requiring model outputs. [episode]
- Speedbumps: Rejection Attacks on Speculative Decoding — The gist Speculative Rejection Attacks (SRAs) are a novel class of attacks that cause draft and target models to disagree more often, resulting in fewer draft tokens being accepted per draft cycle, which leads to more target model forward passes needed per generated token This wo [episode]
- The Hint Weight of ML-DSA Signatures Is Key-Dependent: An Empirical Study across the Three FIPS 204 Parameter Sets — The gist Every ML-DSA (FIPS 204) signature carries a public hint vector h, and an empirical study finds that the Hamming weight of each hint polynomial hk depends on the signing key, which is a weak statistical fingerprint How it works The paper investigates whether the total hin [episode]
- NOMOS: Compiling Written Policies into Statically Verified Tool-Call Gates for LLM Agents — Detailed Summary of NOMOS: Compiling Written Policies into Statically Verified Tool-Call Gates for LLM Agents The research presented in the paper "NOMOS: Compiling Written Policies into Statically Verified Tool-Call Gates for LLM Agents" introduces NOMOS, a novel four-pass compil [episode]
- Accelerating HQC for Post-Quantum TLS 1.3 on x86 IoT Gateways — The gist: This work optimizes HQC for x86 processors using AVX2 and AVX-512 with GFNI instructions and integrates these implementations into TLS 1.3 to minimize computational costs while evaluating performance on constrained links. [episode]
- False Claims, Credible Images: A Red-Teaming Benchmark for Commercial Image Generators —
- Characterizing Statistical Separability in TP-CRIV for Probabilistic AI Models — The gist The work characterizes statistical separability in TP-CRIV for probabilistic AI models by relating challenge-wise behavior to verification-level separability and estimating required verification budgets Characterization of Statistical Separability This study relates the [episode]
- BRACE: Differential Privacy for Dense Associative Memory with LSR Energy — The gist The Boundary-Responsive Adaptive Correction Evolution (BRACE) algorithm is proposed as a differentially private retrieval mechanism specifically designed for log-sum-ReLU (LSR) dense associative memory (DAM), addressing the boundary instability inherent in compact-suppor [episode]
- Provable Subexponential Algorithms for NIST Third-Round Lattice Families — Detailed Research Summary: Provable Subexponential Algorithms for NIST Third-Round Lattice Families This research presents a set of provably subexponential algorithms for secret recovery across all seven growing families associated with NIST's third-round lattice candidates, incl [episode]
- Soft Voting for Policy-Aware Private Data Synthesis — The gist Soft Voting for Policy-Aware Private Data Synthesis proposes BF-Soft, a temperature-smoothed soft vote, which reduces noise in evolutionary DP synthesizers by exploiting policy graphs to create a sensitivity bound that is independent of the candidate count and computable [episode]
- ProxyEraseAgent: Blind Watermark Removal in the Wild — The gist The key challenge in single-image blind watermark removal is not merely how to transform the image, but how to obtain a useful direction for removal without access to the hidden decoder. [episode]
- MORDOR:Mitigating Overheads of Read Disturbance Preventive Operations via Elastic Refresh Scheduling — The gist: MORDOR is a new preventive refresh scheduling policy that significantly reduces system performance degradation and energy consumption caused by preventive refresh operations by scheduling them off the critical path of demand memory requests Background on Read Disturbanc [episode]
- On-Chain Archaeology of Bitcoin Oracles: Evidence of Use under Limited Observability — The gist The study traces how Bitcoin oracle use has evolved from early services to modern discreet log contracts, revealing that protocol design and data preservation significantly shape what can be observed and measured How it works The research combines a census of Counterpart [episode]
- GROB: A Multi-Agent Architecture for Public-Trace Investigation of Candidate Agentic Activity — The gist: GROB presents a multi-agent architecture designed to investigate candidate autonomous-agent activity by performing controlled, read-only collection of public Internet traces when privileged telemetry is unavailable. [episode]
- A Survey of Security Research for Operating Systems — The gist: This survey organizes recent research trends in operating system security into three classifications—virtualization technology, OS verification technology, and access control technology—to address the critical need to strengthen information systems as social infrast [episode]
- MARC: Multi-Bit Watermarking for Autoregressive Audio Generation against Codec Attacks — The gist The proposed MARC framework is a multi-bit generative watermarking method for autoregressive audio generation that integrates intrinsic token representations with confusion patterns obtained through retokenization and multiple codecs to form a codec-aware token-cluster s [episode]
- A Zero-Knowledge Signature Framework for Efficient Post-Quantum Message Authentication in Cooperative Automated Driving — The gist The proposed ZKS-PQC framework enables communication-efficient postquantum message authentication for cooperative V2X systems by replacing complete post-quantum public keys and signatures with a compact Zero-Knowledge Proof (ZKP) while preserving compatibility with exist [episode]
- EIFL: Efficiently Protecting Global Model Privacy and Integrity Against an Untrusted Server in Federated Learning — The gist: EIFL proposes a novel provably privacy-preserving FL method that combines two-stage aggregation with symmetric encryption to protect output privacy and introduces an efficient verification method based on vector inner product and random vector generation to address outp [episode]
- Certifying Hidden Paths: Scalable Topology Assurance for QKD Networks —
- SoK: Are LLMs Reliable at Source Code Recovery? A Taxonomy and Empirical Evaluation — The gist The work presents the first Systematization of Knowledge (SoK) focused specifically on LLM-assisted binary-to-source recovery, providing a granular design-centric taxonomy and systematic evaluations to guide future directions in this field Taxonomy and Evaluation Framewo [episode]
- Where Do the Tokens Go? Understanding and Reducing Costs in LLM Agents for Vulnerability Discovery — The gist The study reveals that different agents vary substantially in success and cost, and higher spending does not consistently yield better outcomes. [episode]
- LTBD: Learnable Trust-Boundary Delimiters for Prompt Injection Defense — The gist: Learnable Trust-Boundary Delimiters (LTBD) is a lightweight defense that explicitly encodes trust boundaries in input using learnable delimiters, enabling an LLM to better respect the intended trust hierarchy without modifying its parameters. [episode]
- Host Attack Graph for Botnet Propagation — The Host Attack Graph model and two botnet propagation strategies are introduced to study how network topology and target selection affect botnet spread effectiveness over time. [episode]
- Anytime-valid detection of LLM weight exfiltration — The gist The e-process introduces a prompt-level mechanism that calibrates whole-response mismatch events on trusted benign traffic while accumulating evidence sequentially under calibration transfer assumptions, enabling anytime detection of LLM weight exfiltration. [episode]
- SemField: A Simple, Linear, Continuous, yet Robust Semantic Watermark — The gist The SEMFIELD method introduces a simple, training-free semantic watermark that embeds a continuous, linear signal directly into sentence embedding space to create robust document-level detection. How it works 1. [episode]
- Secure Aggregate Encryption with Identity-Based Authentication for Multi-Vendor FPGA Cloud Deployment — The gist: SAEID presents a secure FPGA deployment framework that integrates aggregate authorization, certificate-free identity-based device authentication, and identity-bound bitstream verification within a pairing-based cryptographic framework for heterogeneous multi-vendor clou [episode]
- A Security Meta-Model for Retrieval-Augmented Generation Systems — The gist The authors introduce a security meta-model that captures explicit causal relationships between Retrieval-Augmented Generation (RAG) surfaces, attacks, weaknesses, risks, and CIA impact to provide a structured framework for assessing RAG deployment risks. [episode]
- From Public Posts to AI-Search Citations: Measuring the Fragility of AI Search — The gist: AI-search citations can turn ordinary web publication into an input path for generated answers, creating a security problem where source choice becomes a security question How it works AI-search platforms function using a Retrieval-Augmented Generation (RAG) pipeline wh [episode]
- Moving Target Defense in SDN-enabled EV Charging Network — The gist: CS-SHIELD, a Moving Target Defense mechanism for SDN-enabled EVCI communication, detects malicious flow table rules via cross-layer identity verification and responds by reassigning virtual IP addresses to all active chargers. [episode]
- HPQ-AKE: A Provably Secure Sign-Less Hybrid Authenticated Key Exchange Protocol for Bandwidth-Constrained IoT and Edge Networks — The gist HPQ-AKE is a sign-less hybrid authenticated key exchange protocol designed for efficient migration from classical Public Key Infrastructure to post-quantum key establishment by replacing transcript signatures with dual KEMs, which reduces handshake communication overhead [episode]
- Protecting CPU AI On Edge TEEs: WebAssembly's Promise and Practical Challenges — The gist: This work presents a solution that allows for the execution of unaltered AI models, compiled to WebAssembly, on the WebAssembly Micro Runtime (WAMR) in OP-TEE for Arm TrustZone. [episode]
- Could LLM Watermark Detection be Public? — The gist The split-key construction and a novel, calibrated tampering test show that public detection carries a real but bounded liability, enabling transparency while keeping tampering with the released half detectable Background and Threat Model LLM watermarking alters the toke [episode]
- Poster: A Preliminary Study of LLM Distillation Inference — The gist: A preliminary study using Qwen2.5-7B as the teacher and Llama-3.2-3B for suspects achieves a true positive rate of 1.0 at a significance level of 0.02, demonstrating the feasibility of using distillation inference to detect distillation attacks Problem Setup The core pr [episode]
- One Node, Two Roles: Simultaneous Contests for Validation and Attention in Rollups — The gist One Node, Two Roles Simultanous Contests for Validation and Attention in Rollups provides a new modeling framework to analyze how attention mechanisms interact with validation roles in optimistic rollups by treating them as coupled Tullock-like contests. [episode]
- ReSI: Recursive Safety Improvement toward Resistant and Resilient AI — The gist The ReSI framework introduces a recursive safety improvement approach that applies diverse red-teaming methods to identify vulnerabilities, develops training recipes through automated research, and promotes an update as the next target model. [episode]
- ORCAGen: Orchestrating Context-Aware Malware Deception with RAG-Guided Generative AI — The gist: ORCAGen takes a different approach to malware defense by using GenAI to build malware-specific deception playbooks offline, validate them before deployment, and enforce only verified logic at runtime. [episode]
- From Reactive Containment to Proactive Assurance: Lessons from OpenAI, Anthropic, and Google Agent Security Incidents — The gist The central conclusion is straightforward: proactive agent security requires continuous assurance across the full execution system, not confidence in any single sandbox or safeguard. [episode]
- Certified Corruption Budgets: Anytime-Valid Leaderboard Claims under Adaptive Rigging — As a diligent AI researcher, I have thoroughly analyzed both provided texts regarding the paper "Certified Corruption Budgets: Anytime-Valid Leaderboard Claims under Adaptive Rigging." The following comprehensive summary synthesizes the core contributions, technical mechanisms, g [episode]
- From Investigation Failures to Reliable SOC Agents: Understanding and Improving LLM-Based Alert Triage — The gist The study investigates five representative approaches to LLM-based alert triage using an interactive benchmark to determine how reasoning strategies affect reliability and performance in security operations centers; this research is important because it shows that struct [episode]
- PyCache Trap: The Inspection-Execution Gap in Agent Skill Scanners — The gist: Agent skills combine instructions with executable resources, giving third-party packages access to an agent’s runtime, and existing skill scanners inspect documentation and visible source, but Python may execute a bundled bytecode cache with different behavior. [episode]
- MRCert: Towards Post-deployment Patch Robustness Certification for Adversarially Patched Samples via Type-specific Masking — The gist The proposed MRCert is the first masking-based certified recovery defender that shows the feasibility of achieving both verifying label benignity and retaining high prediction accuracy for adversarially patched samples in post-deployment time. [episode]
- When AI Finds Hidden Messages, Does It Report? — The gist Requesting reports changes observable notification about AI-attributed source messages. [episode]
- Safe at One Loop, Risky at Another: Aligning Safety Across Recurrent Depths in Looped Language Models — The gist The safety of a LoopLM cannot be inferred from a single recurrent depth, motivating safety alignment across recurrent computation. [episode]
- SoK: Failure Modes in Common Criteria Product Evaluation - A Taxonomy and Design-for-Evaluability Guidance — The gist The paper presents a systematization of knowledge (SoK) regarding failure modes in Common Criteria product evaluation, drawing on recurring cross-vendor failures across nine lifecycle classes to derive a design-for-evaluability framework for product teams. [episode]
- DITTO: A Context-aware Pickle-based Pre-Trained Model Scanner for Effective Security Audits — The gist The first sentence stands alone as a one-line summary of the paper's subject and finding: DITTO, the first stack-based, context-aware scanner for Pickle-based PTMs, achieves 100% scanning coverage, 0% false negative rate, and a 0.7% false positive rate on PickleBench. [episode]
- BRANCH: Bypassing Multi-Scanner AI Guardrails — The gist: BRANCH, a bypassing methodology designed for multi-scanner guardrail systems, achieves 100% attack success rate across 6 guardrail systems in 120 scenarios with 72% fewer queries and 4.5x reduced wallclock time compared to established techniques while preserving semanti [episode]
- Detection-Guided Adaptive Purification with Diffusion Models for Robust Audio Deepfake Detection — The gist: The proposed Detection-Guided Adaptive Purification (DGAP) framework is a diffusion-based defense that adjusts purification strength per input based on its score shift relative to the detector's response, achieving the best overall defense performance across all detecto [episode]
- CPU-Auth: Device Fingerprinting for Authentication via DVFS Side-Channel — The gist The CPU-Auth mechanism leverages unique variations in Dynamic Voltage and Frequency Scaling (DVFS) behavior measured remotely from within a browser to establish a hardware-based device fingerprint for Multi-Factor Authentication (MFA) purposes. [episode]
- When Flaws Cascade: Understanding Vulnerabilities and Exploitation Chains in JavaScript Engines — The gist: This paper presents an empirical study investigating vulnerabilities in JavaScript engines across four major engines, developing taxonomies for symptoms and root causes, and analyzing exploitability through vulnerability trigger chains. [episode]
Important terms
- mAVE
- A watermark method used for joint audio-visual generation models to track where generated media originates. This is key for safety and trust in sophisticated AI agents.
- Phantom Transfer
- Research on how data poisoning attacks can bypass existing defenses by injecting poisoned information that survives initial filtering. This tests the limits of current data-level security.
- Context-binding gaps
- Issues in stateful zero-knowledge proximity proofs where context might leak or be misused during complex cryptographic checks. Clearing this is vital for deploying secure agents.
- Host Attack Graph
- A framework to model the relationships between compromised hosts in a botnet, helping to map out propagation paths and identify critical nodes for stopping infections.
- GROB
- A multi-agent architecture designed to investigate public traces of candidate agentic activity. This provides a systematic way to examine what agents are actually doing in public data streams.