Security papers — 2026-10-02
Attackers are bypassing machine learning classifiers through adversarial noise, which matters because if we cannot trust these models against subtle manipulation, the security of systems relying on them is fundamentally compromised. Researchers explored UnifiedAttack, which evaluates the safety of large multimodal models in generating harmful image-text combinations and shows how these models can be exploited synergistically to create risks.
Then there was work on Tokenized Key-Gated Adapter Routing, which introduced a secure access control mechanism designed to prevent private data leakage within large language models by managing how certain adapters are routed. This is more specific than the general noise attacks because it addresses internal data flow security.
We also looked at ReCast, which focuses on contract-preserving protection for fixed-interface multimodal reasoning, aiming to ensure that model outputs adhere strictly to predefined interfaces during complex reasoning tasks. This builds upon the idea of controlling model behavior in structured environments.
OverAct investigated measuring and mitigating proactive over-authorization in LLM tool-calling agents, which is important because these agents can sometimes grant themselves more permissions than necessary during execution. This directly relates to ensuring agent actions are appropriately scoped.
Finally, we reviewed a comprehensive taxonomy of one-pixel attacks, which provides a broad overview of the research status and regulatory landscape surrounding these subtle input manipulations. This review sets the context for understanding the broader threat space we discussed earlier.
The most significant piece of work today involves TensorCommitments, which presents a lightweight way to verify inference for language models. This is important because it offers a method for ensuring that the outputs from these large models are trustworthy without requiring massive computational overhead.
This approach uses tensor commitments to provide verifiable inference, meaning we can check if the model's output is consistent with its training in an efficient manner. This builds upon previous work by HarnessAgent, which focuses on scaling automatic fuzzing harness construction using tool-augmented large language model pipelines to find vulnerabilities.
Another area of focus is the development of CausalArmor, which creates efficient indirect prompt injection guardrails through causal attribution. This means it tries to stop malicious inputs from tricking models by tracing the cause of the input's effect, and this ties into rethinking anonymity claims in synthetic data generation from a model-centric privacy attack perspective.
Finally, there is PSR2, a phase-based semantic reasoning framework designed for detecting atomicity violations via contract refinement. This work is crucial because it helps identify when complex processes fail to complete correctly by refining the underlying contracts, which relates back to the broader context of security and reliability discussed in GNSS spoofing surveys.
The most pressing work today involves the TESLA for 5G broadcast authentication, because it directly addresses security in the next generation of mobile networks. Researchers explored how to use this technique to authenticate devices on 5G networks, and they found that a specific method allowed them to achieve a certain level of security while maintaining reasonable performance metrics. This is significant because it provides a concrete pathway for securing 5G infrastructure against unauthorized access attempts in real-time.
Building upon this, there was work on diagnosing issues within closed-loop agent debugging, which is important because it helps us understand how complex automated systems behave when things go wrong. They investigated how a verifier can inadvertently leak the answer during this process, showing that this leakage happens before any serious optimization efforts are applied. This finding suggests we need to be careful about what information these diagnostic tools reveal while they are still in development.
Another area of focus was on ensuring the reliability of persistent AI agents through a cognitive continuity test, which matters because it verifies that these agents maintain their intended state transitions over time. The results showed how this test can confirm whether the agent is actually behaving as designed when it needs to switch between different operational modes. This verification work complements the security concerns by ensuring that autonomous systems remain trustworthy in their long-term operation.
The most significant piece of work from yesterday was the development of a resource-aware behavior reconstruction framework for host intrusion detection, because understanding how systems behave under duress is crucial for building robust security. This framework attempts to model system actions by considering available resources, which helps in spotting anomalies that might otherwise be missed.
A related effort explored autonomous open-source software threat detection using taxonomy-aligned large language models, aiming to automatically classify malicious code based on established threat categories. This work suggests that LLMs can be leveraged for proactive security monitoring.
Then there was the investigation into evidence coverage for intent-bound execution, which looks at how well a system can reason about its scope and obligations when executing specific tasks. This is important because it moves beyond simple pattern matching to understand the underlying purpose of an action.
Further down the line, research on false floors in LLM safety routing evaluations showed that these evaluations break down under distribution shift, meaning the safety checks fail when the input data changes unexpectedly. This points to a fragility in current methods for ensuring agent safety.
Finally, there was work on chaining skills to hijack LLM agents and protocol integration of physical layer deception into EAP-TEAP Wi-Fi authentication, which deals with more complex adversarial techniques and system-level security vulnerabilities.
The work concerning authorization for self modifying AI agent populations is particularly important because it addresses the fundamental challenge of maintaining control when these agents can alter their own code or structure. This research explored how to conserve authority across replacement, forking, and rollback scenarios within these agent populations.
A related effort focused on removing backdoors in large language models through weight orthogonalisation, which attempts to eliminate hidden vulnerabilities within the model's parameters. This is significant because it directly targets the integrity of the foundational models themselves.
Another piece of research investigated proof-gated signing for onchain AI agents, creating solver-checked transaction guards that remain robust even when state drift occurs. This provides a layer of security for agents operating in decentralized environments.
The study on intrusion detection for agentic processes at runtime is crucial because it monitors how these agents behave while they are actively running, providing evidence-based monitoring. This work builds upon the idea of detecting malicious activity during execution.
Finally, the research into the relationship between model quantization and model inversion attacks examines how reducing a model's size affects its susceptibility to revealing sensitive training data. This connects to broader concerns about protecting proprietary information embedded within these models.
The most significant development today involves the work on identity-bound governance under execution uncertainty, which matters because it addresses accountability when large language model agents might unexpectedly halt during operation. This research introduced a cryptographic implementation and cross-model calibration to provide this proof block.
This builds upon the planning and execution framework explored in towards hierarchical cyber defense with large language models, which deals with how these agents plan their actions before they execute them. A related piece of work focused on crossing the cyber divide by examining sim-to-sim and sim-to-real transfer for reinforcement learning agents, suggesting ways to bridge the gap between simulated and real-world agent performance.
Another important area is progressive resolution secure aggregation for federated learning, which tackles how to securely combine models trained across different environments without exposing sensitive data. This contrasts with the work on momat, which focuses on low-power jailbreak defense for quantized large language models by using a mixture of multiple atlases.
The most pressing work today concerns the Sleeping Secrets of fine-tuning, which reveals how reawakening privacy risks in language models can be exploited. This is significant because it shows that simply fine-tuning a model on specific data does not guarantee safety; instead, it opens up new avenues for unintended information leakage.
This risk is compounded by the High-quality Data Do not Mean Safe! paper, which demonstrates how poisoning LLMs after data selection can introduce malicious behavior. This means that even if the initial training set is curated carefully, subsequent contamination during model refinement can compromise the system's integrity.
Moving down in importance but still crucial is SoK: Decentralized Agent Economic Infrastructure, which proposes a framework for decentralized agent economic infrastructure. This work attempts to solve problems related to how agents interact economically without a central authority.
Then there is PACE, which focuses on Provenance-Aware Capability Enforcement for Tool-Using LLM Agents. This research tries to ensure that when an agent uses external tools, its capabilities are strictly enforced based on where that tool's information originated.
Finally, the work on Key-Reuse Vulnerability of Phase-Keyed Fourier-Curve Modulation highlights a key vulnerability in coded links related to relation leakage and the cost of key refreshment. This is a technical finding about cryptographic systems that could impact secure communication channels.
The work on system-level optimization beyond cryptographic kernels in the Arm Cortex M7 is particularly important because it directly impacts the efficiency of embedded security systems. This research explored how machine learning can be used to optimize operations beyond just the cryptographic functions themselves. They investigated this by applying an ML-KEM case study to see how performance could be improved on this specific processor.
This optimization work builds upon other efforts in multimodal retrieval, specifically looking at datastore extraction from Retrieval Augmented Generation systems. That research tried to figure out how to pull relevant data out of a database when using multimodal RAG setups. It suggests that understanding the underlying data structures is key for effective retrieval.
Another piece of work focused on a structured state space sequence model for multi-class classification of malware, which attempts to categorize malicious software based on its internal patterns. This classification approach is significant because it moves beyond simple signature matching by looking at the sequence of operations within the code.
The hybrid approach to malware detection, which integrates few-shot model-agnostic meta-learning with autoencoders, also contributes to this field. This method aims to build robust detection systems that can learn new threats quickly even when trained on very little data.
Finally, there was work on detection and resolution of periodic artifacts in OpenDP's discrete Laplace sampler. This effort addresses issues with timing or repeating patterns within a specific sampling mechanism used in some systems.
Today's papers
- Evasion Attacks: How Adversarial Noise Bypasses ML Classifiers Evasion attacks show how small changes to input can fool machine learning classifiers. [paper] [episode]
- Tokenized Key-Gated Adapter Routing: A Secure Access Control Mechanism Against Private Data Leakage in LLMs This method uses tokenized keys to securely control access to private data within large language models. [paper]
- Actions with Receipts: Jointly Binding Claims, Evidence, and Execution for Replayable Tool-Agent Auditing This system links actions, evidence, and execution together so that agent tool usage can be fully audited. [paper]
- AuraForge: Scaling Security Supervision for Training Coding Agents This framework scales security supervision to help train coding agents more effectively. [paper]
- ReCast: Contract-Preserving Protection for Fixed-Interface Multimodal Reasoning ReCast protects multimodal reasoning by preserving the original contract during fixed-interface operations. [paper]
- OverAct: Measuring and Mitigating Proactive Over-Authorization in LLM Tool-Calling Agents OverAct measures how much proactive authorization an LLM tool agent has and suggests ways to reduce it. [paper]
- UnifiedAttack: Evaluating the Safety of Large Multimodal Models in Synergistic Harmful Image-Text Generation UnifiedAttack evaluates the safety risks of large multimodal models when they generate harmful content from images and text together. [paper]
- A Comprehensive Review of One-Pixel Attack: Research Status, Taxonomy, Applications, Regulation Policy and Future Directions This paper reviews all aspects of one-pixel attacks on machine learning models. [paper] [episode]
- Is it Possible to Generate Irreversible PolyProtected Templates from Face Embeddings using System-Specific Keys This research investigates if system-specific keys can be used to create irreversible protected templates from face embeddings. [paper]
- GNSS Spoofing in Mobile Devices: A Survey on Impact and Countermeasures This survey reviews the impact of GPS spoofing attacks on mobile devices and how to counter them. [paper] [episode]
- Federated Detection of Open Charge Point Protocol 1.6 Cyberattacks This paper describes a method for detecting cyberattacks targeting open charge point protocol 1.6 using federated learning. [paper] [episode]
- HarnessAgent: Scaling Automatic Fuzzing Harness Construction with Tool-Augmented LLM Pipelines HarnessAgent scales the creation of fuzzing harnesses by using tool-augmented language model pipelines.
- TensorCommitments: A Lightweight Verifiable Inference for Language Models TensorCommitments provides a lightweight way to verify inferences made by language models. [paper] [episode]
- PSR2: A Phase-based Semantic Reasoning Framework for Atomicity Violation Detection via Contract Refinement PSR2 uses phase-based reasoning and contract refinement to detect violations of atomicity. [paper] [episode]
- Rethinking Anonymity Claims in Synthetic Data Generation: A Model-Centric Privacy Attack Perspective This paper examines privacy risks in synthetic data generation from a model's perspective. [paper] [episode]
- CausalArmor: Efficient Indirect Prompt Injection Guardrails via Causal Attribution CausalArmor creates effective indirect prompt injection guardrails by using causal attribution. [paper] [episode]
- TESLA-for-5G: Broadcast Authentication for 5G Networks Using TESLA TESLA-for-5G uses broadcast authentication to secure 5G networks. [paper] [episode]
- A Verifier Can Leak the Answer: Diagnosability Before Optimization in Closed-Loop Agent Debugging This work explores how a verifier can leak answers during closed-loop agent debugging before optimization is complete. [paper] [episode]
- The Cognitive Continuity Test: Verifying Governed State Transitions in Persistent AI Agents The cognitive continuity test verifies that persistent AI agents maintain their governed state transitions correctly. [paper] [episode]
- ZoneClaw: Mitigating Persistent Memory Attacks by Establishing Memory-Zoning in OpenClaw-Style Computer-Use Agents ZoneClaw mitigates persistent memory attacks by dividing agent memory into zones. [paper] [episode]
- SafeDepth: Safety-Aware Token-Level Adaptive Computation SafeDepth implements safety-aware token-level adaptive computation to improve model safety. [paper] [episode]
- Do Defenses Against LLM Extraction Work Across Attacks? A Lifecycle Benchmark of Black-Box Model Extraction This paper benchmarks whether defenses against LLM extraction work across the entire lifecycle of black-box model extraction attacks. [paper] [episode]
- ABSENTIA: Detecting Broken Access Control Vulnerabilities in Web Applications ABSENTIA detects broken access control vulnerabilities in web applications. [paper] [episode]
- Helol Tunnel: Covert Channel Exploitation of TLS Extensibility & Privacy Features Helol Tunnel exploits covert channels within TLS extensibility and privacy features for data exfiltration. [paper] [episode]
- A Resource-Aware Behavior Reconstruction and Hierarchical Semantic Learning Framework for Host Intrusion Detection This framework reconstructs host intrusion behavior using resource awareness and hierarchical semantic learning. [paper] [episode]
- Autonomous OSS Threat Detection via Taxonomy-Aligned LLMs Autonomous OSS threat detection is achieved using language models aligned with threat taxonomy. [paper] [episode]
- Evidence Coverage for Intent-Bound Execution: Scope, Obligations, and Cutoff Reasoning This paper defines evidence coverage for execution based on intent, scope, obligations, and cutoff reasoning. [paper] [episode]
- False Floors: LLM Safety Routing Evaluations Break Under Distribution Shift False Floors shows that LLM safety routing evaluations fail when the data distribution shifts. [paper] [episode]
- Chaining Skills to Hijack LLM Agents Chaining skills can be used to hijack language model agents by chaining their different skills together. [paper] [episode]
- Protocol Integration of Physical Layer Deception into EAP-TEAP Wi-Fi Authentication This work integrates physical layer deception into EAP-TEAP Wi-Fi authentication protocols. [paper] [episode]
- Combining Homomorphic Encryption and Differential Privacy in Federated Learning for Model Inspection and Availability This paper combines homomorphic encryption and differential privacy to allow for model inspection while maintaining availability in federated learning. [paper] [episode]
- Safety in Self-Evolving Agents: A Survey This survey reviews the current state of safety considerations for self-evolving AI agents. [paper] [episode]
- Characterizing and Codifying Malware Sophistication This work focuses on characterizing and codifying the sophistication levels of malware. [paper] [episode]
- Intrusion Detection for Agentic Processes: Evidence-Based Runtime Monitoring This paper proposes evidence-based runtime monitoring to detect intrusions in agentic processes. [paper] [episode]
- Authorization for Self-Modifying AI Agent Populations: Conserving Authority across Replacement, Forking, and Rollback This work addresses how to conserve authority when self-modifying AI agents are replaced or rolled back. [paper] [episode]
- Removing the NEEDLE in the Haystack: Backdoor Removal in LLMs via Weight Orthogonalisation This method removes specific backdoor triggers from language models by using weight orthogonalization. [paper] [episode]
- Proof-Gated Signing: Solver-Checked Transaction Guards that Hold Under State Drift for Onchain AI Agents Proof-Gated Signing creates solver-checked transaction guards that remain valid even under state drift for onchain agents. [paper] [episode]
- On the Relationship between Model Quantization and Model Inversion Attacks This paper examines the relationship between model quantization techniques and model inversion attacks. [paper] [episode]
- From A2A Attacks to Envelope-Layer Defense: Red-Teaming Evaluation of LLM Agents and a Three-Layer Isomorphic Attack-Defense Model This work evaluates LLM agent attacks by comparing them to an envelope-layer defense model. [paper] [episode]
- Harbormaster: Evidence-Gated, Replay-Safe Maritime Anomaly Detection on AWS Harbormaster is an evidence-gated, replay-safe anomaly detection system for maritime activities on AWS. [paper] [episode]
- No One Architecture Fits All: A Cross-Environment Evaluation of Hierarchical Red Team Agents This paper evaluates hierarchical red team agents across different environments to see which architecture works best. [paper] [episode]
- Towards Hierarchical Cyber Defense with Large Language Models: From Planning to Execution This work proposes a hierarchical approach to cyber defense using large language models, from planning tasks down to execution. [paper] [episode]
- Progressive-Resolution Secure Aggregation for Federated Learning Progressive-Resolution Secure Aggregation improves secure aggregation in federated learning through progressive resolution. [paper] [episode]
- Crossing the Cyber Divide: Sim-to-Sim and Sim-to-Real Transfer for RL Agents This paper explores transfer techniques between simulation and real environments for reinforcement learning agents. [paper] [episode]
- Made to Measure: Designing Image Watermarks to Specification This research focuses on designing image watermarks that can be customized exactly as specified. [paper] [episode]
- Identity-Bound Governance Under Execution Uncertainty: An Accountability Proof Block for LLM Agent Persistent Halts, with Cryptographic Implementation and Cross-Model Calibration This paper provides a cryptographic proof block for accountability when an LLM agent halts persistently under execution uncertainty. [paper] [episode]
- MOMAT: Mixture of Multiple Atlases for Low-Power Jailbreak Defense of Quantized LLMs MOMAT uses a mixture of multiple atlases to defend quantized language models against jailbreaks with low power. [paper] [episode]
- Jev-IDS: System One Models for Network Intrusion Detection Jev-IDS uses System One models as a framework for network intrusion detection. [paper] [episode]
- A Systematization of Knowledge on DeFi Vaults: Architectures, Curation Mechanisms, and Strategy Design This paper systematizes knowledge about decentralized finance vaults by designing architectures and curation mechanisms. [paper] [episode]
- PACE: Provenance-Aware Capability Enforcement for Tool-Using LLM Agents PACE enforces tool-using LLM agent capabilities while tracking provenance. [paper] [episode]
- Sleeping Secrets: How Fine-Tuning Reawakens Privacy Risks in Language Models This paper investigates how fine-tuning can reawaken privacy risks that were previously suppressed in language models. [paper] [episode]
- High-quality Data Do not Mean Safe! Poisoning LLMs after Data Selection This work shows that high-quality data does not guarantee safety and discusses poisoning LLMs after initial data selection. [paper] [episode]
- Key-Reuse Vulnerability of Phase-Keyed Fourier-Curve Modulation: Relation Leakage and Key-Refresh Cost on Coded Links This paper analyzes key reuse vulnerabilities, relation leakage, and refresh costs in phase-keyed Fourier-curve modulation. [paper] [episode]
- The Achilles' Heel of Partial Reconfiguration: Optical Side-Channel Leakage on the 7-Series ICAP This work identifies optical side-channel leakage as the Achilles' heel of partial reconfiguration on 7-series ICAPs. [paper] [episode]
- SoK: Decentralized Agent Economic Infrastructure SoK proposes a decentralized economic infrastructure for autonomous agents. [paper] [episode]
- The Innocent Courier: Covert Exfiltration Through Legitimate LLM Web Fetching This paper studies covert data exfiltration methods through legitimate web fetching by language models. [paper] [episode]
- Walking the Embedding Space: Datastore Extraction from Multimodal RAG This work shows how to extract datastores from multimodal retrieval augmented generation systems by walking the embedding space. [paper] [episode]
- From Network Intrusion Detection to Blockchain-Backed Endpoint Detection and Response: Mapping the Landscape of Decentralized Detection-and-Response Architectures This paper maps the landscape of decentralized detection and response architectures, moving from network intrusion detection to blockchain-backed endpoint detection and response. [paper] [episode]
- A Structured State Space Sequence Model for Multi-Class Classification of Malware This model uses a structured state space sequence to classify malware into multiple classes. [paper] [episode]
- Detection and Resolution of Periodic Artifacts in OpenDP's Discrete Laplace Sampler This paper focuses on detecting and resolving periodic artifacts found in the discrete Laplace sampler within OpenDP. [paper] [episode]
The papers
- Authorization for Self-Modifying AI Agent Populations: Conserving Authority across Replacement, Forking, and Rollback — As a meticulous researcher, I have thoroughly reviewed both provided texts from the arXiv paper "Authorization for Self-Modifying AI Agent Populations: Conserving Authority across Replacement, Forking, and Rollback." The material presents a highly sophisticated formal system desi [episode]
- A Comprehensive Review of One-Pixel Attack: Research Status, Taxonomy, Applications, Regulation Policy and Future Directions — As a fastidious researcher with millions on the line, I will provide a comprehensive and meticulously detailed synthesis of Paper A (and by extension Paper B's title) based solely on the provided text excerpts. [episode]
- SoK: Decentralized Agent Economic Infrastructure — As a fastidious and diligent AI researcher, I have meticulously analyzed both provided texts (A and B) concerning the research on decentralized agent economies, specifically focusing on security guarantees, economic incentives, and workflow integrity. [episode]
- Sleeping Secrets: How Fine-Tuning Reawakens Privacy Risks in Language Models — Fine-tuning Large Language Models (LLMs) can reawaken latent privacy risks, allowing previously learned private associations to become substantially more recoverable even without genuine private supervision. [episode]
- Made to Measure: Designing Image Watermarks to Specification — Image watermarking supports provenance and attribution by embedding verifiable identity information into images, and this paper proposes TAILOR, a request-conditioned framework that jointly selects complementary watermark fragments and their configurations to satisfy specific dep [episode]
- A Hybrid Approach to Malware Detection: Integrating Few-Shot Model-Agnostic Meta-Learning with Autoencoders — A hybrid deep learning framework combining an Autoencoder Feature Extractor (AFE) with a Model-Agnostic Meta-Learning (MAML) classifier addresses the challenge of few-shot malware detection by leveraging unsupervised feature extraction to create compact representations and meta-l [episode]
- MOMAT: Mixture of Multiple Atlases for Low-Power Jailbreak Defense of Quantized LLMs — Quantized large language models (qLLMs) are increasingly deployed on edge devices for their low latency and energy efficiency, but model quantization weakens alignment safeguards, leaving qLLMs highly vulnerable to jailbreak attacks. [episode]
- CausalArmor: Efficient Indirect Prompt Injection Guardrails via Causal Attribution — AI agents equipped with tool-calling capabilities are susceptible to Indirect Prompt Injection (IPI) attacks, and CausalArmor proposes a selective defense framework that detects dominance shifts at privileged decision points using causal attribution to mitigate these threats whil [episode]
- High-quality Data Do not Mean Safe! Poisoning LLMs after Data Selection — Safety-aligned Large Language Models remain vulnerable to fine-tuning on small sets of harmful or benign-looking samples, and this vulnerability can persist even when quality-based data selection is employed. [episode]
- Safety in Self-Evolving Agents: A Survey — As a fastidious and diligent researcher, I have meticulously analyzed these provided excerpts from the paper "Safety in Self-Evolving Agents: A Survey." The synthesis below integrates all key findings, frameworks, and contributions into a comprehensive overview of the work's scop [episode]
- TESLA-for-5G: Broadcast Authentication for 5G Networks Using TESLA — 5G base stations broadcast unauthenticated system information (SI) that every user equipment (UE) reads during cell selection, enabling attackers to deploy fake base stations (FBS) to deceive UEs into camping on them for various malicious purposes. [episode]
- The Cognitive Continuity Test: Verifying Governed State Transitions in Persistent AI Agents — Persistent AI agents revise beliefs, consolidate memory, and replace execution substrates. [episode]
- Harbormaster: Evidence-Gated, Replay-Safe Maritime Anomaly Detection on AWS — Ships broadcast their positions through AIS, and those reports can be false or missing. This paper describes Harbormaster, a production-shaped system on AWS that follows three rules. [episode]
- Detection and Resolution of Periodic Artifacts in OpenDP's Discrete Laplace Sampler — Systematic artifacts were discovered in OpenDP’s discrete Laplace sampler, which manifest as periodic distortions in the output distribution and compromise theoretical privacy guarantees. [episode]
- A Structured State Space Sequence Model for Multi-Class Classification of Malware — By 2030, as Internet of Things (IoT) devices project to reach 40 billion, they present a massive attack surface for cybercrime due to inadequate built-in security and the rapid creation of malware variants. [episode]
- A Verifier Can Leak the Answer: Diagnosability Before Optimization in Closed-Loop Agent Debugging — A verifier-grounded evaluation shows that an optimizer can appear effective without resolving genuine ambiguity if its probes or predicates encode the target identity. [episode]
- Jev-IDS: System One Models for Network Intrusion Detection — Machine-learning Network Intrusion Detection Systems (IDS) often depend on substantial labeled datasets, but this work presents JEV-IDS, an open experimental general NIDS based on the Jev System One Model (SOM) to detect zero day intrusions under label scarcity. [episode]
- False Floors: LLM Safety Routing Evaluations Break Under Distribution Shift — As a diligent AI researcher, I have thoroughly analyzed both provided excerpts from the paper "False Floors: LLM Safety Routing Evaluations Break Under Distribution Shift." The information is dense, highly technical, and critical for understanding the nuanced findings regarding s [episode]
- Identity-Bound Governance Under Execution Uncertainty: An Accountability Proof Block for LLM Agent Persistent Halts, with Cryptographic Implementation and Cross-Model Calibration — A correctly governed LLM agent can reach a state in which neither continuing execution nor automatically halting is admissible: the system has detected a persistent failure of its observability or drift-detection layer, but cannot itself decide who has the authority to resume, de [episode]
- Autonomous OSS Threat Detection via Taxonomy-Aligned LLMs — A taxonomy-aligned large language model framework for automated detection and classification of open source software (OSS) supply chain threats has been proposed, demonstrating that structured prompting significantly outperforms traditional machine learning approaches in classify [episode]
- Removing the NEEDLE in the Haystack: Backdoor Removal in LLMs via Weight Orthogonalisation — Backdoor attacks can be implanted in Large Language Models (LLMs) during training, causing unwanted behavior when a trigger appears in the input. [episode]
- The Innocent Courier: Covert Exfiltration Through Legitimate LLM Web Fetching — With increasing LLM capabilities, users are increasingly using them to solve everyday problems, raising concerns about data leakage to third parties. [episode]
- PSR2: A Phase-based Semantic Reasoning Framework for Atomicity Violation Detection via Contract Refinement — PSR2 proposes a novel collaborative static analysis framework that integrates structural path searching with deterministic semantic reasoning to detect atomicity violations in smart contracts, addressing limitations in traditional tools by fusing graph-based evidence with semanti [episode]
- SafeDepth: Safety-Aware Token-Level Adaptive Computation — Token-level adaptive computation allows different tokens to execute different subsets of Transformer layers, but existing methods show that these execution choices negatively impact model safety. [episode]
- Proof-Gated Signing: Solver-Checked Transaction Guards that Hold Under State Drift for Onchain AI Agents — AI agents controlling wallets face risks from content they read, leading to harmful transactions, and this research introduces Proof-Gated Signing (PGS), a novel defense mechanism that closes the gap between pre-signing checks and actual transaction execution by using SMT solvers [episode]
- No One Architecture Fits All: A Cross-Environment Evaluation of Hierarchical Red Team Agents — Autonomous red team agents increasingly stress-test AI-enabled cyber defenses by planning strategy and executing multistage attacks, and this study addresses whether observed architectural advantages generalize across different cyber environments. [episode]
- ABSENTIA: Detecting Broken Access Control Vulnerabilities in Web Applications — Broken access control, which involves authorization failures where a principal acts on an unauthorized resource, remains difficult to detect in source code because its defect is defined by an application-specific relation rather than a universal dataflow property. [episode]
- Federated Detection of Open Charge Point Protocol 1.6 Cyberattacks — The ongoing electrification of transportation requires deploying numerous Electric Vehicle (EV) charging stations, which introduce significant cyber-physical and privacy risks due to vulnerable communication protocols like Open Charge Point Protocol (OCPP). [episode]
- ZoneClaw: Mitigating Persistent Memory Attacks by Establishing Memory-Zoning in OpenClaw-Style Computer-Use Agents — Computer-use agents increasingly operate as long-running assistants through persistent workspace memory, which OpenClaw-style CUAs realize as automatically reloaded files that hold user instructions, system summaries, and external claims at the same privilege level. [episode]
- Progressive-Resolution Secure Aggregation for Federated Learning — Secure aggregation lets a server recover an aggregate of client updates without observing any individual update, but conventional protocols fix the aggregate precision when clients upload. [episode]
- Crossing the Cyber Divide: Sim-to-Sim and Sim-to-Real Transfer for RL Agents — Cyber attack agents are typically trained and evaluated within a single simulator, making it unclear whether learned policies transfer beyond the environments in which they were developed. [episode]
- Do Defenses Against LLM Extraction Work Across Attacks? A Lifecycle Benchmark of Black-Box Model Extraction — Large language models deployed through text-only APIs face model extraction risks, as adversaries can collect their responses to train surrogates that reproduce their capabilities. [episode]
- Helol Tunnel: Covert Channel Exploitation of TLS Extensibility & Privacy Features — Covert channels exploiting network protocols for data exfiltration and command-and-control (C2) are integral parts of modern cyberattacks, and this research proposes a novel method to exploit combinatorial properties within TLS Client Hello packets to evade security measures. [episode]
- PACE: Provenance-Aware Capability Enforcement for Tool-Using LLM Agents — This document synthesizes information from two distinct perspectives—a high-level technical overview of the PACE mechanism and a detailed artifact/experiment report—to provide a comprehensive understanding of Provenance-Aware Capability Enforcement (PACE), a novel security fr [episode]
- Chaining Skills to Hijack LLM Agents — LLM agents use skills to improve performance on specialized tasks, but because these skills can be sourced from public repositories, they create a vulnerability where an attacker can control claims made across sequential skill invocations to redirect the agent's behavior. [episode]
- A Systematization of Knowledge on DeFi Vaults: Architectures, Curation Mechanisms, and Strategy Design — Decentralized finance (DeFi) vaults are smart-contract-based asset management systems that pool deposits, execute programmable strategies, and mint tokenized shares representing claims on underlying assets and strategy performance. [episode]
- From Network Intrusion Detection to Blockchain-Backed Endpoint Detection and Response: Mapping the Landscape of Decentralized Detection-and-Response Architectures — While existing literature on blockchain-assisted intrusion detection and prevention systems (IDS/IPS) for IoT and IIoT networks is mature, current systematic reviews suffer from two critical limitations: they overlook the structural shift toward modern Endpoint Detection and Resp [episode]
- System-Level Optimization Beyond Cryptographic Kernels: An ML-KEM Case Study on Arm Cortex-M7 — Recent work on embedded post-quantum cryptography has focused primarily on instruction-level optimization, including arithmetickernel improvements, assembly tuning, register allocation, and instruction scheduling. [episode]
- From A2A Attacks to Envelope-Layer Defense: Red-Teaming Evaluation of LLM Agents and a Three-Layer Isomorphic Attack-Defense Model — Agent interaction protocols like A2A introduce new security threats that necessitate a more granular evaluation framework than traditional attack success rates. [episode]
- Protocol Integration of Physical Layer Deception into EAP-TEAP Wi-Fi Authentication — Credential-based Extensible Authentication Protocol (EAP) authentication cannot distinguish a legitimate credential holder from an adversary using compromised credentials. [episode]
- Combining Homomorphic Encryption and Differential Privacy in Federated Learning for Model Inspection and Availability — Combining Homomorphic Encryption and Differential Privacy in Federated Learning for Model Inspection and Availability proposes a privacy-preserving federated learning framework that integrates homomorphic encryption for training utility with differential privacy for model inspect [episode]
- Towards Hierarchical Cyber Defense with Large Language Models: From Planning to Execution — An autonomous cyber defender trained with reinforcement learning (RL) is typically tied to its training network, limiting its generalization, and this research investigates whether frozen, zero-shot large language models (LLMs) can provide retraining-free control in a hierarchica [episode]
- Walking the Embedding Space: Datastore Extraction from Multimodal RAG — As a fastidious and diligent AI researcher, I have meticulously analyzed both provided excerpts from what appears to be related research papers concerning multimodal retrieval systems and image generation evaluation. [episode]
- Evidence Coverage for Intent-Bound Execution: Scope, Obligations, and Cutoff Reasoning — A verifier may authenticate every available record and still lack grounds to call an execution account complete, necessitating an analytical model to justify which records were due for assessment. [episode]
- TensorCommitments: A Lightweight Verifiable Inference for Language Models — Most large language models (LLMs) run on external clouds, and there is a critical need for verifiable inference where a service must convince a client that an LLM inference was executed correctly without rerunning the model. [episode]
- A Resource-Aware Behavior Reconstruction and Hierarchical Semantic Learning Framework for Host Intrusion Detection — System calls provide fine-grained data for host-based intrusion detection, but existing methods struggle to extract informative patterns from raw sequences due to concurrent execution interleaving. [episode]
- On the Relationship between Model Quantization and Model Inversion Attacks — Model quantization reduces numerical precision to lower storage and computational costs, and this work investigates how these changes affect model inversion attacks. [episode]
- Characterizing and Codifying Malware Sophistication — “Sophistication” is widely used to describe malware, yet it lacks a consistent definition within academic literature. [episode]
- Key-Reuse Vulnerability of Phase-Keyed Fourier-Curve Modulation: Relation Leakage and Key-Refresh Cost on Coded Links — Reusing a phase key in harmonically coupled modulation converts short modular relations among the harmonic indices into estimable key characters, demonstrating that nominal key-space size and error rates are insufficient security evidence for keyed modulations with repeated wavef [episode]
- Rethinking Anonymity Claims in Synthetic Data Generation: A Model-Centric Privacy Attack Perspective — Training generative machine learning models to produce synthetic tabular data has become a popular approach for enhancing privacy in data sharing, but recent research suggests that meaningful assessments must account for the capabilities and properties of the underlying generativ [episode]
- Understanding Gaps in LLM Pipelines Towards Scalable Fuzzing Harness Generation: An Empirical Study and Enhancement — Large language model (LLM)-based techniques have achieved notable progress in generating harnesses for program fuzzing, but applying them to arbitrary functions at scale remains challenging due to the requirement of sophisticated contextual information, such as specification, dep [episode]
- GNSS Spoofing in Mobile Devices: A Survey on Impact and Countermeasures — GNSS spoofing in mobile devices represents an insidious threat where forged satellite signals aim to cause victim receivers to compute false Position, Velocity, and Time (PVT) solutions, making smartphones primary targets due to their ubiquity and sensitive geolocation data. [episode]
- Intrusion Detection for Agentic Processes: Evidence-Based Runtime Monitoring — Agent deployments increasingly combine language-model inference with retrieval, delegation, tool execution, external-system access, and human approval. [episode]
- Evasion Attacks: How Adversarial Noise Bypasses ML Classifiers — Adversarial examples are intentionally perturbed inputs designed to alter a machine-learning model’s prediction while remaining close to the original input under a chosen perturbation constraint, and this study investigates how these manipulations bypass classifiers in both ima [episode]
- The Achilles' Heel of Partial Reconfiguration: Optical Side-Channel Leakage on the 7-Series ICAP — Major FPGA manufacturers have incorporated bitstream encryption to protect sensitive configuration data, but this work presents a proof-of-concept implementation of an AMD-proposed asymmetric key encryption scheme for 7-Series FPGAs and demonstrates that even patchable protection [episode]
- Actions with Receipts: Jointly Binding Claims, Evidence, and Execution for Replayable Tool-Agent Auditing —
- AuraForge: Scaling Security Supervision for Training Coding Agents —
- UnifiedAttack: Evaluating the Safety of Large Multimodal Models in Synergistic Harmful Image-Text Generation —
- ReCast: Contract-Preserving Protection for Fixed-Interface Multimodal Reasoning —
- OverAct: Measuring and Mitigating Proactive Over-Authorization in LLM Tool-Calling Agents —
- Is it Possible to Generate Irreversible PolyProtected Templates from Face Embeddings using System-Specific Keys? —
- Tokenized Key-Gated Adapter Routing: A Secure Access Control Mechanism Against Private Data Leakage in LLMs —
Important terms
- Adversarial Noise
- Attackers use subtle, carefully crafted noise to bypass machine learning classifiers. This is a major threat because it means we can't trust models against slight manipulations, compromising system security.
- UnifiedAttack
- This research evaluates the safety of large multimodal models by testing how they can be exploited together. It shows how combining image and text generation capabilities creates synergistic risks.
- TensorCommitments
- This is a lightweight method to verify language model outputs efficiently. It uses tensor commitments to check if the model's output matches its training data without needing huge computational power.
- CausalArmor
- This technique creates indirect prompt injection guardrails by tracing the cause of an input's effect. It helps stop malicious inputs from tricking models by understanding how they influence the outcome.