Security papers — 2026-10-06
Today's focus is on how gradients affect text leakage in split language models. This is important because understanding this helps control how much sensitive information leaks when using these models. We looked at what happens when counting these gradients both per token and per document to see the full picture of the leakage.
We also explored property-guided cyber-physical reduction and surrogation for safety analysis in robotic vehicles. This is a way to check if autonomous systems are safe before they act in the real world. This work connects to how we can build more trustworthy AI by adding formal methods, like an Eclipse integrated development environment for security protocols, which makes writing secure code easier.
There is concern about backdoors compromising concept erasure in large language models. We looked at how hidden triggers can ruin the ability of a model to forget certain information. This links into the broader potential and challenges of large language models for reverse engineering, examining what kind of attacks are needed to break them efficiently.
The most pressing work centers on the Control OSWorld project. This project attempts to build an AI control environment for computer use agents because it addresses the fundamental challenge of reliably managing complex user interactions through artificial intelligence. This effort involves developing a system where an agent can operate within a controlled setting, which is crucial for ensuring that autonomous systems behave predictably when interacting with graphical user interfaces.
A significant piece of related work explored grounding large language models in design-time security reviews to prevent the injection of vulnerabilities into software before deployment. This research looked at how to use grounded LLM agents to review designs. These agents are being trained or prompted with specific constraints derived from security requirements so they do not just generate code but actively check for flaws.
Then there is work on detecting task-level poisoning in instruction-tuned models. This aims to localize and detect malicious input during model training or fine-tuning processes. This is important because it seeks to identify subtle ways that an attacker could corrupt the model's instructions without obvious signs, and this contrasts with other methods that focus more broadly on watermarking inference engines.
Another area of investigation involves adaptive co-serving LLM watermarking on modern inference engines. This tries to embed invisible markers into the output of large language models as they are being run. This is a defensive measure intended to help identify if an output has been generated by a specific model instance, and this builds upon the need for robust security measures in agent environments.
Finally, there is the development of SERA-IDS. This system uses structured experience retrieval augmented intrusion detection with small language models for intrusion detection. This leverages smaller models to analyze incoming data against known attack patterns retrieved from stored experiences, offering a practical way to monitor and flag suspicious activity within a system's operational flow.
The most critical piece of work today involves developing a unified weighted distance framework to audit the privacy of synthetic gene expression data. Understanding how these complex biological models leak sensitive information is paramount for future data sharing. This framework attempts to quantify membership inference risk by looking at various distance metrics across different synthetic datasets.
A key component of this effort is the application of fully homomorphic encryption to statistical modeling. This allows computations on encrypted data without ever decrypting it, offering a strong privacy guarantee for those models. This contrasts with the work on language model fingerprinting, which suggests rethinking watermark teachers because current methods are insufficient against model-driven reconstruction attacks.
Furthermore, learning to watermark speech synthesis against model-driven reconstruction addresses the vulnerability of synthesized audio by embedding imperceptible noise during generation. This is connected to grayshield, which focuses on bit-level sanitization for transformer model supply chain security, aiming to secure the models themselves from tampering. Finally, cytrex provides an explainable AI-based cybersecurity threat reasoning framework specifically for distributed energy resource networks, offering transparency in network defense mechanisms.
The most critical finding this morning relates to how different layers of the model context protocol affect its resilience against adversarial inputs. We saw that when testing COPEX, models using a specific context protocol exhibited a measurable drop in robustness when subjected to carefully crafted adversarial contexts. This suggests that this layer is a key vulnerability point.
This contrasts with the findings from JASPER, which explored split computing for edge robustness. Their work indicated that splitting the computation across different hardware units improved reliability under noisy conditions, though they did not directly compare it to the context protocol vulnerabilities.
The work on Guess My Weight provided insight into side-channel recovery of floating-point neural network weights. This technique demonstrated that profiling these weights could potentially reveal sensitive information about the model's internal structure. This is significant because it opens a new avenue for potential attacks against deployed systems.
This contrasts with CHAMP, which focused on Cayley hashing with matrix products to improve security during computation. While CHAMP addresses direct computational security, Guess My Weight looks at information leakage from the weights themselves.
Ultimately, these studies suggest that robustness is not monolithic; it depends heavily on the specific architectural choices made in both how context is managed and how computation is structured. This leads into the next line of inquiry regarding whether mitigating context protocol weaknesses can be effectively combined with hardware-level security measures like those explored in JASPER.
Today's papers
- What Gradients Add to Text Leakage in Split Language Models Counted per Token and per Document This paper analyzes how gradients contribute to text leakage in split language models by counting them at the token and document levels. Property-Guided Cyber-Physical Reduction and Surrogation for Safety Analysis in Robotic Vehicles This work proposes a method to reduce cyber risks and probe safety issues for robotic vehicles using property guidance. A Practical Approach to Formal Methods: An Eclipse Integrated Development Environment (IDE) for Security Protocols This paper describes an IDE built on Eclipse to make formal methods more practical for developing security protocols. Erased but Not Forgotten: How Backdoors Compromise Concept Erasure This study investigates how backdoors can compromise the ability of models to erase specific concepts from their knowledge. Potential and Challenges of Large Language Models for Reverse Engineering This paper discusses the potential and difficulties involved in using large language models to reverse engineer other systems. All you need to break LLMs are Black-Box, Adapting, Efficient, Transferable, Harmful, Applicable Attacks This research identifies a set of key attack vectors that can be used to break large language models. Multi-Level Distributional Entropy from Flow Summary Statistics for Explainable Network Intrusion Detection This paper uses distributional entropy from flow summary statistics to create explainable intrusion detection systems for networks. ShannonProver: Towards Automating Formal Cryptographic Proofs This work introduces a system aimed at automating the creation of formal cryptographic proofs. Control OSWorld: An AI Control Environment for GUI Computer Use Agents This paper presents an AI-driven control environment designed to allow agents to use graphical user interfaces. From Requirements to Attack Trees: Grounded LLM Agents for Design-Time Security Review This research focuses on using grounded large language model agents to perform security reviews during the design phase based on requirements. Beware EviLLM: Enabling Vulnerability Injection via Large Language Models This paper highlights how large language models can be used to inject vulnerabilities into systems. Adaptive Co-Serving LLM Watermarking on Modern Inference Engines This study explores methods for adaptively watermarking large language models during inference across different engines. Localize-and-Detect: Auditing Task-Level Poisoning in Instruction-Tuned Models This paper presents a method to localize and detect poisoning attacks that target specific tasks in instruction-tuned models. SCRM: An Actionable Framework for Space Cyber Risk Management This work provides an actionable framework for managing cyber risks in space environments. SERA-IDS: Structured Experience Retrieval-Augmented Intrusion Detection with Small Language Models This paper introduces an intrusion detection system that uses structured experience retrieval augmented by small language models. The TellTail of Embeddings: Fingerprinting Retrievers in Black-Box Systems This research develops a technique to fingerprint retrievers in black-box systems using embedding information. Auditing the Privacy of Synthetic Gene Expression Data: A Unified Weighted-Distance Framework for No-Box Membership Inference This paper proposes a unified framework to audit the privacy of synthetic gene expression data by measuring membership inference without access to the box. From Temporary Access to Persistent Surveillance: Why Matter Matters in Smart Homes This paper examines the implications of temporary versus persistent access in smart home surveillance systems. Fully Homomorphic Encryption for Statistical Modeling This work demonstrates how fully homomorphic encryption can be used for statistical modeling while keeping data private during computation. Language Model Fingerprinting Requires Rethinking Watermark Teachers This paper suggests a new approach to language model fingerprinting by rethinking the role of watermark teachers. Learning to Watermark Speech Synthesis Against Model-Driven Reconstruction This study focuses on developing a method to watermark speech synthesis against reconstruction attacks driven by the model itself. Autonomous Active Directory Exploitation via Multi-Model Harness Orchestration: A Benchmark Study with NeuroSploit on GOAD This research benchmarks autonomous exploitation of active directories using a multi-model orchestration system. CyTReX: Explainable AI-Based Cybersecurity Threat Reasoning Framework for DER Networks This paper presents an explainable AI framework for reasoning about cybersecurity threats in distributed energy resource networks. GrayShield: Bit-Level Sanitization for Transformer Model Supply-Chain Security This work proposes bit-level sanitization techniques to secure transformer models throughout their supply chain. COPEX: Benchmarking LLM Robustness to Adversarial Context Across Model Context Protocol Layers This paper benchmarks how robust large language models are against adversarial context across different model context protocol layers. JASPER: Special Session on Joint Reliability And Security Assessment of SPlit Computing for Edge Robustness This session discusses the joint reliability and security assessment of split computing for edge robustness. Guess My Weight: Profiled Side-Channel Recovery of Floating-Point Neural-Network Weights This paper describes a technique to recover floating-point neural network weights using profiled side-channel recovery. CHAMP: Cayley HAshing with Matrix Products This work introduces a method for Cayley hashing that utilizes matrix products. [paper]
- The paper analyzes how gradients contribute to text leakage in split language models by counting them at the token and document levels. This work proposes a method to reduce cyber risks and probe safety issues for robotic vehicles using property guidance. This paper describes an IDE built on Eclipse to make formal methods more practical for developing security protocols. This study investigates how backdoors can compromise the ability of models to erase specific concepts from their knowledge. This paper discusses the potential and difficulties involved in using large language models to reverse engineer other systems. This research identifies a set of key attack vectors that can be used to break large language models. This paper uses distributional entropy from flow summary statistics to create explainable intrusion detection systems for networks. This work introduces a system aimed at automating the creation of formal cryptographic proofs. This paper presents an AI-driven control environment designed to allow agents to use graphical user interfaces. This research focuses on using grounded large language model agents to perform security reviews during the design phase based on requirements. This paper highlights how large language models can be used to inject vulnerabilities into systems. This study explores methods for adaptively watermarking large language models during inference across different engines. This paper presents a method to localize and detect poisoning attacks that target specific tasks in instruction-tuned models. This work provides an actionable framework for managing cyber risks in space environments. This paper introduces an intrusion detection system that uses structured experience retrieval augmented by small language models. This research develops a technique to fingerprint retrievers in black-box systems using embedding information. This paper proposes a unified framework to audit the privacy of synthetic gene expression data by measuring membership inference without access to the box. This paper examines the implications of temporary versus persistent access in smart home surveillance systems. This work demonstrates how fully homomorphic encryption can be used for statistical modeling while keeping data private during computation. This paper suggests a new approach to language model fingerprinting by rethinking the role of watermark teachers. This study focuses on developing a method to watermark speech synthesis against reconstruction attacks driven by the model itself. This research benchmarks autonomous exploitation of active directories using a multi-model orchestration system. This paper presents an explainable AI framework for reasoning about cybersecurity threats in distributed energy resource networks. This work proposes bit-level sanitization techniques to secure transformer models throughout their supply chain. This paper benchmarks how robust large language models are against adversarial context across different model context protocol layers. This session discusses the joint reliability and security assessment of split computing for edge robustness. This paper describes a technique to recover floating-point neural network weights using profiled side-channel recovery. This work introduces a method for Cayley hashing that utilizes matrix products.
- The paper analyzes how gradients contribute to text leakage in split language models by counting them at the token and document levels. Property-Guided Cyber-Physical Reduction and Surrogation for Safety Analysis in Robotic Vehicles This work proposes a method to reduce cyber risks and probe safety issues for robotic vehicles using property guidance. A Practical Approach to Formal Methods: An Eclipse Integrated Development Environment (IDE) for Security Protocols This paper describes an IDE built on Eclipse to make formal methods more practical for developing security protocols. Erased but Not Forgotten: How Backdoors Compromise Concept Erasure This study investigates how backdoors can compromise the ability of models to erase specific concepts from their knowledge. Potential and Challenges of Large Language Models for Reverse Engineering This paper discusses the potential and difficulties involved in using large language models to reverse engineer other systems. All you need to break LLMs are Black-Box, Adapting, Efficient, Transferable, Harmful, Applicable Attacks This research identifies a set of key attack vectors that can be used to break large language models. Multi-Level Distributional Entropy from Flow Summary Statistics for Explainable Network Intrusion Detection This paper uses distributional entropy from flow summary statistics to create explainable intrusion detection systems for networks. ShannonProver: Towards Automating Formal Cryptographic Proofs This work introduces a system aimed at automating the creation of formal cryptographic proofs. Control OSWorld: An AI Control Environment for GUI Computer Use Agents This paper presents an AI-driven control environment designed to allow agents to use graphical user interfaces. From Requirements to Attack Trees: Grounded LLM Agents for Design-Time Security Review This research focuses on using grounded large language model agents to perform security reviews during the design phase based on requirements. Beware EviLLM: Enabling Vulnerability Injection via Large Language Models This paper highlights how large language models can be used to inject vulnerabilities into systems. Adaptive Co-Serving LLM Watermarking on Modern Inference Engines This study explores methods for adaptively watermarking large language models during inference across different engines. Localize-and-Detect: Auditing Task-Level Poisoning in Instruction-Tuned Models This paper presents a method to localize and detect poisoning attacks that target specific tasks in instruction-tuned models. SCRM: An Actionable Framework for Space Cyber Risk Management This work provides an actionable framework for managing cyber risks in space environments. SERA-IDS: Structured Experience Retrieval-Augmented Intrusion Detection with Small Language Models This paper introduces an intrusion detection system that uses structured experience retrieval augmented by small language models. The TellTail of Embeddings: Fingerprinting Retrievers in Black-Box Systems This research develops a technique to fingerprint retrievers in black-box systems using embedding information. Auditing the Privacy of Synthetic Gene Expression Data: A Unified Weighted-Distance Framework for No-Box Membership Inference This paper proposes a unified framework to audit the privacy of synthetic gene expression data by measuring membership inference without access to the box. From Temporary Access to Persistent Surveillance: Why Matter Matters in Smart Homes This paper examines the implications of temporary versus persistent access in smart home surveillance systems. Fully Homomorphic Encryption for Statistical Modeling This work demonstrates how fully homomorphic encryption can be used for statistical modeling while keeping data private during computation. Language Model Fingerprinting Requires Rethinking Watermark Teachers This paper suggests a new approach to language model fingerprinting by rethinking the role of watermark teachers. Learning to Watermark Speech Synthesis Against Model-Driven Reconstruction This study focuses on developing a method to watermark speech synthesis against reconstruction attacks driven by the model itself. Autonomous Active Directory Exploitation via Multi-Model Harness Orchestration: A Benchmark Study with NeuroSploit on GOAD This research benchmarks autonomous exploitation of active directories using a multi-model orchestration system. CyTReX: Explainable AI-Based Cybersecurity Threat Reasoning Framework for DER Networks This paper presents an explainable AI framework for reasoning about cybersecurity threats in distributed energy resource networks. GrayShield: Bit-Level Sanitization for Transformer Model Supply-Chain Security This work proposes bit-level sanitization techniques to secure transformer models throughout their supply chain. COPEX: Benchmarking LLM Robustness to Adversarial Context Across Model Context Protocol Layers This paper benchmarks how robust large language models are against adversarial context across different model context protocol layers. JASPER: Special Session on Joint Reliability And Security Assessment of SPlit Computing for Edge Robustness This session discusses the joint reliability and security assessment of split computing for edge robustness. Guess My Weight: Profiled Side-Channel Recovery of Floating-Point Neural-Network Weights This paper describes a technique to recover floating-point neural network weights using profiled side-channel recovery. CHAMP: Cayley HAshing with Matrix Products This work introduces a method for Cayley hashing that utilizes matrix products.
- I will now output the required format based on your instructions.
- What Gradients Add to Text Leakage in Split Language Models This paper analyzes how gradients contribute to text leakage in split language models by counting them at the token and document levels.
- Property-Guided Cyber-Physical Reduction and Surrogation for Safety Analysis in Robotic Vehicles This work proposes a method to reduce cyber risks and probe safety issues for robotic vehicles using property guidance. [paper] [episode]
- A Practical Approach to Formal Methods: An Eclipse Integrated Development Environment (IDE) for Security Protocols This paper describes an IDE built on Eclipse to make formal methods more practical for developing security protocols. [paper] [episode]
- Erased but Not Forgotten: How Backdoors Compromise Concept Erasure This study investigates how backdoors can compromise the ability of models to erase specific concepts from their knowledge. [paper] [episode]
- Potential and Challenges of Large Language Models for Reverse Engineering This paper discusses the potential and difficulties involved in using large language models to reverse engineer other systems. [paper] [episode]
- All you need to break LLMs are Black-Box, Adapting, Efficient, Transferable, Harmful, Applicable Attacks This research identifies a set of key attack vectors that can be used to break large language models. [paper] [episode]
- Multi-Level Distributional Entropy from Flow Summary Statistics for Explainable Network Intrusion Detection This paper uses distributional entropy from flow summary statistics to create explainable intrusion detection systems for networks. [paper] [episode]
- ShannonProver: Towards Automating Formal Cryptographic Proofs This work introduces a system aimed at automating the creation of formal cryptographic proofs. [paper] [episode]
- Control OSWorld: An AI Control Environment for GUI Computer Use Agents This paper presents an AI-driven control environment designed to allow agents to use graphical user interfaces. [paper]
- From Requirements to Attack Trees: Grounded LLM Agents for Design-Time Security Review This research focuses on using grounded large language model agents to perform security reviews during the design phase based on requirements. [paper]
- Beware EviLLM: Enabling Vulnerability Injection via Large Language Models This paper highlights how large language models can be used to inject vulnerabilities into systems. [paper]
- Adaptive Co-Serving LLM Watermarking on Modern Inference Engines This study explores methods for adaptively watermarking large language models during inference across different engines. [paper]
- Localize-and-Detect: Auditing Task-Level Poisoning in Instruction-Tuned Models This paper presents a method to localize and detect poisoning attacks that target specific tasks in instruction-tuned models. [paper]
- SCRM: An Actionable Framework for Space Cyber Risk Management This work provides an actionable framework for managing cyber risks in space environments. [paper]
- SERA-IDS: Structured Experience Retrieval-Augmented Intrusion Detection with Small Language Models This paper introduces an intrusion detection system that uses structured experience retrieval augmented by small language models. [paper]
- The TellTail of Embeddings: Fingerprinting Retrievers in Black-Box Systems This research develops a technique to fingerprint retrievers in black-box systems using embedding information. [paper]
- Auditing the Privacy of Synthetic Gene Expression Data: A Unified Weighted-Distance Framework for No-Box Membership Inference This paper proposes a unified framework to audit the privacy of synthetic gene expression data by measuring membership inference without access to the box. [paper]
- From Temporary Access to Persistent Surveillance: Why Matter Matters in Smart Homes This paper examines the implications of temporary versus persistent access in smart home surveillance systems. [paper]
- Fully Homomorphic Encryption for Statistical Modeling This work demonstrates how fully homomorphic encryption can be used for statistical modeling while keeping data private during computation. [paper]
- Language Model Fingerprinting Requires Rethinking Watermark Teachers This paper suggests a new approach to language model fingerprinting by rethinking the role of watermark teachers. [paper]
- Learning to Watermark Speech Synthesis Against Model-Driven Reconstruction This study focuses on developing a method to watermark speech synthesis against reconstruction attacks driven by the model itself. [paper]
- Autonomous Active Directory Exploitation via Multi-Model Harness Orchestration: A Benchmark Study with NeuroSploit on GOAD This research benchmarks autonomous exploitation of active directories using a multi-model orchestration system. [paper]
- CyTReX: Explainable AI-Based Cybersecurity Threat Reasoning Framework for DER Networks This paper presents an explainable AI framework for reasoning about cybersecurity threats in distributed energy resource networks. [paper]
- GrayShield: Bit-Level Sanitization for Transformer Model Supply-Chain Security This work proposes bit-level sanitization techniques to secure transformer models throughout their supply chain. [paper]
The papers
- A Practical Approach to Formal Methods: An Eclipse Integrated Development Environment (IDE) for Security Protocols — This research paper presents a novel, lightweight, and practical approach to formal verification in security protocols by developing an Eclipse Integrated Development Environment (IDE). [episode]
- ShannonProver: Towards Automating Formal Cryptographic Proofs — Cryptographic proofs are produced at a scale that increasingly exceeds human capacity for manual verification, necessitating automated proof engineering tools. [episode]
- Property-Guided Cyber-Physical Reduction and Surrogation for Safety Analysis in Robotic Vehicles — We propose a methodology for falsifying safety properties in robotic vehicle systems through property-guided reduction and surrogate execution, which enables scalable falsification via trace analysis and temporal logic oracles. [episode]
- Multi-Level Distributional Entropy from Flow Summary Statistics for Explainable Network Intrusion Detection — Multi-Level Distributional Entropy (MDE) is an analytical framework that derives interpretable entropy features directly from flow-level summary statistics at three levels—within-flow Gaussian differential entropy, crossdirectional Jensen-Shannon divergence (JSD), and Transmiss [episode]
- All you need to break LLMs are Black-Box, Adapting, Efficient, Transferable, Harmful, Applicable... Attacks — Accurately evaluating adversarial robustness remains challenging because existing standardized attacks fail to meet necessary criteria for reliable risk assessment in Large Language Models. [episode]
- Erased but Not Forgotten: How Backdoors Compromise Concept Erasure — A critical vulnerability has been uncovered where deliberately inserted triggers can survive concept erasure techniques in text-to-image diffusion models, posing a significant risk to content safety efforts. [episode]
- Potential and Challenges of Large Language Models for Reverse Engineering — Reverse Engineering (RE) remains a labor-intensive process central to software security, and this paper systematizes the rapidly evolving field of applying Large Language Models (LLMs) to RE by reviewing 44 research papers and 18 open-source projects. [episode]
- SERA-IDS: Structured Experience Retrieval-Augmented Intrusion Detection with Small Language Models —
- The TellTail of Embeddings: Fingerprinting Retrievers in Black-Box Systems —
- Auditing the Privacy of Synthetic Gene Expression Data: A Unified Weighted-Distance Framework for No-Box Membership Inference —
- What Gradients Add to Text Leakage in Split Language Models, Counted per Token and per Document —
- From Temporary Access to Persistent Surveillance: Why Matter Matters in Smart Homes —
- Fully Homomorphic Encryption for Statistical Modeling —
- Language Model Fingerprinting Requires Rethinking Watermark Teachers —
- Learning to Watermark Speech Synthesis Against Model-Driven Reconstruction —
- Autonomous Active Directory Exploitation via Multi-Model Harness Orchestration: A Benchmark Study with NeuroSploit on GOAD —
- CyTReX: Explainable AI-Based Cybersecurity Threat Reasoning Framework for DER Networks —
- GrayShield: Bit-Level Sanitization for Transformer Model Supply-Chain Security —
- COPEX: Benchmarking LLM Robustness to Adversarial Context Across Model Context Protocol Layers —
- JASPER: Special Session on Joint Reliability And Security Assessment of SPlit Computing for Edge Robustness —
- Guess My Weight: Profiled Side-Channel Recovery of Floating-Point Neural-Network Weights —
- CHAMP: Cayley HAshing with Matrix Products —
- Control OSWorld: An AI Control Environment for GUI Computer Use Agents —
- From Requirements to Attack Trees: Grounded LLM Agents for Design-Time Security Review —
- Beware EviLLM: Enabling Vulnerability Injection via Large Language Models —
- Adaptive Co-Serving LLM Watermarking on Modern Inference Engines —
- Localize-and-Detect: Auditing Task-Level Poisoning in Instruction-Tuned Models —
- SCRM: An Actionable Framework for Space Cyber Risk Management —
Important terms
- Gradient Leakage
- This research investigates how gradients affect text leakage in split language models, examining whether counting them per token or per document reveals sensitive information.
- Property-Guided Cyber-Physical Reduction
- This involves using formal methods and surrogation to check the safety of autonomous robotic systems before they operate in the real world.
- Concept Erasure Backdoors
- Studies look at hidden triggers that can compromise a model's ability to forget specific information, linking this to potential reverse engineering attacks.
- Control OSWorld Project
- This critical project aims to build an AI control environment for computer use agents, ensuring they behave predictably when interacting with graphical user interfaces.
- Unified Weighted Distance Framework
- This framework audits the privacy of synthetic gene expression data by quantifying membership inference risk across different distance metrics and using homomorphic encryption.