Security papers — 2026-09-24
Today we are looking at how to make machine learning models better at spotting network intrusions, which is crucial because the more sophisticated the attacks get, the harder it is to defend against them. The main focus is on developing lightweight adversarial agents trained through reinforcement learning that can trick existing intrusion detection models. This approach works by training these agents offline using representative NetFlow data to generate evasion strategies that do not require complex gradient calculations when they are actually deployed in a real network environment.
This method has shown some promising results regarding efficiency and effectiveness against different types of models. For instance, the agents managed to achieve up to fifty-eight point one percent attack success at a very fast rate of zero point three one milliseconds per attack, which translates to over a thousand times the improvement in throughput compared to gradient-based methods. Even with a small policy configuration requiring only nineteen kilobytes of memory and four thousand nine hundred thirty-one parameters, the agent achieved forty-six percent attack success at zero point one eight milliseconds per attack.
Furthermore, these agents proved surprisingly resilient when tested against non-differentiable models. Traditional gradient methods lost over fifty-nine percent of their effectiveness due to surrogate transfer. The RL agent, however, evaluated these models directly and still achieved a twenty-nine point eight percent attack success rate without any marginal transferability penalty. This suggests that the way these agents learn an evasion strategy is robust across various model types.
We also looked at how well these learned strategies generalize when moving between different training scenarios. The findings indicated that the agents retained attack success under model transfer at a median of twelve point two percent, dataset transfer at eleven point four percent, and full transfer at nine point one percent. This shows a degree of practical robustness in their learned policies across different network conditions.
Finally, the study highlighted that the success of these attacks really depends on the type of attack and how much feature budget is available. Volumetric attacks like denial of service or brute force were most sensitive to small changes in byte and packet budgets. Malware and persistence attacks remained quite robust even when those budgets were extremely constrained. The authors conclude that these lightweight policies are a practical tool for evaluating ML robustness across different intrusion detection settings, though they caution that the defender-side benefits currently seem to outweigh any potential advantages for an attacker.
The most significant piece of work today involved exploring how to defeat federated learning servers using strategic gradient manipulation. This is crucial because it addresses a major vulnerability in decentralized machine learning systems. Researchers investigated how to orchestrate these manipulations to break the integrity of the learning process.
One line of inquiry focused on the like trap, which examines multi-stage poisoning against agents within similarity-based recommendation systems. This work suggests that subtle, layered attacks can compromise these systems effectively. This is related to another study looking at multi-view fusion for encrypted command and control detection, which explores leakage-controlled measurements to find evaluation pitfalls in those same environments.
Further down the list, there was a look at reliable federated tinyml deployment for IoT security. This aims to ensure that machine learning models can operate securely on small devices. This contrasts with work enhancing multiclass malware classification in resource-constrained environments, which focuses on improving how we identify malicious software when computational power is limited.
The research also touched upon issuer-sovereign agentic payments, which deals with the security implications of autonomous financial transactions. This connects to a more foundational question about extending the chains of trust in infrastructure firmware using Python. Finally, there was an examination of the hidden life of signals, which investigates time-domain inferences and other privacy attacks on everyday devices.
The work on Control-Token Injection Suppresses Chain-of-Thought and Defeats Reasoning-Based Oversight in Tool-Using Agents is particularly important because it directly addresses the safety concerns surrounding autonomous agents using large language models for complex reasoning. This research explored a method where injecting specific control tokens can suppress the chain of thought process, effectively defeating reasoning that relies on oversight mechanisms when agents are using tools.
This technique was tested against agent name collision attacks in multi-agent systems, showing that this token injection successfully disrupted the intended reasoning flow. Furthermore, the work on FedCoT-VQA presents a federated learning and unlearning framework designed for chain-of-thought planners in video question answering tasks. This framework aims to improve planning accuracy while maintaining privacy across distributed datasets.
The SAGEGAN paper introduces style-based anomaly detection using Gaussian embeddings within generative adversarial networks. This is significant for identifying subtle deviations in visual data. This contrasts with the MDRC work, which focuses on a deployable state-recovery defense for traffic signal control when sensors are corrupted, suggesting a focus on real-world system robustness.
The CCR paper proposes a common quality-gated CACAO integrations registry for European cybersecurity automation. This seeks to standardize and improve the quality of various integrations in this domain. Finally, ACTS evaluates large language model cipher identification under controlled blind conditions, providing insight into the security vulnerabilities of these models themselves.
The work on strengthening clean-label backdoor attacks against malware detectors is particularly important because it directly addresses the integrity of security systems that rely on machine learning for threat detection. Researchers explored RAMP, which reverses adversarial perturbations to make these backdoor attacks less effective. This involves taking the malicious input and trying to reverse the small changes made to it so that the detector can no longer be tricked by the hidden trigger.
This effort builds upon previous work concerning retrieval-augmented generation where diverse distributed poisoning was used for retrieval augmentation. This suggests a parallel interest in manipulating model inputs maliciously. Furthermore, there is ongoing investigation into how hardware fuzzing can be improved by rethinking oracles and guidance mechanisms to find specific vulnerabilities. This connects to the broader theme of improving system robustness across different domains.
A separate line of inquiry focused on cryptographic security gaps within decentralized dark pools, examining the privacy issues present in these systems. Simultaneously, efforts are underway to enhance verifiable LLM inference from models like GPT-2 up to 70 billion parameters by using sampled layerwise proofs. This verification method aims to provide a way to prove the output is correct without needing full access to the model's internal workings.
The most significant piece of work today involves the EVAGE project, which explores autonomous mechanisms for generating and adapting maximal extraction value in decentralized systems. This matters because it tackles how agents can intelligently navigate complex environments to find opportunities that others miss.
One key attempt was developing the EVAGE framework, which uses multi-agent harnessing to achieve this autonomous MEV generation and adaptation. This means the system learns how different agents should cooperate to maximize profit opportunities without constant human intervention.
Following that, there was research into extracting convolutional neural networks from unknown architectures in a setting where feedback is not available. This is important because it shows a new way to understand complex neural network structures even when you do not have the usual training signals.
Another area of focus was investigating how models leak information through residual streams during large language model operations. This work addresses a critical security concern by looking at covert information transfer that might bypass standard defenses.
Then, there is the effort to create a bulletproof method for detecting infrastructure-as-a-service offerings on Telegram. This aims to build reliable systems for identifying potentially risky services in public channels.
Safety is also being addressed through safety-aware zero trust enforcement designed for IoT and cyber-physical systems. This work focuses on building robust security protocols where every device must be constantly verified before interacting with the network.
Finally, there is the development of GUIAuditor, which enables post-hoc child safety forensics by using action-guided GUI provenance on mobile devices. This tool allows researchers to trace user actions back to the interface itself after an incident has occurred.
Today's papers
- The Role of Learning in Attacking ML-based Network Intrusion Detection The authors develop lightweight adversarial agents trained via reinforcement learning to evade intrusion detection models without needing gradient computation at deployment. [paper] [episode]
- SilentLedger: Privacy-Preserving Auditing for Blockchains with Complete Non-Interactivity This paper proposes a system that allows users to transact and auditors to audit blockchain data without needing direct interaction between them. [paper] [episode]
- Lightweight, Practical Encrypted Face Recognition with GPU Support This paper presents a lightweight face recognition system that can be deployed efficiently using GPU support. [paper] [episode]
- Optimizing watermarks for large language models This research focuses on how to best place watermarks in large language models to ensure their integrity and traceability. [paper]
- WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents This paper introduces a benchmark to test how well systems detect prompt injection attacks against web agents. [paper]
- Comparative Evaluation of Static Embedding Models for HTTP Request Anomaly Detection This study compares different static embedding models for detecting unusual HTTP requests. [paper]
- Ajar: Measuring Open Privilege in Agent Defenses This paper measures the level of open privilege that agents possess within their security defenses. [paper]
- Topological Signatures of Cyber-Attack Classes in Natural Visibility Graph Representations of Network Traffic This research uses network traffic graphs to find topological patterns that represent different types of cyber attacks. [paper]
- When Clients Are Orchestrated: Strategic Gradient Manipulation to Defeat Federated Learning Servers with Efficient Defense This paper shows how manipulating gradients can be used strategically to defeat federated learning servers. [paper]
- The Like Trap: Multi-Stage Poisoning against Agents in Similarity-based Recommendation Systems This study investigates how multi-stage poisoning attacks can target agents in recommendation systems based on similarity. [paper]
- Multi-View Fusion for Encrypted C2 Detection: A Leakage-Controlled Measurement Study of Evaluation Pitfalls This paper examines how fusing multiple views helps detect command and control communications while controlling information leakage. [paper]
- Issuer-Sovereign Agentic Payments This paper explores the concept of payments where the issuer has sovereign control over agentic transactions. [paper]
- Reliable Federated TinyML Deployment for IoT Security This work discusses reliable methods for deploying tiny machine learning models on Internet of Things devices in a federated manner. [paper]
- Enhancing Multiclass Malware Classification in Resource-Constrained Environments This research focuses on improving malware classification accuracy when resources are very limited. [paper]
- Trouble at the top: can Python extend the chains of trust in infrastructure firmware? This paper investigates whether Python can be used to extend the chain of trust within infrastructure firmware. [paper]
- The hidden life of signals: Time-domain inferences and other privacy attacks on everyday devices This study analyzes privacy risks associated with time-domain signal inferences on common electronic devices. [paper]
- FedCoT-VQA: A Federated Learning and Unlearning Framework for Chain-of-Thought Planners in VideoQA This paper introduces a framework for federated learning to improve chain-of-thought planners in video question answering tasks. [paper]
- Anti-Localization Uplink Communications in Satellite-Terrestrial Systems This research focuses on developing anti-localization techniques for communications between satellite and terrestrial systems. [paper]
- SAGEGAN: Style-Based Anomaly Detection with Gaussian Embeddings using Generative Adversarial Networks This paper proposes a style-based anomaly detection method that uses generative adversarial networks with Gaussian embeddings. [paper]
- MDRC: A Deployable State-Recovery Defense for Traffic Signal Control under Sensor Corruption This work presents a deployable defense mechanism for recovering the state of traffic signals when sensors become corrupted. [paper]
- Control-Token Injection Suppresses Chain-of-Thought and Defeats Reasoning-Based Oversight in Tool-Using Agents This paper shows that injecting control tokens can suppress reasoning and oversight in agents that use tools. [paper]
- CCR: Towards a Common, Quality-Gated CACAO Integrations Registry for European Cybersecurity Automation This paper proposes a registry to standardize and gate the quality of integrations for European cybersecurity automation. [paper]
- Agent Name Collision Attacks in Multi-Agent Systems This study investigates security vulnerabilities arising from name collisions between agents in multi-agent systems. [paper]
- ACTS: A multi-tier benchmark evaluating LLM cipher identification under controlled blind conditions This paper creates a benchmark to test how well large language models can identify ciphers under controlled, blind conditions. [paper]
- Improving Service Availability in KubeEdge-Based Architectures Using Lightweight Intrusion Detection This research focuses on using lightweight intrusion detection to improve service availability in KubeEdge architectures. [paper]
- Divide and Doubt: Diverse Distributed Poisoning for Retrieval-Augmented Generation This paper explores how diverse distributed poisoning can be used to affect retrieval-augmented generation systems. [paper]
- Solidity Meets LLMs: A Transformer-Based Approach to Smart Contract Vulnerability Detection This study proposes using a transformer model to detect vulnerabilities in smart contracts written in Solidity. [paper]
- Cryptographic Security Is Not Enough: Privacy Gaps in the Renegade Decentralized Dark Pool This paper discusses the privacy gaps that exist even when strong cryptographic security is applied to decentralized dark pools. [paper]
- SoK: You Find What You Seek: Rethinking Oracles, Guidance, and Input Generation in Hardware Fuzzing This paper rethinks how oracles, guidance, and input generation should be used in hardware fuzzing efforts. [paper]
- Seal, Then Sample: Sampled Layerwise Proofs for Verifiable LLM Inference from GPT-2 to 70B This work provides verifiable proofs for large language model inference by using sampled layerwise proofs across different sizes of models. [paper]
- Only Pay What You Must Spend: On-Demand Privacy Budget Payment for Differentially Private RAG This paper proposes an on-demand privacy budget payment system for retrieval-augmented generation systems that use differential privacy. [paper]
- RAMP: Reversing Adversarial Perturbations to Strengthen Clean-Label Backdoor Attacks against Malware Detectors This research focuses on reversing adversarial perturbations to make malware detectors more robust against clean-label backdoor attacks. [paper]
- EVAGE: Autonomous MEV Generation and Adaptation via Multi-Agent Harness This paper describes a multi-agent harness that enables autonomous mining of maximum extractable value in decentralized finance. [paper]
- Extracting CNNs in the Unknown-Architecture and Feedback-Agnostic Setting This study demonstrates how to extract convolutional neural networks even when the underlying architecture is unknown and without feedback. [paper]
- A Bulletproof Business? Towards Detecting Infrastructure-as-a-Service Offerings on Telegram This paper proposes a method for detecting infrastructure as a service offerings advertised on Telegram. [paper]
- Your Model Is Leaking: Covert Information Transfer through LLM Residual Streams This research analyzes how covert information transfer can occur through the residual streams of large language models. [paper]
- No Place to Hide: An Analysis on Protected Order Flow Sandwich Attacks This paper analyzes sandwich attacks that exploit protected order flow mechanisms in financial systems. [paper]
- Safety-Aware Zero Trust Enforcement for IoT and Cyber-Physical Systems This work proposes a zero trust enforcement strategy specifically designed for IoT and cyber-physical systems that prioritizes safety. [paper]
- GUIAuditor: Enabling Post-hoc Child Safety Forensics via Action-Guided GUI Provenance on Mobile Devices This paper enables post-hoc child safety forensics by using action-guided GUI provenance on mobile devices. [paper]
- Do Electromagnetic Side-Channel Attacks Threaten Electronic Polling Stations? Scenarios and Recommendations This study assesses the threat posed by electromagnetic side-channel attacks to electronic polling stations. [paper]
- Pinpointing Super-Quadratic Quantum Enumeration Speedups: Exact and Certified Evaluation of the Guessing-Moment Exponent under Product-Distribution Advice This paper precisely evaluates quantum enumeration speedups using exact and certified methods under product-distribution advice. [paper]
- MimicSat: A Reconfigurable Cyber-Physical Testbed For Small Satellite Systems and Cybersecurity Research This paper describes a reconfigurable testbed for small satellite systems used in cybersecurity research. [paper]
- Physalia: Redistribution-Resistant Content Protection for Decentralized Storage This work proposes content protection techniques that are resistant to redistribution in decentralized storage systems. [paper]
- A Gmail-Based Phishing Detection Prototype for Nigerian Fintech Emails Using Sender Checks and BiLSTM Classification This paper presents a prototype for detecting phishing emails using Gmail data and a Bidirectional LSTM classifier. [paper]
- Safeguarding LLM Agents against Long-Horizon Threats via Shadow Memory This research proposes shadow memory as a defense mechanism to safeguard large language model agents against long-horizon threats. [paper]
- A Non-Invasive Cloud-Based Migration Strategy for Post-Quantum Cybersecurity in Smart HVAC Systems: Architecture, Implementation, and Empirical Evaluation This paper presents a non-invasive cloud migration strategy for post-quantum cybersecurity in smart HVAC systems. [paper]
- ChronosAttack: Adversarial Tool Scheduling Attacks on LLM Agents This paper investigates adversarial scheduling attacks targeting the tools used by large language model agents. [paper]
- Shedding Light on Complex Bitcoin Mixer Transactions: 67-Fold Reduction in Unclassified Cases This study analyzes complex Bitcoin mixer transactions and achieves a significant reduction in unclassified cases. [paper]
- MixGuard: Towards Detecting and Understanding Mixer Laundering on Ethereum This paper focuses on detecting and understanding the laundering process of cryptocurrency through Ethereum mixers. [paper]
- Security and Privacy in Large-Model-Driven Embodied Agents: Attacks, Defenses, and Future Directions This paper provides a comprehensive overview of attacks, defenses, and future directions concerning security and privacy in embodied agents driven by large models. [paper]
- Agentic AI Cybersecurity Framework This paper proposes a comprehensive framework for securing agentic artificial intelligence systems. [paper]
The papers
- The Role of Learning in Attacking ML-based Network Intrusion Detection — The authors develop "lightweight adversarial agents trained via reinforcement learning (RL) that decouples the cost of learning an evasion strategy from the cost of executing it." These agents "learn offline to perturb malicious NetFlow records to evade surrogate intrusion detect [episode]
- SilentLedger: Privacy-Preserving Auditing for Blockchains with Complete Non-Interactivity — The paper "SilentLedger: Privacy-Preserving Auditing for Blockchains with Complete Non-Interactivity" proposes a transaction system designed to reconcile blockchain privacy with compliance auditing through "complete non-interactivity." The authors identify that existing auditable [episode]
- Lightweight, Practical Encrypted Face Recognition with GPU Support — "" [episode]
- FedCoT-VQA: A Federated Learning and Unlearning Framework for Chain-of-Thought Planners in VideoQA —
- Comparative Evaluation of Static Embedding Models for HTTP Request Anomaly Detection —
- ACTS: A multi-tier benchmark evaluating LLM cipher identification under controlled blind conditions —
- Ajar: Measuring Open Privilege in Agent Defenses —
- Topological Signatures of Cyber-Attack Classes in Natural Visibility Graph Representations of Network Traffic —
- Improving Service Availability in KubeEdge-Based Architectures Using Lightweight Intrusion Detection —
- Divide and Doubt: Diverse Distributed Poisoning for Retrieval-Augmented Generation —
- Solidity Meets LLMs: A Transformer-Based Approach to Smart Contract Vulnerability Detection —
- Cryptographic Security Is Not Enough: Privacy Gaps in the Renegade Decentralized Dark Pool —
- When Clients Are Orchestrated: Strategic Gradient Manipulation to Defeat Federated Learning Servers with Efficient Defense —
- The Like Trap: Multi-Stage Poisoning against Agents in Similarity-based Recommendation Systems —
- Reliable Federated TinyML Deployment for IoT Security —
- Anti-Localization Uplink Communications in Satellite-Terrestrial Systems —
- SoK: You Find What You Seek: Rethinking Oracles, Guidance, and Input Generation in Hardware Fuzzing —
- Multi-View Fusion for Encrypted C2 Detection: A Leakage-Controlled Measurement Study of Evaluation Pitfalls —
- SAGEGAN: Style-Based Anomaly Detection with Gaussian Embeddings using Generative Adversarial Networks —
- Seal, Then Sample: Sampled Layerwise Proofs for Verifiable LLM Inference from GPT-2 to 70B —
- Only Pay What You Must Spend: On-Demand Privacy Budget Payment for Differentially Private RAG —
- RAMP: Reversing Adversarial Perturbations to Strengthen Clean-Label Backdoor Attacks against Malware Detectors —
- EVAGE: Autonomous MEV Generation and Adaptation via Multi-Agent Harness —
- Extracting CNNs in the Unknown-Architecture and Feedback-Agnostic Setting —
- A Bulletproof Business? Towards Detecting Infrastructure-as-a-Service Offerings on Telegram —
- Issuer-Sovereign Agentic Payments —
- MDRC: A Deployable State-Recovery Defense for Traffic Signal Control under Sensor Corruption —
- Control-Token Injection Suppresses Chain-of-Thought and Defeats Reasoning-Based Oversight in Tool-Using Agents —
- CCR: Towards a Common, Quality-Gated CACAO Integrations Registry for European Cybersecurity Automation —
- Agent Name Collision Attacks in Multi-Agent Systems —
- Trouble at the top: can Python extend the chains of trust in infrastructure firmware? —
- The hidden life of signals: Time-domain inferences and other privacy attacks on everyday devices —
- MixGuard: Towards Detecting and Understanding Mixer Laundering on Ethereum —
- A Non-Invasive Cloud-Based Migration Strategy for Post-Quantum Cybersecurity in Smart HVAC Systems: Architecture, Implementation, and Empirical Evaluation —
- Security and Privacy in Large-Model-Driven Embodied Agents: Attacks, Defenses, and Future Directions —
- Agentic AI Cybersecurity Framework —
- ChronosAttack: Adversarial Tool Scheduling Attacks on LLM Agents —
- Shedding Light on Complex Bitcoin Mixer Transactions: 67-Fold Reduction in Unclassified Cases —
- Enhancing Multiclass Malware Classification in Resource-Constrained Environments —
- Your Model Is Leaking: Covert Information Transfer through LLM Residual Streams —
- No Place to Hide: An Analysis on Protected Order Flow Sandwich Attacks —
- Safety-Aware Zero Trust Enforcement for IoT and Cyber-Physical Systems —
- GUIAuditor: Enabling Post-hoc Child Safety Forensics via Action-Guided GUI Provenance on Mobile Devices —
- Do Electromagnetic Side-Channel Attacks Threaten Electronic Polling Stations? Scenarios and Recommendations —
- Pinpointing Super-Quadratic Quantum Enumeration Speedups: Exact and Certified Evaluation of the Guessing-Moment Exponent under Product-Distribution Advice —
- MimicSat: A Reconfigurable Cyber-Physical Testbed For Small Satellite Systems and Cybersecurity Research —
- Physalia: Redistribution-Resistant Content Protection for Decentralized Storage —
- Optimizing watermarks for large language models —
- A Gmail-Based Phishing Detection Prototype for Nigerian Fintech Emails Using Sender Checks and BiLSTM Classification —
- WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents —
- Safeguarding LLM Agents against Long-Horizon Threats via Shadow Memory —
Important terms
- Reinforcement Learning Adversarial Agents
- Lightweight agents trained via reinforcement learning to trick intrusion detection models offline, generating evasion strategies that are fast and don't need complex gradient calculations when deployed in a real network.
- Model Transfer Robustness
- The ability of learned attack strategies to maintain effectiveness when moving between different training scenarios or datasets, showing practical robustness across various network conditions.
- Control-Token Injection
- A method used to suppress the chain-of-thought process in autonomous agents using LLMs by injecting specific control tokens, effectively defeating reasoning that relies on oversight mechanisms.
- Clean-Label Backdoor Attacks
- Attacks designed to hide malicious triggers within malware detectors, and RAMP is a technique used to reverse these adversarial perturbations to make the backdoor attacks less effective.
- EVAGE Project (Maximal Extraction Value)
- Autonomous mechanisms in decentralized systems that use multi-agent harnessing to intelligently generate and adapt maximal extraction value, allowing agents to find opportunities without constant human intervention.