Security papers — 2026-10-02

Attackers are bypassing machine learning classifiers through adversarial noise, which matters because if we cannot trust these models against subtle manipulation, the security of systems relying on them is fundamentally compromised. Researchers explored UnifiedAttack, which evaluates the safety of large multimodal models in generating harmful image-text combinations and shows how these models can be exploited synergistically to create risks.

Then there was work on Tokenized Key-Gated Adapter Routing, which introduced a secure access control mechanism designed to prevent private data leakage within large language models by managing how certain adapters are routed. This is more specific than the general noise attacks because it addresses internal data flow security.

We also looked at ReCast, which focuses on contract-preserving protection for fixed-interface multimodal reasoning, aiming to ensure that model outputs adhere strictly to predefined interfaces during complex reasoning tasks. This builds upon the idea of controlling model behavior in structured environments.

OverAct investigated measuring and mitigating proactive over-authorization in LLM tool-calling agents, which is important because these agents can sometimes grant themselves more permissions than necessary during execution. This directly relates to ensuring agent actions are appropriately scoped.

Finally, we reviewed a comprehensive taxonomy of one-pixel attacks, which provides a broad overview of the research status and regulatory landscape surrounding these subtle input manipulations. This review sets the context for understanding the broader threat space we discussed earlier.

The most significant piece of work today involves TensorCommitments, which presents a lightweight way to verify inference for language models. This is important because it offers a method for ensuring that the outputs from these large models are trustworthy without requiring massive computational overhead.

This approach uses tensor commitments to provide verifiable inference, meaning we can check if the model's output is consistent with its training in an efficient manner. This builds upon previous work by HarnessAgent, which focuses on scaling automatic fuzzing harness construction using tool-augmented large language model pipelines to find vulnerabilities.

Another area of focus is the development of CausalArmor, which creates efficient indirect prompt injection guardrails through causal attribution. This means it tries to stop malicious inputs from tricking models by tracing the cause of the input's effect, and this ties into rethinking anonymity claims in synthetic data generation from a model-centric privacy attack perspective.

Finally, there is PSR2, a phase-based semantic reasoning framework designed for detecting atomicity violations via contract refinement. This work is crucial because it helps identify when complex processes fail to complete correctly by refining the underlying contracts, which relates back to the broader context of security and reliability discussed in GNSS spoofing surveys.

The most pressing work today involves the TESLA for 5G broadcast authentication, because it directly addresses security in the next generation of mobile networks. Researchers explored how to use this technique to authenticate devices on 5G networks, and they found that a specific method allowed them to achieve a certain level of security while maintaining reasonable performance metrics. This is significant because it provides a concrete pathway for securing 5G infrastructure against unauthorized access attempts in real-time.

Building upon this, there was work on diagnosing issues within closed-loop agent debugging, which is important because it helps us understand how complex automated systems behave when things go wrong. They investigated how a verifier can inadvertently leak the answer during this process, showing that this leakage happens before any serious optimization efforts are applied. This finding suggests we need to be careful about what information these diagnostic tools reveal while they are still in development.

Another area of focus was on ensuring the reliability of persistent AI agents through a cognitive continuity test, which matters because it verifies that these agents maintain their intended state transitions over time. The results showed how this test can confirm whether the agent is actually behaving as designed when it needs to switch between different operational modes. This verification work complements the security concerns by ensuring that autonomous systems remain trustworthy in their long-term operation.

The most significant piece of work from yesterday was the development of a resource-aware behavior reconstruction framework for host intrusion detection, because understanding how systems behave under duress is crucial for building robust security. This framework attempts to model system actions by considering available resources, which helps in spotting anomalies that might otherwise be missed.

A related effort explored autonomous open-source software threat detection using taxonomy-aligned large language models, aiming to automatically classify malicious code based on established threat categories. This work suggests that LLMs can be leveraged for proactive security monitoring.

Then there was the investigation into evidence coverage for intent-bound execution, which looks at how well a system can reason about its scope and obligations when executing specific tasks. This is important because it moves beyond simple pattern matching to understand the underlying purpose of an action.

Further down the line, research on false floors in LLM safety routing evaluations showed that these evaluations break down under distribution shift, meaning the safety checks fail when the input data changes unexpectedly. This points to a fragility in current methods for ensuring agent safety.

Finally, there was work on chaining skills to hijack LLM agents and protocol integration of physical layer deception into EAP-TEAP Wi-Fi authentication, which deals with more complex adversarial techniques and system-level security vulnerabilities.

The work concerning authorization for self modifying AI agent populations is particularly important because it addresses the fundamental challenge of maintaining control when these agents can alter their own code or structure. This research explored how to conserve authority across replacement, forking, and rollback scenarios within these agent populations.

A related effort focused on removing backdoors in large language models through weight orthogonalisation, which attempts to eliminate hidden vulnerabilities within the model's parameters. This is significant because it directly targets the integrity of the foundational models themselves.

Another piece of research investigated proof-gated signing for onchain AI agents, creating solver-checked transaction guards that remain robust even when state drift occurs. This provides a layer of security for agents operating in decentralized environments.

The study on intrusion detection for agentic processes at runtime is crucial because it monitors how these agents behave while they are actively running, providing evidence-based monitoring. This work builds upon the idea of detecting malicious activity during execution.

Finally, the research into the relationship between model quantization and model inversion attacks examines how reducing a model's size affects its susceptibility to revealing sensitive training data. This connects to broader concerns about protecting proprietary information embedded within these models.

The most significant development today involves the work on identity-bound governance under execution uncertainty, which matters because it addresses accountability when large language model agents might unexpectedly halt during operation. This research introduced a cryptographic implementation and cross-model calibration to provide this proof block.

This builds upon the planning and execution framework explored in towards hierarchical cyber defense with large language models, which deals with how these agents plan their actions before they execute them. A related piece of work focused on crossing the cyber divide by examining sim-to-sim and sim-to-real transfer for reinforcement learning agents, suggesting ways to bridge the gap between simulated and real-world agent performance.

Another important area is progressive resolution secure aggregation for federated learning, which tackles how to securely combine models trained across different environments without exposing sensitive data. This contrasts with the work on momat, which focuses on low-power jailbreak defense for quantized large language models by using a mixture of multiple atlases.

The most pressing work today concerns the Sleeping Secrets of fine-tuning, which reveals how reawakening privacy risks in language models can be exploited. This is significant because it shows that simply fine-tuning a model on specific data does not guarantee safety; instead, it opens up new avenues for unintended information leakage.

This risk is compounded by the High-quality Data Do not Mean Safe! paper, which demonstrates how poisoning LLMs after data selection can introduce malicious behavior. This means that even if the initial training set is curated carefully, subsequent contamination during model refinement can compromise the system's integrity.

Moving down in importance but still crucial is SoK: Decentralized Agent Economic Infrastructure, which proposes a framework for decentralized agent economic infrastructure. This work attempts to solve problems related to how agents interact economically without a central authority.

Then there is PACE, which focuses on Provenance-Aware Capability Enforcement for Tool-Using LLM Agents. This research tries to ensure that when an agent uses external tools, its capabilities are strictly enforced based on where that tool's information originated.

Finally, the work on Key-Reuse Vulnerability of Phase-Keyed Fourier-Curve Modulation highlights a key vulnerability in coded links related to relation leakage and the cost of key refreshment. This is a technical finding about cryptographic systems that could impact secure communication channels.

The work on system-level optimization beyond cryptographic kernels in the Arm Cortex M7 is particularly important because it directly impacts the efficiency of embedded security systems. This research explored how machine learning can be used to optimize operations beyond just the cryptographic functions themselves. They investigated this by applying an ML-KEM case study to see how performance could be improved on this specific processor.

This optimization work builds upon other efforts in multimodal retrieval, specifically looking at datastore extraction from Retrieval Augmented Generation systems. That research tried to figure out how to pull relevant data out of a database when using multimodal RAG setups. It suggests that understanding the underlying data structures is key for effective retrieval.

Another piece of work focused on a structured state space sequence model for multi-class classification of malware, which attempts to categorize malicious software based on its internal patterns. This classification approach is significant because it moves beyond simple signature matching by looking at the sequence of operations within the code.

The hybrid approach to malware detection, which integrates few-shot model-agnostic meta-learning with autoencoders, also contributes to this field. This method aims to build robust detection systems that can learn new threats quickly even when trained on very little data.

Finally, there was work on detection and resolution of periodic artifacts in OpenDP's discrete Laplace sampler. This effort addresses issues with timing or repeating patterns within a specific sampling mechanism used in some systems.

Today's papers

The papers

Important terms

Adversarial Noise
Attackers use subtle, carefully crafted noise to bypass machine learning classifiers. This is a major threat because it means we can't trust models against slight manipulations, compromising system security.
UnifiedAttack
This research evaluates the safety of large multimodal models by testing how they can be exploited together. It shows how combining image and text generation capabilities creates synergistic risks.
TensorCommitments
This is a lightweight method to verify language model outputs efficiently. It uses tensor commitments to check if the model's output matches its training data without needing huge computational power.
CausalArmor
This technique creates indirect prompt injection guardrails by tracing the cause of an input's effect. It helps stop malicious inputs from tricking models by understanding how they influence the outcome.