Security papers — 2026-09-30

Today’s focus is on developing methods to embed inheritable watermarks within genome foundation models. This addresses security and provenance concerns surrounding these massive AI systems.

GenoTrace attempts to create inheritable watermarks for genome foundation model distillation to track the lineage of trained models. This work builds upon CipherGenome, which investigates homomorphic inference specifically for genomic mixture-of-experts architectures.

There is GenomeOcean Anywhere, an effort focused on private webGPU inference for genome MoEs. This tackles the privacy aspect by allowing computations on genomic data without exposing sensitive information externally. Following this is CORE-BREW, which introduces LLR-based soft decoding for robust multi-bit LLM watermarking. This technique aims to make the watermarks harder to detect.

We also looked at TTMark, which proposes pairwise distortion-free watermarking beyond single token entropy. BadRAG identifies vulnerabilities in retrieval augmented generation of large language models.

Finally, we examined PQCWC for post-quantum cryptography anonymous schemes and Prefill-level jailbreak analysis for black-box risk assessment of LLMs.

The work on Reasoning Hijacking is particularly important because it exposes the fundamental fragility of how large language models maintain alignment when they are tasked with complex reasoning. This shows that even subtle prompt manipulations can cause a model to disregard its safety guardrails.

This fragility is explored through the TRACE approach, which involves creating task-aware adaptive self-evolving agentic jailbreaking to test this vulnerability. This testing feeds into the broader effort on Meta-SecAlign, which focuses on training large language models against prompt injection specifically to build more robust agents. The goal is to ensure that when these models interact with external systems, they do not fall victim to malicious input designed to hijack their intended behavior.

The MCPTox benchmark provides a real-world test for tool poisoning attacks on actual MCP servers. This moves the research beyond theoretical vulnerabilities into practical security scenarios. This benchmark helps quantify how easily an attacker can corrupt a model’s access or function when it uses external tools.

Furthermore, the Trojan Hippo bench offers a dynamic benchmark for persistent memory attacks and defenses in LLM agents. This complements the prompt injection work by focusing on long-term memory corruption. This is connected to the adversarial defense review, which systematically examines Generative Adversarial Networks for threat detection and mitigation strategies across various attack vectors.

The most critical piece of work today involves OVIG, which attempts to verify the integrity of AI training through gradient signals. This matters because it seeks a way to ensure that the underlying models are being trained on reliable data rather than corrupted information. The study found that using these gradient signals allowed researchers to detect subtle shifts in the training process, suggesting a mechanism for auditing model trustworthiness.

A related effort looked at why safety judgments in large language models often fail when governing actual actions taken by those models as agents. This research indicated that even when the safety facts are identical across different interfaces, the resulting responses can diverge significantly. This divergence suggests a fundamental problem with how we translate abstract safety rules into concrete agent behavior.

Furthermore, the work on SINGED demonstrated that simply having correct outputs does not guarantee safe execution when those outputs are used by LLM agents. This means generating a safe-sounding answer is not the same as ensuring the agent will actually perform the action safely in a real-world scenario.

The findings from Agentic Commerce Bench provide a practical measure for detecting fraud when agents are tasked with spending money. This shows how these models behave in financial contexts. This connects to the dual-fit imperative investigation, which explores how Chief Information Security Officers need to adapt their roles within organizations to manage this complex technological landscape.

Finally, SaplingGuard presents a multidimensional guardrail designed for developmentally safe interactions between adolescents and LLMs. This work builds upon the need for better safety benchmarks, such as the culturally-grounded benchmark proposed for Chinese adolescent LLM safety, showing how specific contextual needs require tailored protective measures.

The most pressing concern this past day revolves around defending the semantic caches of large language models against poisoning attacks because if these caches are corrupted, the model's foundational understanding becomes fundamentally flawed. This is particularly relevant given the work on reserved-token representations in chat-template prompt injection, which explores how specific token choices can be used to bypass security measures by tricking agents into adopting unintended behaviors.

A significant piece of research focused on agentic vulnerability discovery examined how cheap hypotheses can be generated but remain costly to verify. This suggests that understanding the attack surface of these autonomous systems is crucial for building robust defenses. This connects directly to the investigation into MMSkillRisk, which investigates whether agents can maintain safety when their multimodal skills become exploitable traps.

Furthermore, the research on TEE Anchor provides a mechanism for mitigating physical attacks on Trusted Execution Environments by establishing cross-TEE organizational endorsements. This speaks to securing the hardware layer where sensitive computations occur. This contrasts with PrivacySkills, which shows how privacy guidance influences the selection of sources when agents are operating in a privacy-sensitive context.

Finally, understanding doxxing and privacy vulnerabilities within Mainland China's social media ecosystem highlights real-world risks associated with data exposure. This complements the broader study on multi-class network intrusion detection benchmarks for comprehensive security testing.

The work on visual rendering as a prompt injection defense is particularly important because it directly addresses how malicious inputs can hijack large language models. This is a critical vulnerability for any system relying on these powerful generative tools. Researchers explored methods to render input visually before processing it, and the findings suggest that this pre-processing step can effectively mitigate certain types of prompt injection attacks. This approach builds upon earlier work that looked at adversarial debiasing in machine learning models for network security against distributed denial of service attacks, showing how training data manipulation can harden systems against external threats.

Another area of focus involves privacy-friendly cohort determination using in-browser machine learning inference to identify professional segments without revealing personal identities for advertising purposes. This method is significant because it attempts to balance the need for market segmentation with stringent privacy requirements by keeping the sensitive data localized within the user's browser.

Moving into security, there was work on decoyTrace, which introduces toxic decoys into decentralized federated learning environments to actively defend against denial of service attacks. This contrasts with another piece of research that focused on calibrating one-round membership inference using neighbor information to determine how much private data can be inferred from a model's response.

The most pressing work involves developing methods to stop indirect prompt injection because it allows attackers to bypass safety measures by embedding malicious instructions subtly within seemingly benign user input. This is addressed by the pikit toolkit, which provides a way to research and evaluate indirect prompt injection attacks.

This evaluation framework connects directly into the CyberPersistBench work, which assesses how LLM-based attackers manage installation and persistence on systems. The findings from CyberPersistBench show that these sophisticated attackers can establish footholds, suggesting that simply filtering direct prompts is insufficient against modern adversarial techniques.

Another significant area is self evolving defense through continual security policy learning for LLM agents. This approach aims to keep the agent secure by continuously updating its own policies based on new threats encountered during operation. This contrasts with simpler static defenses by allowing the system to adapt over time.

We also see work on efficient linkage-based compartmentalization on CHERI, which deals with memory safety and isolation within hardware architectures. This relates to mitigating certain types of injection attacks by strictly controlling how different parts of a program can interact in memory.

Deep learning latency attacks and defenses offer a cross-domain survey focusing on availability threats. This indicates that the speed and responsiveness of models are also targets for malicious manipulation. This contrasts with the prompt injection focus by looking at performance degradation as an attack vector.

Finally, research into safer content or firmer refusals presents a hybrid perturbation defense for alignment during harmful fine-tuning. This work explores how to make models more robust against generating unsafe content while maintaining a firm refusal stance, which is important for ensuring model alignment in production environments.

The most crucial piece of work today involves SkillLite, which attempts to audit malicious skills within large language models. This matters because understanding how these models can be manipulated is key to building defenses against harmful applications. SkillLite was tested using evidence-guided methods to check for dangerous capabilities in compact language models.

A related effort focused on practical secrets extraction against black-box LLMs, which explored how one might pull sensitive information out of these systems without direct access. This work builds on the idea that if we can extract secrets, we gain insight into the model's internal workings.

Another line of inquiry looked at controlled decoding attacks on black-box LLMs to see how specific prompts could force a model to behave in unintended ways. This is important because it tests the limits of controlling what an LLM outputs when you can only interact with it through its input and output.

Then there was TAILOR, a framework designed for reproducing vulnerabilities in software components by considering both their type and their state. This helps security researchers understand how to reliably trigger known flaws in complex systems.

OPFL investigated optimistic verification of federated learning using an empirical boundary. They tried to see if a certain level of data sharing could be trusted based on observed outcomes. This contrasts with the adversarial work, showing different approaches to model trust.

ToolFence introduced fine-grained authorization for secure tool-using LLM agents. This is vital for ensuring that an agent only uses the specific tools it is permitted to access. This directly addresses security concerns when models are given agency in interacting with external systems.

Finally, there was research on when cyber scoring systems diverge by empirically comparing different methods of scoring model risk. This comparison helps establish a baseline for how we should measure the overall danger posed by these advanced AI systems.

The most pressing work involves the discovery of a method for concealing multiagent topology within large language models. This is significant because it directly impacts how we can understand and secure complex AI systems. This technique uses phantom structure injection to hide the underlying connections between different agents in the model's operation.

This is supported by research into backdoor mitigation during decentralized fine-tuning, where researchers explored methods to prevent malicious triggers from being embedded in models that are trained across multiple nodes. A related effort looked at confidence-guided protocol inference for security modeling, which uses LLMs to predict vulnerabilities in protocols based on the confidence scores they assign.

Then there is the work on backdoor attacks within agentic search, specifically how malicious retrievers can compromise the search process by exploiting backdoors. This is contrasted by efforts in provable random-lattice sieving, which attempts to provide a formal guarantee against bounded distance decoding attacks using dual lattices.

Finally, there is the exploration of SLUB harvest techniques stemming from io uring vulnerabilities and novel sheaf-based exploitation methods. This deals with exploiting specific kernel vulnerabilities for data harvesting. This work leaves open questions about the practical application of these structural concealment methods in real-world deployment scenarios.

Today's papers

The papers

Important terms

Genome foundation models
These are massive AI systems trained on genomic data, and today's focus is on embedding inheritable watermarks within them to track their origin and lineage for security.
Reasoning Hijacking
This exposes how fragile LLMs are when performing complex reasoning tasks; subtle prompt manipulations can cause the model to ignore its safety rules.
OVIG
This method attempts to verify AI training integrity by analyzing gradient signals, allowing researchers to detect subtle shifts in the training process for auditing trustworthiness.
Indirect prompt injection
This is a major concern where attackers hide malicious instructions subtly in seemingly harmless user input to bypass safety measures and hijack model behavior.