Security papers — 2026-09-30
Today’s focus is on developing methods to embed inheritable watermarks within genome foundation models. This addresses security and provenance concerns surrounding these massive AI systems.
GenoTrace attempts to create inheritable watermarks for genome foundation model distillation to track the lineage of trained models. This work builds upon CipherGenome, which investigates homomorphic inference specifically for genomic mixture-of-experts architectures.
There is GenomeOcean Anywhere, an effort focused on private webGPU inference for genome MoEs. This tackles the privacy aspect by allowing computations on genomic data without exposing sensitive information externally. Following this is CORE-BREW, which introduces LLR-based soft decoding for robust multi-bit LLM watermarking. This technique aims to make the watermarks harder to detect.
We also looked at TTMark, which proposes pairwise distortion-free watermarking beyond single token entropy. BadRAG identifies vulnerabilities in retrieval augmented generation of large language models.
Finally, we examined PQCWC for post-quantum cryptography anonymous schemes and Prefill-level jailbreak analysis for black-box risk assessment of LLMs.
The work on Reasoning Hijacking is particularly important because it exposes the fundamental fragility of how large language models maintain alignment when they are tasked with complex reasoning. This shows that even subtle prompt manipulations can cause a model to disregard its safety guardrails.
This fragility is explored through the TRACE approach, which involves creating task-aware adaptive self-evolving agentic jailbreaking to test this vulnerability. This testing feeds into the broader effort on Meta-SecAlign, which focuses on training large language models against prompt injection specifically to build more robust agents. The goal is to ensure that when these models interact with external systems, they do not fall victim to malicious input designed to hijack their intended behavior.
The MCPTox benchmark provides a real-world test for tool poisoning attacks on actual MCP servers. This moves the research beyond theoretical vulnerabilities into practical security scenarios. This benchmark helps quantify how easily an attacker can corrupt a model’s access or function when it uses external tools.
Furthermore, the Trojan Hippo bench offers a dynamic benchmark for persistent memory attacks and defenses in LLM agents. This complements the prompt injection work by focusing on long-term memory corruption. This is connected to the adversarial defense review, which systematically examines Generative Adversarial Networks for threat detection and mitigation strategies across various attack vectors.
The most critical piece of work today involves OVIG, which attempts to verify the integrity of AI training through gradient signals. This matters because it seeks a way to ensure that the underlying models are being trained on reliable data rather than corrupted information. The study found that using these gradient signals allowed researchers to detect subtle shifts in the training process, suggesting a mechanism for auditing model trustworthiness.
A related effort looked at why safety judgments in large language models often fail when governing actual actions taken by those models as agents. This research indicated that even when the safety facts are identical across different interfaces, the resulting responses can diverge significantly. This divergence suggests a fundamental problem with how we translate abstract safety rules into concrete agent behavior.
Furthermore, the work on SINGED demonstrated that simply having correct outputs does not guarantee safe execution when those outputs are used by LLM agents. This means generating a safe-sounding answer is not the same as ensuring the agent will actually perform the action safely in a real-world scenario.
The findings from Agentic Commerce Bench provide a practical measure for detecting fraud when agents are tasked with spending money. This shows how these models behave in financial contexts. This connects to the dual-fit imperative investigation, which explores how Chief Information Security Officers need to adapt their roles within organizations to manage this complex technological landscape.
Finally, SaplingGuard presents a multidimensional guardrail designed for developmentally safe interactions between adolescents and LLMs. This work builds upon the need for better safety benchmarks, such as the culturally-grounded benchmark proposed for Chinese adolescent LLM safety, showing how specific contextual needs require tailored protective measures.
The most pressing concern this past day revolves around defending the semantic caches of large language models against poisoning attacks because if these caches are corrupted, the model's foundational understanding becomes fundamentally flawed. This is particularly relevant given the work on reserved-token representations in chat-template prompt injection, which explores how specific token choices can be used to bypass security measures by tricking agents into adopting unintended behaviors.
A significant piece of research focused on agentic vulnerability discovery examined how cheap hypotheses can be generated but remain costly to verify. This suggests that understanding the attack surface of these autonomous systems is crucial for building robust defenses. This connects directly to the investigation into MMSkillRisk, which investigates whether agents can maintain safety when their multimodal skills become exploitable traps.
Furthermore, the research on TEE Anchor provides a mechanism for mitigating physical attacks on Trusted Execution Environments by establishing cross-TEE organizational endorsements. This speaks to securing the hardware layer where sensitive computations occur. This contrasts with PrivacySkills, which shows how privacy guidance influences the selection of sources when agents are operating in a privacy-sensitive context.
Finally, understanding doxxing and privacy vulnerabilities within Mainland China's social media ecosystem highlights real-world risks associated with data exposure. This complements the broader study on multi-class network intrusion detection benchmarks for comprehensive security testing.
The work on visual rendering as a prompt injection defense is particularly important because it directly addresses how malicious inputs can hijack large language models. This is a critical vulnerability for any system relying on these powerful generative tools. Researchers explored methods to render input visually before processing it, and the findings suggest that this pre-processing step can effectively mitigate certain types of prompt injection attacks. This approach builds upon earlier work that looked at adversarial debiasing in machine learning models for network security against distributed denial of service attacks, showing how training data manipulation can harden systems against external threats.
Another area of focus involves privacy-friendly cohort determination using in-browser machine learning inference to identify professional segments without revealing personal identities for advertising purposes. This method is significant because it attempts to balance the need for market segmentation with stringent privacy requirements by keeping the sensitive data localized within the user's browser.
Moving into security, there was work on decoyTrace, which introduces toxic decoys into decentralized federated learning environments to actively defend against denial of service attacks. This contrasts with another piece of research that focused on calibrating one-round membership inference using neighbor information to determine how much private data can be inferred from a model's response.
The most pressing work involves developing methods to stop indirect prompt injection because it allows attackers to bypass safety measures by embedding malicious instructions subtly within seemingly benign user input. This is addressed by the pikit toolkit, which provides a way to research and evaluate indirect prompt injection attacks.
This evaluation framework connects directly into the CyberPersistBench work, which assesses how LLM-based attackers manage installation and persistence on systems. The findings from CyberPersistBench show that these sophisticated attackers can establish footholds, suggesting that simply filtering direct prompts is insufficient against modern adversarial techniques.
Another significant area is self evolving defense through continual security policy learning for LLM agents. This approach aims to keep the agent secure by continuously updating its own policies based on new threats encountered during operation. This contrasts with simpler static defenses by allowing the system to adapt over time.
We also see work on efficient linkage-based compartmentalization on CHERI, which deals with memory safety and isolation within hardware architectures. This relates to mitigating certain types of injection attacks by strictly controlling how different parts of a program can interact in memory.
Deep learning latency attacks and defenses offer a cross-domain survey focusing on availability threats. This indicates that the speed and responsiveness of models are also targets for malicious manipulation. This contrasts with the prompt injection focus by looking at performance degradation as an attack vector.
Finally, research into safer content or firmer refusals presents a hybrid perturbation defense for alignment during harmful fine-tuning. This work explores how to make models more robust against generating unsafe content while maintaining a firm refusal stance, which is important for ensuring model alignment in production environments.
The most crucial piece of work today involves SkillLite, which attempts to audit malicious skills within large language models. This matters because understanding how these models can be manipulated is key to building defenses against harmful applications. SkillLite was tested using evidence-guided methods to check for dangerous capabilities in compact language models.
A related effort focused on practical secrets extraction against black-box LLMs, which explored how one might pull sensitive information out of these systems without direct access. This work builds on the idea that if we can extract secrets, we gain insight into the model's internal workings.
Another line of inquiry looked at controlled decoding attacks on black-box LLMs to see how specific prompts could force a model to behave in unintended ways. This is important because it tests the limits of controlling what an LLM outputs when you can only interact with it through its input and output.
Then there was TAILOR, a framework designed for reproducing vulnerabilities in software components by considering both their type and their state. This helps security researchers understand how to reliably trigger known flaws in complex systems.
OPFL investigated optimistic verification of federated learning using an empirical boundary. They tried to see if a certain level of data sharing could be trusted based on observed outcomes. This contrasts with the adversarial work, showing different approaches to model trust.
ToolFence introduced fine-grained authorization for secure tool-using LLM agents. This is vital for ensuring that an agent only uses the specific tools it is permitted to access. This directly addresses security concerns when models are given agency in interacting with external systems.
Finally, there was research on when cyber scoring systems diverge by empirically comparing different methods of scoring model risk. This comparison helps establish a baseline for how we should measure the overall danger posed by these advanced AI systems.
The most pressing work involves the discovery of a method for concealing multiagent topology within large language models. This is significant because it directly impacts how we can understand and secure complex AI systems. This technique uses phantom structure injection to hide the underlying connections between different agents in the model's operation.
This is supported by research into backdoor mitigation during decentralized fine-tuning, where researchers explored methods to prevent malicious triggers from being embedded in models that are trained across multiple nodes. A related effort looked at confidence-guided protocol inference for security modeling, which uses LLMs to predict vulnerabilities in protocols based on the confidence scores they assign.
Then there is the work on backdoor attacks within agentic search, specifically how malicious retrievers can compromise the search process by exploiting backdoors. This is contrasted by efforts in provable random-lattice sieving, which attempts to provide a formal guarantee against bounded distance decoding attacks using dual lattices.
Finally, there is the exploration of SLUB harvest techniques stemming from io uring vulnerabilities and novel sheaf-based exploitation methods. This deals with exploiting specific kernel vulnerabilities for data harvesting. This work leaves open questions about the practical application of these structural concealment methods in real-world deployment scenarios.
Today's papers
- GenoTrace: Inheritable Watermarks for Genome Foundation Model Distillation. [paper]
- GenomeOcean Anywhere: Private WebGPU Inference for Genome MoEs. [paper]
- CipherGenome: Homomorphic Inference for Genomic Mixture-of-Experts. [paper]
- BadRAG: Identifying Vulnerabilities in Retrieval Augmented Generation of Large Language Models. [paper] [episode]
- CORE-BREW: LLR-Based Soft Decoding for Robust Multi-Bit LLM Watermarking. [paper] [episode]
- TTMark: Pairwise Distortion-Free Watermarking Beyond Single-Token Entropy. [paper]
- Post-Quantum Cryptography Anonymous Scheme -- PQCWC: Post-Quantum Cryptography Winternitz-Chen. [paper] [episode]
- Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models. [paper] [episode]
- Meta-SecAlign: Training LLMs against Prompt Injection for Robust Agents. [paper] [episode]
- MCPTox: A Benchmark for Tool Poisoning Attack on Real-World MCP Servers. [paper] [episode]
- Adversarial Defense in Cybersecurity: A Systematic Review of GANs for Threat Detection and Mitigation. [paper] [episode]
- NonTextual Target Attack. [paper] [episode]
- Reasoning Hijacking: The Fragility of Reasoning Alignment in Large Language Models. [paper] [episode]
- Benign Fine-Tuning Breaks Safety Alignment in Audio LLMs. [paper] [episode]
- Trojan Hippo Bench: A Dynamic Benchmark for Persistent Memory Attacks and Defenses in LLM Agents. [paper] [episode]
- TRACE: Task-Aware Adaptive Self-Evolving Agentic Jailbreaking. [paper] [episode]
- OVIG: Optimistic Verification of AI Training Integrity via Gradient Signals. [paper] [episode]
- Says Block, Still Acts: Why LLM Safety Judgments Fail to Govern Action in LLM Agents. [paper]
- SameFact: The Same Safety Facts Lead to Different Responses Across Interfaces. [paper]
- Agentic Commerce Bench: Measuring Fraud Detection for Agents That Spend Money. [paper]
- Dual-Fit Imperative in Security Leadership: A Grounded Theory Investigation of CISO Role Enactment in Modern Organisations. [paper]
- SINGED: Correct Outputs Do Not Certify Safe Execution in LLM Agents. [paper]
- SaplingGuard: A Multidimensional-Profile-Aware Multi-Agent Guardrail for Developmentally Safe Adolescent-LLM Interaction. [paper]
- Raising the Bar for Chinese Adolescent LLM Safety: A Culturally-Grounded, Fine-Grained Benchmark. [paper]
- Similarity Is Not Validity: Defending LLM Semantic Caches Against Poisoning. [paper]
- Cheap to Hypothesize, Costly to Verify: The Defense Surface of Agentic Vulnerability Discovery. [paper]
- MMSkillRisk: Can Agents Stay Safe When Multimodal Skills Become Traps?. [paper]
- TEE Anchor: Cross-TEE Organizational Endorsement for Mitigating TEE Physical Attacks. [paper]
- Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection. [paper]
- PrivacySkills: How Privacy Guidance Shapes Source Selection in LLM Agents. [paper]
- When Privacy Becomes a Weapon: Understanding Doxxing and Privacy Vulnerabilities in Mainland China's Social Media Ecosystem. [paper]
- Multi-Class, Multi-Tier Network Intrusion Detection: A Comprehensive and Reproducible Benchmark. [paper]
- Render Before Reading: Visual Rendering as a Prompt Injection Defense. [paper]
- Privacy-Friendly Cohort Determination: Sealed, CSP-Independent In-Browser ML Inference of Professional Segments for Identity-Less Advertising. [paper]
- Adversarial Debiasing of Machine Learning Models for Enhanced Network Security against DDoS Attacks. [paper]
- DecoyTrace: Toxic Decoys for Active Defense in Decentralized Federated Learning. [paper]
- Calibrating One-Round Membership Inference with Neighbors. [paper]
- Audience-Bound Persistent Memory: Authorization Across the Memory Lifecycle. [paper]
- Quantization Enables Private Dense Retrieval against Malicious Service Providers. [paper]
- Know the Normal, Track the Attack: Context-Grounded and Stateful LLM Investigation over System Provenance. [paper]
- CounterSteer: Suppressing Indirect Prompt Injection with Activation Steering. [paper]
- CyberPersistBench: Evaluating LLM-Based Cyber Attackers on Installation and Persistence. [paper]
- Self-Evolving Defense: Continual Security Policy Learning for LLM Agents. [paper]
- Efficient Linkage-Based Compartmentalization on CHERI. [paper]
- Deep Learning Latency Attacks and Defenses: A Cross-Domain Survey of Availability Threats. [paper]
- pikit: A Composable Toolkit for Indirect Prompt Injection Research and Evaluation. [paper]
- Does the Unsafe Gradient Survive a Conversation? On the Fragility of Gradient-Based Jailbreak Detection in Multi-Turn Dialogue. [paper]
- Safer Content or Firmer Refusals? A Hybrid Perturbation Defense for Alignment under Harmful Fine-tuning. [paper]
- SKILLLITE: Evidence-Guided Malicious Skill Auditing with Compact LLMs. [paper]
- Practical Secrets Extraction against Black-box LLMs. [paper]
- Controlled Decoding Attacks on Black-Box LLMs. [paper]
- One Pipeline Does Not Fit All: TAILOR, a Type- and State-Aware Framework for CVE Reproduction. [paper]
- OPFL: Optimistic Verification of Federated Learning via Empirical Boundary. [paper]
- ToolFence: Fine-Grained Authorization for Secure Tool-Using LLM Agents. [paper]
- When Cyber Scoring Systems Diverge: An Empirical Comparison. [paper]
- Beyond Semantic Narrowing: Robust and Efficient LLM Watermarking with Hamming Neighborhoods. [paper]
- AutoMark: Enabling Autoresearch to Discover Better LLM Watermarks. [paper]
- Backdoor Mitigation in Decentralized LLM Fine-Tuning. [paper]
- Confidence-Guided Protocol IR for LLM-Aided Security Protocol Modeling. [paper]
- Backdoor in the Loop: Compromising Agentic Search via Malicious Retrievers. [paper]
The papers
- Meta-SecAlign: Training LLMs against Prompt Injection for Robust Agents — META SECALIGN: A Secure Foundation LLM Against Prompt Injection Attacks Prompt injection attacks, where untrusted data contains an injected prompt to manipulate the system, have been listed as the top security threat to LLM-integrated applications. [episode]
- NonTextual Target Attack — Existing gradient-based jailbreak attacks on Large Language Models (LLMs) typically optimize adversarial suffixes to align the LLM output with predefined target responses, but restricting this objective inherently constrains the adversarial search space, limiting overall attack e [episode]
- CORE-BREW: LLR-Based Soft Decoding for Robust Multi-Bit LLM Watermarking — As a fastidious and diligent AI researcher, I have meticulously analyzed both provided texts regarding "CORE-BREW." The goal is to synthesize these descriptions into a single, comprehensive, long, and detailed summary that accurately reflects the technical depth of the work while [episode]
- Trojan Hippo Bench: A Dynamic Benchmark for Persistent Memory Attacks and Defenses in LLM Agents — Trojan Hippo is a class of persistent memory attacks that operates in a more realistic threat model than prior memory poisoning work: "the attacker plants a dormant payload into an agent’s long-term memory via a single untrusted tool call (e.g., a crafted email), which activate [episode]
- Reasoning Hijacking: The Fragility of Reasoning Alignment in Large Language Models — Current LLM safety research predominantly focuses on mitigating Goal Hijacking, preventing attackers from redirecting a model’s high-level objective (e.g., from “summarizing emails” to “phishing users”). [episode]
- Benign Fine-Tuning Breaks Safety Alignment in Audio LLMs — Prior work shows that fine-tuning aligned models on benign data degrades safety in text and vision modalities, and that proximity to harmful content in representation space predicts which samples cause the most damage. [episode]
- OVIG: Optimistic Verification of AI Training Integrity via Gradient Signals — The rapid growth of AI has increased the demand for domain-specific post-training, while cost and specialization of accelerator infrastructure push many model owners to outsource this process. [episode]
- Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models — Abstract—Warning: this paper includes examples that may be offensive or harmful. Large Language Models face security threats from jailbreak attacks. [episode]
- TRACE: Task-Aware Adaptive Self-Evolving Agentic Jailbreaking — The rise of LLM agents introduces a new threat by enabling planning, coding, and even end-toend execution of expert-level attack workflows. [episode]
- BadRAG: Identifying Vulnerabilities in Retrieval Augmented Generation of Large Language Models — Retrieval-Augmented Generation (RAG) systems, which combine external data retrieval with large language models, introduce new security risks because their databases are often sourced from public data, making them susceptible to poisoning attacks. [episode]
- MCPTox: A Benchmark for Tool Poisoning Attack on Real-World MCP Servers — By providing a standardized interface for LLM agents to interact with external tools, the Model Context Protocol (MCP) is quickly becoming a cornerstone of the modern autonomous agent ecosystem. However, it creates novel attack surfaces due to untrusted external tools. [episode]
- Adversarial Defense in Cybersecurity: A Systematic Review of GANs for Threat Detection and Mitigation — Machine learning-based cybersecurity systems are highly vulnerable to adversarial attacks, while Generative Adversarial Networks (GANs) act as both powerful attack enablers and promising defenses. [episode]
- Post-Quantum Cryptography Anonymous Scheme -- PQCWC: Post-Quantum Cryptography Winternitz-Chen — "Due to quantum computing technology becoming mature, it will threaten the security of current mainstream asymmetric cryptography methods (including RSA cryptography and Elliptic Curve Cryptography). [episode]
- MMSkillRisk: Can Agents Stay Safe When Multimodal Skills Become Traps? —
- TEE Anchor: Cross-TEE Organizational Endorsement for Mitigating TEE Physical Attacks —
- Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection —
- PrivacySkills: How Privacy Guidance Shapes Source Selection in LLM Agents —
- When Privacy Becomes a Weapon: Understanding Doxxing and Privacy Vulnerabilities in Mainland China's Social Media Ecosystem —
- Multi-Class, Multi-Tier Network Intrusion Detection: A Comprehensive and Reproducible Benchmark —
- Render Before Reading: Visual Rendering as a Prompt Injection Defense —
- Privacy-Friendly Cohort Determination: Sealed, CSP-Independent In-Browser ML Inference of Professional Segments for Identity-Less Advertising —
- Adversarial Debiasing of Machine Learning Models for Enhanced Network Security against DDoS Attacks —
- DecoyTrace: Toxic Decoys for Active Defense in Decentralized Federated Learning —
- Calibrating One-Round Membership Inference with Neighbors —
- TTMark: Pairwise Distortion-Free Watermarking Beyond Single-Token Entropy —
- Audience-Bound Persistent Memory: Authorization Across the Memory Lifecycle —
- Quantization Enables Private Dense Retrieval against Malicious Service Providers —
- Know the Normal, Track the Attack: Context-Grounded and Stateful LLM Investigation over System Provenance —
- CounterSteer: Suppressing Indirect Prompt Injection with Activation Steering —
- CyberPersistBench: Evaluating LLM-Based Cyber Attackers on Installation and Persistence —
- Self-Evolving Defense: Continual Security Policy Learning for LLM Agents —
- Efficient Linkage-Based Compartmentalization on CHERI —
- Deep Learning Latency Attacks and Defenses: A Cross-Domain Survey of Availability Threats —
- pikit: A Composable Toolkit for Indirect Prompt Injection Research and Evaluation —
- Does the Unsafe Gradient Survive a Conversation? On the Fragility of Gradient-Based Jailbreak Detection in Multi-Turn Dialogue —
- Safer Content or Firmer Refusals? A Hybrid Perturbation Defense for Alignment under Harmful Fine-tuning —
- SKILLLITE: Evidence-Guided Malicious Skill Auditing with Compact LLMs —
- Practical Secrets Extraction against Black-box LLMs —
- Controlled Decoding Attacks on Black-Box LLMs —
- One Pipeline Does Not Fit All: TAILOR, a Type- and State-Aware Framework for CVE Reproduction —
- OPFL: Optimistic Verification of Federated Learning via Empirical Boundary —
- ToolFence: Fine-Grained Authorization for Secure Tool-Using LLM Agents —
- When Cyber Scoring Systems Diverge: An Empirical Comparison —
- Beyond Semantic Narrowing: Robust and Efficient LLM Watermarking with Hamming Neighborhoods —
- AutoMark: Enabling Autoresearch to Discover Better LLM Watermarks —
- Backdoor Mitigation in Decentralized LLM Fine-Tuning —
- Confidence-Guided Protocol IR for LLM-Aided Security Protocol Modeling —
- Backdoor in the Loop: Compromising Agentic Search via Malicious Retrievers —
- Smooth Sailing through Spherical Shells: Provable Random-Lattice Sieving in Time 2 0.292n —
- Dual lattice attacks for bounded distance decoding, revisited —
- Concealing LLM-Based Multi-Agent Topology via Phantom Structure Injection —
- Harvest Season for SLUB: From io uring vulnerability to Novel Sheaf-Based Exploitation Techniques —
- She Spoofed Sea Ships by the Sea Shore: Measuring Large-Scale GPS Spoofing in Global Maritime Traffic —
- Where Do LLMs Decide to Break the Rules? Mechanistic Localization of Prompt Injection Compliance —
- Lights, Camera, Attack: Exploiting Temporal HDR Fusion with Pulsed Light —
- Selective Channel Restoration for Backdoored Vision-Language Models —
- Making Duplicate Reimbursement Unrepresentable: A Verified Ethereum E-Invoice System for Humans and AI Agents —
- Dagger: Decoupling-based Model Stealing Attack against Graph Neural Networks —
- A Function-level Dataset of Vulnerable and Fixed Source Code in JavaScript and TypeScript —
- Says Block, Still Acts: Why LLM Safety Judgments Fail to Govern Action in LLM Agents —
- SameFact: The Same Safety Facts Lead to Different Responses Across Interfaces —
- GenoTrace: Inheritable Watermarks for Genome Foundation Model Distillation —
- GenomeOcean Anywhere: Private WebGPU Inference for Genome MoEs —
- CipherGenome: Homomorphic Inference for Genomic Mixture-of-Experts —
- Agentic Commerce Bench: Measuring Fraud Detection for Agents That Spend Money —
- Dual-Fit Imperative in Security Leadership: A Grounded Theory Investigation of CISO Role Enactment in Modern Organisations —
- SINGED: Correct Outputs Do Not Certify Safe Execution in LLM Agents —
- SaplingGuard: A Multidimensional-Profile-Aware Multi-Agent Guardrail for Developmentally Safe Adolescent-LLM Interaction —
- Raising the Bar for Chinese Adolescent LLM Safety: A Culturally-Grounded, Fine-Grained Benchmark —
- Similarity Is Not Validity: Defending LLM Semantic Caches Against Poisoning —
- Cheap to Hypothesize, Costly to Verify: The Defense Surface of Agentic Vulnerability Discovery —
Important terms
- Genome foundation models
- These are massive AI systems trained on genomic data, and today's focus is on embedding inheritable watermarks within them to track their origin and lineage for security.
- Reasoning Hijacking
- This exposes how fragile LLMs are when performing complex reasoning tasks; subtle prompt manipulations can cause the model to ignore its safety rules.
- OVIG
- This method attempts to verify AI training integrity by analyzing gradient signals, allowing researchers to detect subtle shifts in the training process for auditing trustworthiness.
- Indirect prompt injection
- This is a major concern where attackers hide malicious instructions subtly in seemingly harmless user input to bypass safety measures and hijack model behavior.