MOMAT: Mixture of Multiple Atlases for Low-Power Jailbreak Defense of Quantized LLMs
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: Today's paper: "MOMAT: Mixture of Multiple Atlases for Low-Power Jailbreak Defense of Quantized LLMs".
Elias: Quantized large language models (qLLMs) are increasingly deployed on edge devices for their low latency and energy efficiency, but model quantization weakens alignment safeguards,
Nadia: First, who's behind it and why it matters.
Title and authors: Nadia: So, we're diving into "MOMAT: Mixture of Multiple Atlases for Low-Power Jailbreak Defense of Quantized LLMs." This paper tackles the big problem that when you shrink a large language model down to run on smaller edge devices, the safety guardrails that keep it aligned get weaker. It focuses specifically on how model quantization can open up these quantized large language models to jailbreak attacks.
Elias: I'm interested in the authors right away; seeing who is tackling this specific vulnerability in the context of deployment is important for understanding the real-world applicability of this research. What do you think about how they framed this as a problem needing a new approach?
Priya: From my side, I’m curious about what data they actually used to show that quantization specifically degrades the representational fidelity needed for safety layers, because I want to know if their findings hold up under real-world measurement conditions.
Nadia: Exactly, Priya. The core idea is that aggressive quantization messes with the feature geometry of the embedding space, which makes it harder for a model to tell a harmful prompt from a benign one when you're pushing those models onto edge hardware like mobile chips.
Elias: And that's where I get intrigued; if quantization flattens those safety-critical gradients in deeper transformer layers, it sounds like the very mechanisms meant to distinguish intent become less effective. I wonder what kind of numerical precision issues they found most detrimental to those distinctions.
Priya: Looking at the context, it seems like they are pointing out that weight rounding errors and activation outlier clipping directly compress semantic distances, which is a measurable way that fidelity drops during quantization. That compression is what makes separating harmful from benign intent so difficult for the safety layers to do effectively.
Nadia: Right, and this leads us into the core of their proposed solution: the MOMAT framework, which is essentially a hardware-enhanced safety system designed specifically for low-power edge deployment. It's about moving beyond those traditional RAG approaches that struggle with large, diverse safety databases.
Elias: A hardware-enhanced approach sounds promising because it addresses the latency and energy concerns inherent in running complex safety checks on resource-constrained devices where we already have so much to contend with. What kind of architectural components are they relying on for this enhancement?
Priya: The paper suggests a few key architectural shifts, like using Compute-in-Memory arrays to handle similarity searches directly in the memory, which sounds like a direct answer to the energy efficiency challenge. I’m hoping their data will confirm that this hardware acceleration provides the speedup they claim for low-power queries.
Nadia: They build this system around a few modules, including a CiM retrieval engine and a lightweight Mixture of Experts detector. This structure is designed to handle the complexity of safety checks without bogging down the inference process on edge hardware. How does this modular approach solve the issue they identified with large, heterogeneous safety databases?
Title and authors: Elias: By organizing those samples into semantically coherent subspaces called atlases, they are essentially tackling that curse of dimensionality by confining the safety processing to appropriate domains rather than searching a single massive database indiscriminately. That seems like a smart way to manage complexity.
Priya: The offline construction process where they ICL-generate cross-domain harmful and benign pairs and then recursively fan them out into dense variant sets sounds like a robust way to build these atlases without having to manually label every single edge case from scratch, which is a real practical concern for any safety system.
Nadia: That pipeline is crucial because it’s how they create those high-density neighborhoods that reduce retrieval noise, allowing the subsequent search to be much more precise when an actual prompt comes in. It's about creating structure where there was previously just a flat sea of data points.
Elias: And then you have this Mixture-of-Experts architecture where each expert specializes in a particular neighborhood signature, which suggests they are trying to route the query very specifically based on its embedding and precomputed atlas features. That routing mechanism sounds sophisticated for handling the diverse query neighborhoods they described.
Priya: The way they construct that rich feature vector, incorporating raw similarity distributions and statistical summaries per atlas, tells me that the system isn't just doing a simple lookup; it's actually calculating something about the context of the query relative to all known safety patterns in a very detailed manner.
Nadia: And finally, when harmful intent is detected, they have this targeted second pass retrieval mechanism that pulls directly from the most activated atlas to surface a safe response template. That’s how they move from detection to action quickly without needing another heavy generative step at inference time.
Elias: The efficiency gains cited are quite substantial; they report a "four point six nine × one hundred six times speedup" and a reduction in energy consumption of approximately "two point five × one hundred five times" over DRAM-based baselines for the same workload, which is significant when you think about battery life on an edge device. That level of hardware optimization is what makes this viable for deployment.
Priya: It’s interesting because they also report that MOMAT achieves "zero point zero percent ASR across both datasets" for Llama2-7B and Mistral-7B under W4A8 quantization, while maintaining the same False Refusal Rate as the baseline defense, meaning it avoids any extra overrefusal of benign queries. That balance between safety enforcement and utility preservation is something I find very compelling.
Nadia: That zero attack success rate across those specific models under that quantization level shows that their method works reliably even when the underlying model is significantly compressed, which was the central challenge they set out to solve with MOMAT: restoring alignment when quantization weakens safeguards.
Title and authors: Elias: So, looking at the overall implications of this paper on the field, it seems like a major step toward making robust safety defenses practical for resource-limited environments where we can't afford massive safety overhead. It shifts the focus from just having large models to having efficient mechanisms to secure those models at deployment.
Priya: I think the real impact is in showing that co-designing the retrieval structure with the underlying hardware substrate is what really unlocks this level of efficiency and safety, rather than just slapping a standard RAG layer on top. That modularity seems key for future work in this area.
Nadia: Exactly, Priya. The implication here is that we don't have to sacrifice alignment when deploying quantized models; we just need the right structure to handle the complexity locally and efficiently, which is what MOMAT demonstrates by using those CiM arrays and atlases.
Elias: I agree; it suggests that for edge-deployed AI, the future of safety isn't just about bigger safety models, but about creating these highly specialized, low-power retrieval mechanisms that understand the specific constraints of the quantized environment.
Priya: It’s exciting to see how this moves us toward systems where fine-grained intent detection is possible without needing a massive central cloud infrastructure for every single query decision.
Nadia: We've seen that MOMAT uses a structured knowledge retrieval method, specifically leveraging atlases and hardware acceleration to mitigate the issues caused by quantization on edge devices. This entire research effort is focused on making safety defenses practical and efficient for quantized large language models by decoupling prompt encoding from the protected LLM through a lightweight encoder, then using a CiM-accelerated RAG engine to perform real-time multi-atlas similarity searches and MoE scoring.
Elias: The main takeaway here is that they’ve engineered a system where the structure of safety knowledge—the atlases—is tailored to the hardware capabilities via Compute-in-Memory, leading to substantial speedups and energy reductions compared to traditional vector search methods.
Priya: From a measurement standpoint, it's clear that this approach successfully maintains high alignment robustness against jailbreaks at zero attack success rate across specific benchmarks when using W4A8 quantization for models like Llama2-7B and Mistral-7B.
Nadia: So, to wrap up the discussion on MOMAT: they’ve shown that by restructuring retrieval into semantically coherent atlases and executing similarity search on a CiM-accelerated engine, they can restore alignment behaviors sacrificed by quantization.
Elias: It proves that modular defenses can make edge-deployed qLLMs both safer and more energy-efficient through this careful co-design of the retrieval structure and the underlying hardware substrate.
Priya: It’s a significant development because it shows that strong safety enforcement and benign utility preservation aren't mutually exclusive when you use this kind of structured approach.
Nadia: That’s our discussion on MOMAT: a framework designed to tackle the challenges of jailbreak vulnerability in quantized models using hardware acceleration and domain-localized knowledge structures. We’ll be looking at other papers soon, but for now, that covers what we have on this one.
The paper's summary: Nadia: So, to recap, MOMAT is this safety framework that uses structured knowledge in multiple domains—the atlases—combined with super efficient hardware acceleration to protect quantized AI from jailbreaks on edge devices.
Elias: That's right; it’s essentially a way of organizing the safety data so the system doesn't have to sift through everything randomly when a prompt comes in, which is a huge structural improvement over flat RAG systems.
Priya: From a measurement standpoint, I find that their focus on reducing retrieval noise by isolating harmful and benign samples into these specific subspaces seems like the most critical part for maintaining high utility while keeping safety tight.
Nadia: Exactly; the whole point is that when you compress a model with quantization, its internal logic gets muddied, and MOMAT tries to compensate by using this structured lookup system to catch those subtle shifts in intent before they escalate into a harmful response.
Elias: And that structure relies heavily on the Mixture-of-Experts component, which acts like a sophisticated router, directing the query’s embedding to the most relevant safety expertise based on what’s already precomputed about the atlases.
Priya: The data they present is very compelling because it shows that this method doesn't just block everything; it manages to keep the False Refusal Rate identical to a baseline defense while maintaining zero attack success across specific quantized models. That suggests a very good balance between safety enforcement and allowing benign tasks to complete successfully.
Nadia: It’s exciting because if this works reliably on resource-constrained hardware, we could see AI systems deployed in fields like industrial monitoring or remote healthcare get real safety guards without needing massive cloud infrastructure constantly running checks.
Elias: The implication for the world is that it pushes the idea of "on-device" safety from a theoretical challenge to a practical engineering problem, provided you design the retrieval structure to match the hardware constraints.
Priya: I think the biggest impact is showing that we can actually maintain strong alignment guarantees even when we have to make trade-offs in model size and precision for deployment on edge devices.
Nadia: We've seen how this moves beyond just building bigger models; it's about building smarter, more efficient ways to secure the AI we already have deployed everywhere.
Elias: So, moving forward, I wonder if the reliance on that specific CiM hardware architecture means these atlases are inherently tied to certain types of memory structures?
Priya: That’s a good question; it suggests that future work might need to look at how different memory substrates could influence the optimal atlas construction strategy for maximizing safety coverage.
Nadia: We should definitely keep an eye on that; understanding the hardware substrate is key if we want this kind of localized defense to scale up across different types of edge processors.
The paper's improvements: Tom: So, we're talking about how MOMAT suggests we can take this framework and make it even better by adding more specific technical enhancements for real-world application.
Nadia: It seems like they're suggesting we really lean into that offline atlas construction pipeline, which automates the creation of those safety knowledge bases using in-context learning to generate high-density neighborhood samples.
Elias: That pipeline is smart because it tackles the manual effort of creating these domains by letting AI generate the dense variants themselves, which helps reduce human bias in what gets included in those atlases.
Priya: From a privacy researcher’s view, I’m interested in how this automated generation process affects the data we use to build these safety stores; are there concerns about inadvertently creating patterns that could leak information from the training set?
Nadia: That's a valid concern; we need to scrutinize those generated pairs closely, but the benefit is that it allows us to create far more comprehensive and detailed safety knowledge than manual curation ever could.
Elias: And then there’s the MoE scoring mechanism, where they propose constructing a feature vector that includes raw similarity distributions and statistical summaries per atlas for each expert. That level of detail in the input to the router sounds like it’s what allows for such fine-grained routing decisions.
Priya: I see that; incorporating those statistical summaries means we aren't just looking at the prompt embedding, but also how similar that prompt is across *all* domains simultaneously, which should give us a much richer safety decision than before.
Nadia: It’s about moving from simple detection to a nuanced understanding of the query's position within the entire safety knowledge landscape, ensuring we catch subtle jailbreak patterns.
Elias: And then there’s the suggestion for a targeted second pass retrieval when harm is detected, which minimizes inference time by only querying the most relevant atlas directly in memory. That’s a great way to keep that low-power requirement from getting blown out by complex safety checks during live operation.
Priya: The efficiency aspect of that targeted retrieval is key; it means the system doesn't waste power searching irrelevant data once it knows exactly where to look, which aligns perfectly with the low-energy goal we talked about earlier.
Nadia: It really shows how these modular improvements work together: better offline training, smarter routing, and faster on-device retrieval all feeding into a single high-performance defense layer.
Elias: If we could analyze the gating network that produces those expert weights more deeply, I wonder if there are specific parameter settings or data distributions that cause the MoE to become overly specialized in a way that might introduce new vulnerabilities.
Priya: That’s deep theoretical work; we need to check if any of these optimizations inadvertently create new types of adversarial inputs that exploit the structure itself rather than just the model's core weaknesses.
Nadia: It points toward future work focusing on adversarial testing specifically against the routing mechanism, which is a novel part of this defense.
Elias: Right, so we’ve seen how they improve the defense layer through better data generation and smarter routing mechanisms; now we need to look at the robustness of that routing itself.
Conclusion: Nadia: So, to wrap up this discussion on MOMAT: Mixture of Multiple Atlases for Low-Power Jailbreak Defense of Quantized LLMs, we've seen how this framework uses structured knowledge and hardware acceleration to make quantized AI much safer on edge devices.
Elias: We’ve established that the core strength lies in the combination of those domain-localized atlases and the Compute-in-Memory engine for fast retrieval.
Priya: I think what really stands out is how it balances safety enforcement with maintaining utility, given that it achieves zero attack success while keeping the False Refusal Rate consistent with baseline defenses.
Nadia: That balance is exactly what makes this paper significant; we’re seeing a path toward robust security for AI deployed everywhere, not just in massive data centers.
Elias: It confirms that the assumptions about how prompt embeddings are compressed during quantization can be mitigated through this specific architectural design rather than just hoping the model handles it.
Priya: I still have to emphasize that the success hinges on those offline construction steps; if those atlases aren't built with enough diversity, they might not catch every new attack vector we see in the wild.
Nadia: That’s a fair caveat; future work needs to focus heavily on how to keep those atlas generation pipelines updated as new jailbreak techniques emerge.
Elias: Indeed, and we should also look into whether the MoE gating network could be made more robust against adversarial perturbations in the expert weights themselves.
Priya: It really shows that for edge AI, co-designing the retrieval structure with the hardware substrate is what unlocks this level of energy efficiency and safety without needing to scale up massive cloud defenses.
Nadia: So, it’s a powerful demonstration that we can build specialized safety mechanisms for quantized LLMs that are both fast and energy-conscious.
Elias: We’ve seen how MOMAT successfully addresses the issues of low latency and high energy consumption inherent in current safety guardrails for edge AI.
Priya: Ultimately, it sets a very practical standard for how we should think about deploying complex AI systems onto resource-limited hardware while keeping them aligned.
Nadia: That’s our final look at MOMAT, showing us a tangible way to improve safety in quantized models through structured retrieval and CiM acceleration.
Elias: We’ll keep an eye on those suggestions for future work regarding adversarial testing of the MoE routing mechanism.
Boyang Lia, Bingyu Shenb, Weihao Honga, Zhiyuan Jianga, Xinlei Guana, Yan Maa, Miles Q. Lid, Yi Shenge and Ruiyang Qinc
Department of Computer Science, Kean University · Department of Computer Science and Engineering, University of Notre Dame · Department of Computing Sciences, Villanova University · McGill University · Department of Computer Science and Engineering, University of South Florida
cs.CR, cs.AI
Submitted: 2026-10-01
Updated: 2026-10-01
Comments: 16 pages, 13 figures
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 92/100
The gist: Quantized large language models (qLLMs) are increasingly deployed on edge devices for their low latency and energy efficiency, but model quantization weakens alignment safeguards, leaving qLLMs
Key concepts
- Atlases
- Instead of one large safety database, MOMAT organizes harmful and benign samples into multiple semantically coherent subspaces called atlases. This structure helps reduce retrieval noise by confining safety processing to specific domains, making the search more accurate.
- CiM (Compute-in-Memory)
- This is a hardware technique that performs similarity searches directly within the memory arrays of the storage device instead of relying on slow data transfers to external DRAM. This allows for extremely fast, low-energy retrieval operations crucial for real-time safety checks.
- Mixture of Experts (MoE) Intent Detector
- This is a lightweight detector that uses a gating network to route queries into specialized 'experts,' where each expert handles a specific neighborhood signature. This allows the system to efficiently process the diverse query space and make an informed safety decision.
- Cross-domain Harmful-Benign Pairs
- These are pairs of prompts generated using initial seed prompts and then recursively varied through techniques like paraphrasing and abstraction. This process creates dense, high-quality neighborhoods that ensure the system can accurately distinguish between safe and unsafe inputs.
Terminology
Summary
Quantized large language models (qLLMs) are increasingly deployed on edge devices for their low latency and energy efficiency, but model quantization weakens alignment safeguards, leaving qLLMs highly vulnerable to jailbreak attacks.
The gist
MOMAT introduces a hardware-enhanced safety framework that combines structured knowledge retrieval with low-power defense acceleration to mitigate the curse of dimensionality in large safety databases and provide domain-localized RetrievalAugmented Generation guarding for edge-deployed qLLMs.
How it works: Framework Overview
MOMAT is a CiM–assisted safety framework that decouples prompt encoding from the protected qLLM using a lightweight, frozen encoder-only embedding model (e.g., BAAI/bge-large-en-v1.5). The system comprises four modules: a CiM retrieval module, atlas-specific safety stores, a lightweight Mixture of Experts (MoE) intent detector, and a CiM-accelerated RAG engine that serves two roles: computing atlas-local top–k similarity features for MoE scoring and retrieving safe response templates directly within memory arrays.
How it works: Atlas Construction
Instead of using a single large safety database, MOMAT organizes harmful and benign samples into multiple semantically coherent subspaces called atlases. The construction process is structured in three steps:
-
Use a small set of seed prompts to ICL-generate cross-domain harmful–benign pairs.
-
Recursively fan out each pair with paraphrasing, syntactic variation, semantic abstraction, and detailed expansions to obtain dense variant sets.
-
Assemble these variants into atlas-specific corpora, yielding domain-labeled, high-density neighborhoods that reduce retrieval noise and align naturally with CiM-executed similarity search.
How it works: Retrieval and Intent Detection
After encoding a prompt into an embedding, the CiM engine retrieves top–k similarity features
from all atlases in parallel. This confines safety processing to appropriate domains, suppressing cross-domain retrieval noise. A compact MoE detector then receives the prompt embedding concatenated with atlas-local top–k precomputed features to produce a safety decision. The online inference phase involves two modes: first, computing atlas-local top–k similarity features for every query to support MoE scoring; second, when harmful intent is detected, performing a targeted second pass retrieval
from the most activated atlas to surface a safe response template.
How it works: Mixture-of-Experts for Pattern-Specific Routing
To handle the diversity of query neighborhoods across the multi-atlas space, MOMAT employs a Mixture-of-Experts (MoE) architecture where each expert specializes in a particular neighborhood signature. The system constructs a rich feature vector, denoted as 3, which captures not only the original query embedding but also raw similarity distributions
and statistical summaries per atlas,
including pairwise comparative features for atlas pairs. A gating network produces expert routing weights conditioned only on the query embedding, and the final risk estimate is computed as a weighted aggregation of expert scores.
How it works: Hardware Acceleration and Efficiency
The core efficiency gain comes from using Compute-in-Memory (CiM) arrays to perform similarity search directly in memory. This eliminates DRAM-based vector search, which requires repeated high-latency data transfers and incurs significant energy overhead. The CiM engine executes parallel Multiply-Accumulate (MAC) operations directly within the storage arrays, enabling sub–millisecond, low–energy retrieval
for both MoE scoring and targeted template retrieval. This results in a 4.69 × 106× speedup
and a reduction in energy consumption of approximately 2.5 × 105× over DRAM-based baselines for the same workload.
How it works: Evaluation Metrics
MOMAT is evaluated on standard jailbreak benchmarks, reporting four metrics: Attack Success Rate (ASR), False Refusal Rate (FRR), p95 end–to–end latency, and energy per query. The results demonstrate that MOMAT achieves 0.0% ASR across both datasets
for Llama2-7B and Mistral-7B under W4A8 quantization, while maintaining the same FRR as the baseline defense, confirming it avoids additional overrefusal of benign queries.
This shows that strong safety enforcement and benign utility preservation are not mutually exclusive.
Conclusion
MOMAT restores alignment behaviors sacrificed by quantization by restructuring retrieval into semantically coherent atlases and executing similarity search on a CiM-accelerated engine, proving that modular defenses can make edge-deployed qLLMs both safer and more energy-efficient. This approach suggests that reliable safety for edgedeployed qLLMs requires co-designing the retrieval structure and the underlying hardware substrate.
References
M. Ali et al.
Improvements for AI systems
Based on the provided research paper, MOMAT: Mixture of Multiple Atlases for Low-Power Jailbreak Defense of Quantized LLMs,
here are specific, actionable improvements that can be implemented in AI systems, along with what those improved systems will be able to do:
) 1. Implement Hardware-Accelerated Retrieval for Safety Guards:
Improve existing Retrieval-Augmented Generation (RAG) or safety filtering mechanisms by integrating a Compute-in-Memory (CiM) engine.
This system will perform similarity search for safety templates and refusal policies directly within the memory arrays, achieving a speedup of up to 4.69 × 106 and energy reduction of up to 2.5 × 105 compared to DRAM-based systems.
This enables the deployment of real-time, low-latency safety checks on edge devices (e.g., Raspberry Pi, mobile hardware) that currently cannot sustain the bandwidth demands of traditional vector databases.
) 2. Restructure Safety Knowledge into Domain-Localized Atlases:
Replace monolithic safety databases with multiple, semantically coherent atlases (Mixture of Multiple Atlases).
This system will organize harmful and benign samples into distinct semantic clusters, mitigating the
curse of dimensionalityand cross-domain interference that plague flat RAG systems.
This allows the system to retrieve highly relevant refusal templates for specific threat categories (e.g., self-harm vs. illegal acts) with high precision, even when the input prompt is semantically ambiguous or uses adversarial paraphrasing.
) 3. Integrate a Lightweight Mixture-of-Experts (MoE) Safety Scorer:
Implement a lightweight MoE architecture that fuses prompt embeddings with domain-localized retrieval features to produce a safety decision.
This system will utilize multiple small MLP experts, each specialized in detecting specific neighborhood patterns within the multi-atlas space, to calculate a final harm probability.
This provides fine-grained intent detection by allowing the system to distinguish between subtle jailbreak attempts that might be misclassified by simpler classifiers, ensuring high accuracy without incurring the computational overhead of large safety models.
) 4. Create an Offline Atlas Construction Pipeline:
Develop a robust offline process for automatically generating these domain-localized atlases using In-Context Learning (ICL).
This pipeline will recursively fan out seed prompts through paraphrasing and semantic abstraction to generate dense, high-density neighborhoods of harmful/benign pairs before assembling them into the final atlas structure.
This automates the creation of safety knowledge bases, ensuring that the system is continuously updated with novel jailbreak techniques without requiring constant retraining of large safety models.
) 5. Use a Frozen, Small Encoder for Prompt Encoding:
Employ a small, frozen encoder-only embedding model (like BAAI/bge-large-en-v1.5) solely to generate stable, full-precision embeddings for safety processing.
This component will be independent of the primary qLLM being defended and remain frozen during inference to ensure stability and low latency.
This decouples the safety mechanism from the core LLM's operational state, allowing it to function as a modular guardrail that can be swapped or updated independently of the main model weights.
) 6. Implement a Safe Response Template
Retrieval Trigger:
Design an inference pipeline where, upon detecting high harm probability, the system triggers a targeted retrieval pass from only the most activated atlas to surface a pre-vetted safe response template.
This will occur entirely in-situ within the CiM arrays when needed, minimizing data movement.
This ensures that when a jailbreak is detected, the system provides an immediate, policy-compliant refusal or redirection using curated templates rather than relying on complex, resource-intensive generative safety layers at inference time.
The resulting improved AI system will be a highly efficient and robust safety layer for quantized LLMs deployed on edge devices. It can perform:
-
Maintain high alignment robustness against jailbreaks, even when the underlying LLM is heavily quantized (e.g., 4-bit/8-bit).
-
Operate with extremely low latency (sub-millisecond) and minimal energy consumption (µJ range) on resource-constrained hardware like edge servers or embedded chips.
-
Provide fine-grained, category-specific safety enforcement by dynamically routing the query to the most relevant safety knowledge domain while ensuring benign user queries are never over-blocked (zero False Refusal Rate).
Sources
- Tone Matters: The Impact of Linguistic Tone on Hallucination in VLMs
- Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation
- LLM-as-Judge Framework for Evaluating Tone-Induced Hallucination in Vision-Language Models
- Investigating the Impact of Quantization Methods on the Safety and Reliability of Large Language Models
- Cosmos: A CXL-Based Full In-Memory System for Approximate Nearest Neighbor Search
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- ZeroQuant-FP: A Leap Forward in LLMs Post-Training W4A8 Quantization Using Floating-Point Formats
- C-Pack: Packed Resources For General Chinese Embeddings
- Beyond the Tip of Efficiency: Uncovering the Submerged Threats of Jailbreak Attacks in Small Language Models
- Universal and Transferable Adversarial Attacks on Aligned Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs