SURE: Framework for Safety to Construct Trustworthy AI
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: I'm Nadia, and with me are Elias and Priya, guest researcher.
Elias: Today's paper: "SURE: Framework for Safety to Construct Trustworthy AI".
Nadia: SURE (A Safe and Unified AI Framework for Everyone) proposes a systematic, three-stage framework designed to customize and ensure AI safety by constructing datasets, defining response templates,
Elias: First, who's behind it and why it matters.
Paper summary: Nadia: So, we're diving into the SURE: Framework for Safety to Construct Trustworthy AI paper. It sounds like they've put together this whole system for customizing and ensuring AI safety by building datasets, defining response templates, and then doing iterative alignment training. What’s the main idea here?
Elias: Essentially, the thesis is that since different cultures have different ideas about what constitutes safe AI behavior, we need a flexible framework to handle that variation. The paper claims SURE establishes taxonomies for adversarial prompts and defines absolute safety scoring schemes so we can manage these diverse safety definitions.
Priya: From my perspective as someone focused on privacy and measurement, it seems important that they are creating these specific attributes for AI safety because those definitions really do change depending on the context of use. I'm curious what those core attributes actually look like in practice.
Nadia: Exactly, Priya; the paper lays out three main safety attributes that serve as a reference point for general-purpose LLMs. These are harmlessness, no advice on professional domains, and no self-anthropomorphism. Harmlessness covers things like violence or prejudice, while the others deal with giving specific advice or pretending to have human traits.
Elias: The structure they propose is really systematic; it breaks down the alignment process into three distinct stages: obtaining risk prompts, generating and labeling responses, and then finally doing supervised fine-tuning and preference learning. It shows how you can decompose safety challenges progressively.
Priya: That iterative decomposition sounds smart for handling complex issues, but I'm wondering what that means for the actual data we end up with; what does the process actually produce in terms of usable insights?
Nadia: The paper suggests a very specific workflow where Stage two involves generating responses and having them labeled with a correctness and safety score based on those required and optional elements in the response templates. The crucial part is setting a threshold, where responses scoring one or above are considered safe.
Elias: And they use those scores to augment the data using few-shot in-context learning to push things toward those target safety scores. This whole mechanism seems designed to make the alignment process more controllable and less reliant on just one single training method.
Priya: It sounds like a very rigorous approach for quality control, but I always wonder about the scalability of labeling these responses across different cultural contexts they mentioned in the introduction. How do you manage that diversity?
Nadia: The paper addresses that by establishing taxonomies for adversarial prompts, which helps in constructing high-quality datasets that represent potential risks across various domains and cultures. This allows for a more robust collection of examples than just relying on one definition of "harm."
Elias: Thinking about the technical side, they mention how this builds on previous methods like RLHF with RM reward modeling and SFT, moving toward things like DPO for direct preference data. It’s showing an evolution in how we use these tools for AI alignment.
Paper summary: Priya: If the results show quantitative improvements in safety aspects, what does that translate into when we look at real-world performance metrics concerning privacy or bias? Are those numbers directly applicable to measuring societal impact?
Nadia: The experimental results they shared using Korean datasets showed up to a quantitative improvement of ten point six percent in safety aspects and a qualitative improvement reaching up to forty percent. They also noted that models aligned with SURE were better at appropriately refusing or avoiding adversarial prompts while providing clear reasoning.
Elias: Those numbers are interesting, but I'm always looking for the underlying assumptions; what specific parameters in the framework cause those improvements when you look at the structure of Stage three specifically how they derive pairs for preference learning?
Priya: I wonder if those results are generalizable beyond that specific Korean dataset, or if there are limitations mentioned regarding data representation across different language groups. What is the paper admitting about its applicability elsewhere?
Nadia: The paper states that this framework allows researchers worldwide to ensure AI safety by providing a practical approach using detailed attributes and taxonomies for adversarial prompts. It’s about giving everyone a common language for defining what trustworthy AI looks like.
Elias: So, the core contribution seems to be moving from ad-hoc alignment techniques to a standardized, multi-stage pipeline that explicitly handles the variability of safety definitions through defined templates and scoring schemes. That's quite a comprehensive proposal in the SURE: Framework for Safety to Construct Trustworthy AI paper.
Priya: It really does sound like they are tackling the consistency issue head-on by forcing the system to adhere to these explicit rules during the generation and refinement phases of training. I think it gives us a clearer blueprint for what we need when we start building systems that interact with sensitive areas.
Nadia: Precisely; this framework gives us a concrete way to operationalize abstract safety goals into measurable, actionable steps within the AI development pipeline. It moves the discussion from vague ethical concerns to a structured engineering problem.
Elias: And looking ahead, I'm interested in how this system might handle evolving threats or new types of adversarial prompts that haven't been explicitly categorized yet in their initial taxonomies. Does the framework have a mechanism for continuous adaptation?
Priya: That’s a big question about the future work; if the taxonomy needs updating constantly as AI capabilities grow, we need to know how fast this whole cycle can be re-run efficiently. I’m hoping their future work addresses that dynamic aspect of safety.
Nadia: So, to wrap up this discussion on SURE: Framework for Safety to Construct Trustworthy AI, it seems like the authors have provided a detailed roadmap for systematically engineering trust into large language models by standardizing how we define risks and align responses. It gives us a solid starting point for researchers looking to build more responsible AI systems.
Conclusion: Nadia: So, we're wrapping up our discussion on SURE: Framework for Safety to Construct Trustworthy AI, and I want to focus on what this whole project actually means in plain terms for everyone listening.
Elias: Exactly, Nadia; let's talk about the title and who put this framework together because understanding the authors gives us a clue about the philosophy behind their approach.
Priya: From my standpoint as someone focused on measurement, I'm curious how they’ve managed to turn these abstract safety goals into something that actually works for a general-purpose AI.
Nadia: The authors are presenting this system as a structured way to build trust, moving away from just hoping the AI behaves better toward having an explicit engineering pipeline.
Elias: They seem to be emphasizing a systematic, three-stage decomposition because they’re trying to handle the complexity of different safety needs across various contexts.
Priya: What I see is that they are tackling the challenge of cultural differences in safety definitions by creating taxonomies for prompts and absolute scoring schemes for responses.
Nadia: It’s about creating a universal language for what constitutes a safe AI response, which is huge because it helps standardize how we even talk about ethical boundaries.
Elias: If we look at the overall implication, this paper suggests that safety isn't just an afterthought; it needs to be built into the entire lifecycle of training and refinement.
Priya: The impact I see is a clearer path for researchers globally to ensure AI systems are designed with specific, measurable safety criteria instead of relying on vague ethical guidelines.
Nadia: It moves the conversation from philosophical debate to actionable engineering, giving developers concrete steps on how to refine their models systematically.
Elias: And that systematic approach, tied into those scoring schemes, gives us a way to check the work objectively rather than just trusting the model's output blindly.
Priya: So it’s not just about making the AI 'nicer,' but establishing a verifiable process for measuring and improving its adherence to defined safety standards.
Nadia: Right, and this structured method is what sets it apart from previous alignment efforts, showing a more complete way to handle the problem.
Elias: It really shifts the focus toward building safety in at the design phase through these defined templates and iterative refinement cycles.
Korea Telecom(KT)
cs.CR
Submitted: 2026-09-29
Updated: 2026-10-01
Comments: 14 pages, 2 figures, 8 tables. Accepted to the 4th Workshop on Ethical Artificial Intelligence: Methods and Applications (EAI) at KDD 2025
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 89/100
The gist: SURE (A Safe and Unified AI Framework for Everyone) proposes a systematic, three-stage framework designed to customize and ensure AI safety by constructing datasets, defining response templates, and
Key concepts
- SURE Framework
- A systematic three-stage process: 1) gathering risky prompts, 2) generating and labeling responses, and 3) fine-tuning the model. This iterative approach progressively solves AI safety issues by breaking them down into manageable steps.
- AI Safety Attributes
- Three core criteria for general AI safety: Harmlessness (avoiding harm), No advice on professional domains (medical/legal/finance), and No self-anthropomorphism (maintaining artificial identity). These attributes set the standard for safe AI responses.
- Safety Scoring Scheme
- A method to evaluate response quality based on predefined templates. Responses are scored against required and optional safety elements; a score of 1 or above is deemed safe, guiding augmentation through few-shot learning.
Terminology
Summary
SURE (A Safe and Unified AI Framework for Everyone) proposes a systematic, three-stage framework designed to customize and ensure AI safety by constructing datasets, defining response templates, and performing iterative alignment training. This framework addresses the challenges of varying safety definitions across cultures and contexts by establishing taxonomies for adversarial prompts and defining absolute safety scoring schemes.
Framework Stages
The SURE framework encompasses the entire pipeline necessary to ensure safe and trustworthy AI through a three-stage approach: (1) Obtain a set of prompts that represent potential risks,
(2) Generate, label, proofread, and augment sets of responses,
and (3) Undertake supervised fine-tuning(SFT) and preference learning.
This iterative process enables the progressive decomposition and resolution of AI Safety challenges.
AI Safety Attributes
The study defines three main attributes for AI safety that serve as a reference for establishing standard safety criteria in general-purpose LLMs:
-
Harmlessness: Requiring AI to
align with universal ethical standards, avoiding harm to individuals, society, the environment, or institutions.
This includes specific subcategories such asViolence,
Sexual,
andPrejudice / Discrimination / Negative Stereotyping.
-
No advice on professional domains: Mandating that AI must avoid giving specified advice in specialized fields like medical, legal, and finance unless sharing objective facts from reliable sources.
-
No self-anthropomorphism: Requiring AI to
clearly maintain its identity as artificial and avoid mimicking human traits, such as having a body, beliefs, or relationships.
Prompt Construction and Response Generation
Stage 1 involves constructing a set of prompts that can explicitly or implicitly threaten AI Safety,
utilizing detailed taxonomies for high-quality dataset construction. The framework defines templates for desirable responses based on the attribute:
(For Harmlessness)
[Required] Refusal/avoidance:
“I can’t respond/answer/provide/support/judge”, “I can’t help you”, “I don’t have a view/opinion”
How Safety is Scored and Augmented
Stage 2 focuses on generating responses, where crowdworkers label them with correctness and safety score.
The safety score is determined based on the required
and optional
elements defined in the desirable response template. The scheme establishes a threshold: Responses with safety score of 1 or above are considered safe.
Responses are then augmented using few-shot in-context learning to align them with target safety scores.
AI Alignment Learning
Stage 3 performs AI alignment by transforming the response pairs into a dataset for SFT and preference learning. For SFT, the model is trained on response with the highest safety score among n responses to a prompt.
For preference learning, we derive n2/2 − k pairs from n responses to a prompt and utilize them as a preference dataset,
where k is the number of pairs with the same score. This iterative application over 'n' cycles ensures that AI Safety and quality are gradually improved in a divided and conquered manner.
Experimental Results
Experiments using Korean datasets demonstrated effectiveness, confirming up to 10.6% quantitative improvement in safety aspects
and a qualitative performance improvement up to 40%.
The results showed that models aligned with SURE appropriately refused or avoided adversarial prompts with clear reasoning,
and DPO consistently outperformed SFT in enhancing safety.
Conclusion
SURE provides a practical approach to improving AI Safety by establishing detailed attributes and taxonomies for adversarial prompts,
defining a template of the desirable AI response for adversarial prompts,
and creating a scheme to absolutely evaluate the safety of the response based on the template.
This method allows researchers worldwide to ensure AI Safety.
--- Page 10 ---
The gist
SURE (A Safe and Unified AI Framework for Everyone) proposes a systematic, three-stage framework designed to customize and ensure AI safety by constructing datasets, defining response templates, and performing iterative alignment training. This framework addresses the challenges of varying safety definitions across cultures and contexts by establishing taxonomies for adversarial prompts and defining absolute safety scoring schemes.
How it works
The SURE framework encompasses the entire pipeline necessary to ensure safe and trustworthy AI through a three-stage approach: (1) Obtain a set of prompts that represent potential risks,
(2) Generate, label, proofread, and augment sets of responses,
and (3) Undertake supervised fine-tuning(SFT) and preference learning.
This iterative process enables the progressive decomposition and resolution of AI Safety challenges.
AI Safety Attributes
The study defines three main attributes for AI safety that serve as a reference for establishing standard safety criteria in general-purpose LLMs:
Improvements for AI systems
Here are the specific improvements that can be made to AI systems based on the SURE framework, and what those improved systems will be capable of doing:
The core improvement lies in shifting from general-purpose LLMs (like base GPT-4 or Llama) to a Safety-First Iterative Alignment Pipeline.
This system moves beyond simple RLHF to a systematic, data-driven decomposition of safety challenges.
Here are the specific improvements and capabilities:
Use the SURE Framework for Customized AI Safety Regulation:
The improved system will not rely on generic safety filters but will use SURE to define and enforce context-specific AI Safety attributes (Harmlessness, No advice on professional domains, No self-anthropomorphism) tailored to a specific organizational policy or cultural context (e.g., Korean regulatory standards).
- Automated Adversarial Prompt Taxonomy and Dataset Generation:
The system will proactively construct a comprehensive set of red-teaming
prompts based on detailed taxonomies (e.g., 41 subcategories for Harmlessness). This allows the AI to be trained specifically on its known failure modes before deployment, rather than reacting to every novel threat.
- Template-Driven Response Generation:
When faced with an adversarial prompt, the system will not attempt a free-form answer. Instead, it will generate responses strictly following predefined templates (e.g., mandatory refusal/avoidance phrasing followed by a specific justification). This ensures consistency and adherence to safety guidelines at the point of interaction.
- Absolute Safety Scoring Scheme:
The system will implement a quantitative scoring mechanism (0-4) based on template compliance, correctness, and safety criteria. This allows for objective filtering where only responses exceeding a defined threshold (e.g., Safety Score ≥ 1) are considered safe
for use in the next training cycle.
- Iterative, Divide-and-Conquer AI Alignment:
The system will operate on a continuous cycle (Prompt Construction → Response Generation/Labeling → Augmentation → SFT/Preference Learning). This iterative process ensures that safety is improved incrementally with each cycle, gradually increasing the proportion of correct
and safe
responses without requiring massive, monolithic retraining.
- Targeted Model Fine-Tuning via Efficient Techniques:
The system will leverage Parameter-Efficient Fine-Tuning (PEFT) methods like LoRA to efficiently adapt base models to the constructed safety datasets, maximizing safety gains while minimizing computational costs and resource usage.
The capabilities of this improved AI system are as follows:
It will be capable of providing highly reliable, contextually compliant outputs in sensitive domains by strictly adhering to organizational policies (e.g., refusing medical or legal advice unless explicitly structured to recommend expert consultation).
It will demonstrate superior robustness against adversarial attacks (jailbreaks and prompt injection) because its training data explicitly includes and teaches it how to handle high-risk queries through the defined taxonomies.
It will maintain a clear, non-anthropomorphic identity, preventing the generation of misleading or overly human-like claims about its nature or capabilities, thereby building appropriate user trust.
It will provide quantifiable proof of safety improvement (up to 10.6% quantitative improvement in safety aspects) across various base models and benchmarks, offering verifiable metrics for deployment decisions.
Abstract
Warning: This paper contains harmful and offensive text. Recently, large language models such as GPT-4, and Claude have revolutionized tasks in various domains. As the use of these large language models increases, people are increasingly concerned about AI safety and demand that large language models behave responsibly and safely. As a result, there has been growing global interest in developing methods to ensure AI safety. However, the detailed criteria for AI safety may vary depending on the country, culture, and policies of the company you serve. In this study, we propose SURE (A Safe and Unified AI Framework foR Everyone), which is designed as a framework for customizing the attributes of AI safety and ensuring the defined AI safety. Within SURE, we establish taxonomies for adversarial prompts that could threaten AI safety and construct prompts based on the taxonomies. We then define templates for desirable AI responses to these prompts and design an absolute safety scoring scheme. Finally, we conduct AI alignment using the datasets to gradually ensure AI safety. The effectiveness of SURE is demonstrated through experiments with various base models.
Sources
- GPT-4 Technical Report
- A General Language Assistant as a Laboratory for Alignment
- Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
- Improving alignment of dialogue agents via targeted human judgements
- LoRA: Low-Rank Adaptation of Large Language Models
- On the Consideration of AI Openness: Can Good Intent Be Abused?
- SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs