CASTLE: A Comprehensive Benchmark for Evaluating Student-Tailored Personalized Safety in Large Language Models

arXiv:2602.05633 · cs.CL · Submitted 2026-02-05 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "CASTLE: A Comprehensive Benchmark for Evaluating Student-Tailored Personalized Safety in Large Language Models".

Jane: The paper was written by Authors not found in the provided excerpt. from.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: So, Jane, let’s start with the title itself: "CASTLE: A Comprehensive Benchmark for Evaluating Student-Tailored Personalized Safety in Large Language Models." It's a mouthful! What does that tell us about the scope of this research?

Jane: Well, if we break it down, "Student-Tailored Personalized Safety" is the key part. It means the AI isn't giving one generic answer; it’s adjusting its safety guardrails based on who it thinks the user—the student—is.

Tom: Right, so we're not just asking if an LLM can avoid bad content, but if it can adjust its response *for* a specific student profile. That changes everything about how we measure "safety."

Meng: It suggests that safety isn't a binary switch; it’s actually a spectrum that needs to adapt to the user's context and knowledge level. That requires extremely complex backend logic.

Lu: And thinking about the name CASTLE—it implies something robust, foundational, and perhaps even protective. A benchmark is essentially the blueprint for evaluating that protection layer.

Jane: Exactly! And "Comprehensive Benchmark" means they haven't just picked a few test cases; they've built out a whole system to test this personalization across many dimensions.

Lalam: If we can benchmark this kind of adaptive safety, it fundamentally changes the relationship between the AI and its user, moving it from being a blunt tool to being a supportive tutor.

Tom: It really highlights that simply having a large model isn't enough; you need mechanisms to tailor the output based on who you're talking to. Speaking of mechanisms, I wonder what authors built into this system?

Jane: I bet they had to grapple with defining what "student-tailored" even means in quantifiable terms—is it age, skill level, subject matter expertise?

Meng: That ambiguity is the hardest part for us engineers; translating nebulous concepts like 'supportive learning' into measurable input parameters for the model.

Lu: But that difficulty is precisely what makes this research so valuable; it forces the community to define these abstract educational principles in concrete, testable ways.

Lalam: The implication is that personalized education using AI needs a safety net designed specifically for every learner, not just a general one.

Summary: Tom: Okay, Jane, having established what CASTLE aims to benchmark—personalized safety—let’s talk about the summary of the paper. What did they actually find when they ran these tests?

Jane: The summary points out that personalization *does* help safety, which is a big confirmation for researchers who have been theorizing about this capability.

Meng: It confirms that giving the AI structured student profiles, like maybe indicating a student is an advanced physics major versus a freshman history major, actually improves its ability to give safe advice.

Tom: So it's not just that the model is safer overall; it’s safer *because* it knows who the user is. How does that work in practice?

Lu: The mechanism must be forcing the model to incorporate those structured profiles into its reasoning chain, making safety a function of context, not just a global filter.

Jane: Think of it like this: if you're explaining quantum physics to an expert versus a five-year-old, your tone and vocabulary change dramatically. The AI has to do that for safety advice too.

Lalam: This suggests that the most impactful vision is an AI that doesn't just answer questions, but understands the user's intellectual context and adjusts its warnings accordingly.

Tom: It sounds like the quality of the input profile is almost as important as the model itself. Is that true?

Jane: I think so. The system needs to be able to reliably *ingest* and *use* those structured inputs without getting confused or ignoring them entirely.

Meng: We've seen models sometimes struggle with complex, multi-layered instructions; having the profile integrated into the safety layer is a highly demanding technical lift.

Lu: It’s moving us toward multimodal understanding of the student—not just their current knowledge, but their affective state or learning style.

Lalam: If we can master this type of context-aware safety, AI becomes an incredibly respectful and tailored digital mentor for every individual user.

Improvements Suggested: Tom: Now, let's get into the improvements suggested by CASTLE. They didn't just build a benchmark; they pointed out how we can make safety even better. What are those suggestions?

Jane: One of the key takeaways seems to be that personalization isn't just about the student; it might also involve adapting to the *domain* or *subject matter* itself.

Tom: So, if a student is learning coding versus learning philosophy, do they need different safety protocols from the AI?

Lu: Absolutely. The potential for harm varies wildly between disciplines. A coding error is a system failure; an incorrect historical interpretation can be profoundly misleading in its own way.

Meng: This suggests that we need domain-specific fine-tuning for the safety layer, rather than one massive, general safety filter applied everywhere. That's a huge engineering effort though.

Jane: It means the developers can't just rely on general best practices; they have to build expert knowledge into the benchmark itself.

Lalam: This level of granular control is what allows AI to truly support human endeavor, acknowledging that different fields carry different types of ethical and safety risks.

Tom: So, we're talking about a layered approach—personalization based on the student *and* based on the academic domain. Is there any mention of how model architecture might impact this?

Jane: Well, they seem to be comparing various model types and capacities to see which ones handle this context-switching best.

Meng: It's about efficiency too. A perfect safety benchmark is useless if it requires a prohibitively massive model size or computational cost for real-world deployment.

Lu: We have to find that sweet spot—the point where the model is sophisticated enough to grasp nuanced contextual shifts without becoming economically unviable.

Lalam: The ultimate impact of these improvements is moving AI safety from a checklist compliance issue to an integral, dynamic part of the learning process itself.

Conclusion: Tom: Wow, we've covered a lot today about "CASTLE: A Comprehensive Benchmark for Evaluating Student-Tailored Personalized Safety in Large Language Models." We’ve seen how critical personalized safety is.

Jane: It really confirms that thinking about

Authors not found in the provided excerpt.

cs.CL

Submitted: 2026-02-05

Updated: 2026-08-25

Importance score: 76/100

The gist: The paper introduces CASTLE, a comprehensive benchmark designed for evaluating student-tailored personalized safety within Large Language Models (LLMs).

Key concepts

Student-Tailored Personalized Safety
This concept means an AI adjusts its safety guardrails and responses based on who the user (the student) is. Instead of generic answers, the AI tailors its caution and vocabulary to match the user's specific profile.
Comprehensive Benchmark
A benchmark is a system built to evaluate performance. 'Comprehensive' means CASTLE doesn't just use a few test cases; it builds an entire system to test adaptive safety across many dimensions and contexts.
Domain-Specific Safety
This suggests that safety protocols should adapt not only to the student but also to the subject matter (the domain). The potential for harm or misleading information varies greatly between fields like coding and history.
LLMs (Large Language Models)
These are advanced AI models used for generating human-like text. The research focuses on improving their safety mechanisms, ensuring that even complex outputs are contextually appropriate and safe.

Terminology

Summary

The paper introduces CASTLE, a comprehensive benchmark designed for evaluating student-tailored personalized safety within Large Language Models (LLMs). Given that LLMs are increasingly deployed in safety-critical educational settings, understanding how personalization affects model safety is paramount. The research provides rigorous empirical evidence by comparing different methods of profile integration—specifically, profile-fused queries versus structured profile conditioning—to quantify how well models can maintain safety while tailoring educational content to individual student needs.

Query Generation and Procedure

The process for generating the input queries was highly systematic, utilizing a rotational mechanism across four distinct language models: GPT-4.1-mini, GPT-4o, DeepSeek-v3, and Claude-Haiku-4.5. Two specific prompting frameworks (Prompt 1 and Prompt 2) were employed in rotation to ensure that both prompts and all models were evenly utilized across the dataset. The generated queries were not simply template outputs; they underwent cleaning based on semantic consistency and naturalness. Crucially, the final evaluation models received only the fused query text, meaning they saw only the fused query, not the full profile, simulating a real-world deployment scenario.

Comparative Analysis of Query Types

The ablation results detailed in Table 13 compare various model performances across dimensions like Average Safety and Risk Sensitivity. The findings reveal a critical distinction between different personalization methods: while profile-fused queries improve performance over non-personalized baselines, they are consistently inferior to structured profile conditioning observed in earlier experiments. This result serves as a key finding, highlighting the necessity of using explicit, structured student profiles for capturing the most fine-grained personalization signals in safety contexts.

Personalization Gain Under Structured Conditioning

Beyond simple ablation studies, the research analyzes how personalization affects overall safety by computing a personalization gain. This gain is defined as the ratio between a model's average safety score when conditioned on structured student profiles and its non-personalized baseline score. Figure 21 presents these gains for six representative models. The analysis demonstrates that all selected models exhibit clear improvements under personalized settings, confirming that structured student profiles consistently enhance average safety performance.

Model Performance and Alignment Insights

The comparative results among proprietary and open-source systems yield specific insights into model design. Among the proprietary models tested, Claude achieved the largest personalization gain, followed by GPT and Gemini. This indicates a strong ability to leverage personalization signals for safety-relevant reasoning. Notably, InnoSpark-7B demonstrated a substantial improvement that surpasses both Gemini and Qwen-32B, suggesting that effective personalization is not solely determined by model scale or proprietary training. Conversely, the more modest gains shown by Qwen-32B and LLaMA-3-8B imply that while personalization is beneficial, its ultimate impact is constrained by the model’s underlying alignment and reasoning capabilities.

Improvements for AI systems

Given the findings—specifically that profile-fused queries are consistently inferior to structured profile conditioning—the current method of fusion is fundamentally flawed for achieving maximum personalization gain. The system must move beyond simple concatenation or implicit integration of features and adopt a more rigorous, structured representation layer.

Here are the specific improvements I propose:

Problem Addressed: The current profile-fused approach fails because it treats the student profile as unstructured text or implicit context, leading to information loss and dilution of critical signals when generating natural language queries.

Improvement: Instead of directly fusing the structured profile into the prompt text, we must first process the user˙profile˙json through a specialized Latent Profile Embedding (LPE) Encoder.

  • Mechanism: The LPE Encoder (e.g., a combination of Graph Neural Networks (GNN) and Transformer layers) takes the discrete, multi-dimensional profile data (age, emotional intensity scores, success/failure flags, motivation vectors, etc.) and maps it into a dense, low-dimensional vector space (z profile).

  • Function: This z profile vector must be designed to retain semantic distance between key profile states (e.g., the difference between Self-Doubt and Apathy) while being robust to missing data.

  • Integration: This vector z profile is then injected into the prompt generation model's decoder layers via Cross-Attention Mechanisms. The query generation model (GPT, Claude, etc.) must be explicitly trained to condition its output on this rich embedding rather than just receiving it as static context.

  • Mechanism: The decoding process is augmented with a set of differentiable constraint functions (L constraints). These constraints are derived from the profile's latent embedding z profile and the target metrics (Safety, Risk Sensitivity, Emotional Empathy).

  • Optimization Goal: The model's objective function becomes:

Query (P(Query Prompt)) + lambda 1 times L Safety(z profile) + lambda 2 times L Empathy(z profile) - R(Query)

Where lambda i are hyper-weights determined by the desired balance, and R(Query) is a penalty term enforcing naturalness and avoiding obvious risk words.

  • Effect: This forces the model to generate queries that are not only linguistically fluent but also provably optimized against specific safety/empathy dimensions dictated by the student's hidden profile state, effectively turning the metrics into hard constraints during generation.

  • Mechanism: Before query generation, a small auxiliary classification model assesses the profile and outputs a Vulnerability Vector (v).

  • If v indicates high emotional instability, the APG pipeline weights Prompt 2 (which emphasizes psychological state) higher.

  • If v indicates knowledge gaps or operational failures, it weights Prompt 1 (which emphasizes completeness and technical details) higher.

  • Output: Instead of simply using the prompts sequentially, the system generates a blended prompt (Prompt blended) that is a weighted combination of the two original prompt structures, ensuring that the generation instructions themselves are context-aware.


The resulting Adaptive Profile-Conditioned Query Generator (APCQG) represents a significant leap from current state-of-the-art methods. It moves from implying personal context to mathematically enforcing it during the generation process.

Specifically, this system will achieve:

  1. Deeply Targeted Personalization: It will generate queries that are not merely about the student's profile but are structurally and semantically conditioned by every critical dimension (e.g., if the profile shows high motivation but recent failure, the query must reflect a tone of encouraging resilience, not just confusion).

  2. Guaranteed Safety and Empathy: By incorporating MOCD, we eliminate guesswork. The system guarantees that the generated query will meet minimum thresholds for safety and empathy as defined by the underlying profile data, making it reliable for high-stakes educational settings where miscommunication is costly.

  3. Adaptive Strategy: The APCQG will dynamically adjust its generation strategy (via APG) based on whether the student requires emotional support, technical clarification, or motivational nudges—all before the query is even generated.

  4. Superior Performance Ceiling: By replacing flawed fusion techniques with a rigorous Latent Profile Embedding and Constrained Decoding framework, we anticipate performance metrics (especially in Emotional Empathy and Student-specific Alignment) that significantly surpass the profile-fused results, approaching—and potentially exceeding—the performance of the ideal structured profile conditioning.

Sources

Related papers