SaplingGuard: A Multidimensional-Profile-Aware Multi-Agent Guardrail for Developmentally Safe Adolescent-LLM Interaction
cs.CR
Submitted: 2026-09-27
Updated: 2026-09-27
Terminology
Sources
- The Age of Curiosity Meets the Age of AI: Benchmarking Child Safety in Large Language Models
- Constitutional AI: Harmlessness from AI Feedback
- The Llama 3 Herd of Models
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Mixtral of Experts
- Safe-Child-LLM: A Developmental Benchmark for Evaluating LLM Safety in Child-LLM Interactions
- Evaluating LLM Safety Across Child Development Stages: A Simulated Agent Approach
- GPT-4 Technical Report
- Qwen2.5 Technical Report
- LLM Safety for Children
- CASE-Bench: Context-Aware SafEty Benchmark for Large Language Models
- Gemini: A Family of Highly Capable Multimodal Models
- GuardAgent: Safeguard LLM Agents by a Guard Agent via Knowledge-Enabled Reasoning
- Qwen3 Technical Report
- Let Them Down Easy! Contextual Effects of LLM Guardrails on User Perceptions and Preferences
- Speak Out of Turn: Safety Vulnerability of Large Language Models in Multi-turn Dialogue
- Embracing AI in Education: Understanding the Surge in Large Language Model Use by Secondary Students
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs