Safety Nudges: User-Facing Interventions for Real-Time AI Risk Awareness
cs.HC, cs.AI, cs.CY, cs.LG
Submitted: 2026-09-22
Updated: 2026-09-24
Code: https://github.com/jtbwedgwood/safety-nudges
Project page: https://open-reflection.com
Terminology
Sources
- Constitutional AI: Harmlessness from AI Feedback
- How AI and Human Behaviors Shape Psychosocial Effects of Extended Chatbot Use: A Longitudinal Randomized Controlled Study
- WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
- Measuring and mitigating overreliance to build human-compatible AI
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Uncovering Deceptive Tendencies in Language Models: A Simulated Company AI Assistant
- Scalable Extraction of Training Data from (Production) Language Models
- Investigating Affective Use and Emotional Well-being on ChatGPT
- ShareChat: A Dataset of Chatbot Conversations in the Wild
- Individual-level interventions against sycophantic AI reduce its appeal but not its persuasiveness
- YouthSafe: A Youth-Centric Safety Benchmark and Safeguard Model for Large Language Models
- ShieldGemma: Generative AI Content Moderation Based on Gemma
- WildChat: 1M ChatGPT Interaction Logs in the Wild
Related papers
- EduGage: A Multimodal Dataset and Benchmark for Sensor-Based Momentary Assessment of Engagement in Self-Guided Video Learning
- EvoDesign: Agentic Editable Diagram Creation via Design Expertise Evolution
- HAGI++: Head-Assisted Gaze Imputation and Generation
- Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving
- Review of Explainable Decision Support and Adaptive Human-Machine Interfaces for Automation Transparency in Maritime Autonomous Surface Ships
- Towards Cognitive Process-Aware Proactive Writing Support