RISA: Response Inspection and Selective Actions for Refusal Calibration in Large Language Models
cs.CR
Submitted: 2026-09-01
Updated: 2026-09-01
Code: https://github.com/ChangWenhan/RISA
Terminology
Sources
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Steering Language Models With Activation Engineering
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Fine-Tuning Language Models from Human Preferences
- A General Language Assistant as a Laboratory for Alignment
- Constitutional AI: Harmlessness from AI Feedback
- Representation Engineering: A Top-Down Approach to AI Transparency
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs