FinRED: An Expert-Guided Benchmark Generation and Evaluation Framework for Financial LLM Red-Teaming
cs.CR, cs.AI
Submitted: 2026-06-18
Updated: 2026-09-22
Comments: Accepted at IEEE ICDM 2026
Code: https://github.com/selectstar-ai/FinRED-paper
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Existing safety benchmarks target general adversarial scenarios but miss finance-specific risks.
Terminology
Abstract
Existing safety benchmarks target general adversarial scenarios but miss finance-specific risks. Financial LLMs face regulatory compliance violations, fraud facilitation, and systemic trust erosion that require targeted evaluation. We introduce FinRED, an expert-guided red-teaming framework for financial LLM safety evaluation developed with financial experts. FinRED uses a novel two-level taxonomy mapping global standards (e.g., FATF and EU DORA) to threats ranging from regulatory evasion to complex fraud, integrated with a scalable pipeline that converts real financial documents into context-rich red-teaming Behavioral Prompts (seeds) through an expert-defined schema. Rigorous expert validation confirms seed plausibility and realism for meaningful LLM safety evaluation. We also provide an expert-validated, finance-specific rubric that goes beyond disclaimer checks, aligns more closely with human experts than static one-size-fits-all rubrics, and reduces critical false negatives from 28 to 12. Aligned with internationally adopted risk-management and information-security standards (e.g., ISO/IEC 27001), FinRED is deployed in South Korea's Financial Security Institute (FSI) regulatory sandbox for generative AI security evaluation in real financial services. To mitigate dual-use risks, the dataset, generation pipeline, prompt template, and evaluation framework are gated for qualified researchers at https://github.com/selectstar-ai/FinRED-paper and https://huggingface.co/datasets/datumo/FinRED.
Sources
- Red teaming ChatGPT via Jailbreaking: Bias, Robustness, Reliability and Toxicity
- AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons
- SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models
- R-Judge: Benchmarking Safety Risk Awareness for LLM Agents
- Is ChatGPT a Financial Expert? Evaluating Language Models on Financial Natural Language Processing
- CFBenchmark: Chinese Financial Assistant Benchmark for Large Language Model
- FinanceBench: A New Benchmark for Financial Question Answering
- Technical Report: Full-Stack Fine-Tuning for the Q Programming Language
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- ALERT: A Comprehensive Benchmark for Assessing Large Language Models' Safety through Red Teaming
- WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
- Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs
- AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies
- SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code
- LongSafety: Evaluating Long-Context Safety of Large Language Models
- RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent
- CASE-Bench: Context-Aware SafEty Benchmark for Large Language Models
- FinQA: A Dataset of Numerical Reasoning over Financial Data
- Improved Techniques for Optimization-Based Jailbreaking on Large Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs