FinRT: Distilling Adaptive Red-Teaming Strategies into Reusable Adversarial Generators in Consumer Finance
cs.CL, cs.CR, cs.LG
Submitted: 2026-09-29
Updated: 2026-09-29
Terminology
Sources
- Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
- Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
- Learning to Conceal Risk: Controllable Multi-turn Red Teaming for LLMs in the Financial Domain
- AutoRed: A Free-form Adversarial Prompt Generation Framework for Automated Red Teaming
- Risk-Adjusted Harm Scoring for Automated Red Teaming for LLMs in Financial Services
- The Vendi Score: A Diversity Evaluation Metric for Machine Learning
- FinSafetyBench: Evaluating LLM Safety in Real-World Financial Scenarios
- Jailbreak-Zero: A Path to Pareto Optimal Red Teaming for Large Language Models
- TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- SAGE-RT: Synthetic Alignment data Generation for Safety Evaluation and Red Teaming
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- gpt-oss-120b & gpt-oss-20b Model Card
- AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs
- CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models
- BloombergGPT: A Large Language Model for Finance
- WizardLM: Empowering large pre-trained language models to follow complex instructions
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- Adaptive Instruction Composition for Automated LLM Red-Teaming
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering