Evaluating the Effects of Prompt Perturbation on Bias and Hallucination in Large Language Models
cs.CL, cs.AI, cs.HC
Submitted: 2026-09-18
Updated: 2026-09-18
Code: https://github.com/fiddler-labs/fiddler-auditor
Terminology
Sources
- GPT-4 Technical Report
- PaLM 2 Technical Report
- Quantifying and Reducing Stereotypes in Word Embeddings
- Uncovering Biases with Reflective Large Language Models
- Measuring Massive Multitask Language Understanding
- Language Models are Few-Shot Learners
- On Faithfulness and Factuality in Abstractive Summarization
- Large Language Models: A Survey
- Beyond Accuracy: Behavioral Testing of NLP models with CheckList
- Aequitas: A Bias and Fairness Audit Toolkit
- Societal Biases in Language Generation: Progress and Challenges
- Moral Mimicry: Large Language Models Produce Moral Rationalizations Tailored to Political Identity
- Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
- Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
- Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models
- Large Language Models are not Fair Evaluators
- DataFrame QA: A Universal LLM Framework on DataFrame Question Answering Without Data Exposure
- CHBias: Bias Evaluation and Mitigation of Chinese Conversational Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering