VEX-Bench: Benchmarking Verification Complexity of LLM-Generated Misinformation
cs.LG, cs.AI, cs.CL, cs.CY, cs.IR
Submitted: 2026-09-28
Updated: 2026-09-28
Terminology
Sources
- On the Trustworthiness of Generative Foundation Models: Guideline, Assessment, and Perspective
- ADMIT: Few-shot Knowledge Poisoning Attacks on RAG-based Fact Checking
- Safety in Embodied AI: A Survey of Risks, Attacks, and Defenses
- Internal Safety Collapse in Frontier Large Language Models
- Fighting an Infodemic: COVID-19 Fake News Dataset
- Evaluating Language Models for Harmful Manipulation
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
- MasterKey: Automated Jailbreak Across Multiple Large Language Model Chatbots
- Confusion is the Final Barrier: Rethinking Jailbreak Evaluation and Investigating the Real Misuse Threat of LLMs
- Generative Large Language Models in Automated Fact-Checking: A Survey
- Large Language Models are Inconsistent and Biased Evaluators
- JAILJUDGE: A Comprehensive Jailbreak Judge Benchmark with Multi-Agent Enhanced Explanation Evaluation Framework
- Counting on Consensus: Selecting the Right Inter-annotator Agreement Metric for NLP Annotation and Evaluation
- Kimi K2.5: Visual Agentic Intelligence
- DeepSeek-V3 Technical Report
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks