A Unified Evaluation Framework for Trustworthy Large Language Models, Agentic AI, and Multimodal Systems
cs.AI
Submitted: 2026-09-17
Updated: 2026-09-18
License: http://creativecommons.org/licenses/by/4.0/
The gist: Benchmark scores alone provide an incomplete basis for assessing the trustworthiness of modern artificial intelligence systems.
Terminology
Abstract
Benchmark scores alone provide an incomplete basis for assessing the trustworthiness of modern artificial intelligence systems. Large language models (LLMs), agentic systems, and multimodal models (MLLMs) require different forms of assessment, yet their evaluation evidence must remain interpretable for development and oversight. We propose a unified framework that connects output-level, trajectory-level, and cross-modal assessment through eight trustworthiness dimensions: capability, robustness, safety, fairness, transparency, governance, oversight, and efficiency. The framework preserves system-specific metrics while mapping native measurements to common performance bands, accompanied by uncertainty estimates and traceable evidence. A meta-evaluation layer examines the validity, reliability, and reproducibility of the evaluation itself. Multidimensional profiles expose strengths and weaknesses, while safety-critical overrides prevent aggregate scores from masking critical failures. Mappings to governance frameworks, international standards, and European Union regulatory requirements connect technical assessment with oversight needs. The framework provides a structured basis for assessing both system performance and the credibility of the evidence supporting it, with empirical validation across deployment contexts remaining an essential next step.
Sources
- Training Verifiers to Solve Math Word Problems
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Program Synthesis with Large Language Models
- MAS-FIRE: Fault Injection and Reliability Evaluation for LLM-Based Multi-Agent Systems
- When Users Change Their Mind: Evaluating Interruptible Agents in Long-Horizon Web Navigation
- Language Models (Mostly) Know What They Know
- Uncertainty-Based Abstention in LLMs Improves Safety and Reduces Hallucinations
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection