Lexara-RF: Reference-Free Metrics for Evaluating Conversational Visual Analytics Agents
cs.HC, cs.AI
Submitted: 2026-09-15
Updated: 2026-09-15
Comments: 4 pages, 1 figure Conversational Visual Analytics, Evaluation Metrics, Visualization Design, Cooperative Communication Principles
Journal ref: 2026 IEEE Visualization and Visual Analytics (VIS) Conference
Code: https://github.com/rapidfuzz/RapidFuzz
License: http://creativecommons.org/licenses/by/4.0/
The gist: Conversational visual analytics (CVA) agents powered by large language models generate visualizations and natural-language explanations from open-ended queries.
Terminology
Abstract
Conversational visual analytics (CVA) agents powered by large language models generate visualizations and natural-language explanations from open-ended queries. Evaluating these multimodal outputs is challenging: curated reference benchmarks are costly to author, cannot comprehensively capture the space of valid responses, and are unavailable in production. Building on the Lexara evaluation framework, we introduce Lexara-RF, a reference-free set of metrics that scores CVA outputs using only the prompt, data, and model response. We reformulate evaluation as verification: 13 metrics operationalize visualization design theory and Gricean cooperative principles as computable consistency, intent-alignment, and design validity checks. On a human-rated corpus of CVA test-cases, Lexara-RF achieves alignment comparable to reference-based formulations, outperforms surface-similarity NLG baselines, and localizes structurally grounded failures with high accuracy.
Sources
- SimVecVis: A Dataset for Enhancing MLLMs in Visualization Understanding
- nvBench 2.0: Resolving Ambiguity in Text-to-Visualization through Stepwise Reasoning
- VIS-Shepherd: Constructing Critic for LLM-based Data Visualization Generation
- Vi(E)va LLM! A Conceptual Stack for Evaluating and Interpreting Generative AI-based Visualizations
Related papers
- EduGage: A Multimodal Dataset and Benchmark for Sensor-Based Momentary Assessment of Engagement in Self-Guided Video Learning
- EvoDesign: Agentic Editable Diagram Creation via Design Expertise Evolution
- HAGI++: Head-Assisted Gaze Imputation and Generation
- Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving
- Review of Explainable Decision Support and Adaptive Human-Machine Interfaces for Automation Transparency in Maritime Autonomous Surface Ships
- Towards Cognitive Process-Aware Proactive Writing Support