TrustDABench: Benchmarking Reliability and Robustness of LLMs for Structured Data Analysis
cs.CL, cs.SE
Submitted: 2026-08-25
Updated: 2026-08-25
Comments: Code&Data: https://github.com/Skyorca/TrustDABench
Code: https://github.com/Skyorca/TrustDABench
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- The Mighty ToRR: A Benchmark for Table Reasoning and Robustness
- Exploring the Robustness of Language Models for Tabular Question Answering via Attention Analysis
- Dr.Spider: A Diagnostic Evaluation Benchmark towards Text-to-SQL Robustness
- PRACTIQ: A Practical Conversational Text-to-SQL dataset with Ambiguous and Unanswerable Queries
- Evaluating the Data Model Robustness of Text-to-SQL Systems Based on Real User Queries
- Towards Robustness of Text-to-SQL Models against Synonym Substitution
- InfiAgent-DABench: Evaluating Agents on Data Analysis Tasks
- DSBench: How Far Are Data Science Agents from Becoming Data Science Experts?
- Same Data, Different Schemas: Robustness of LLM-based Text-to-SQL
- Towards Fair In-Context Learning with Tabular Foundation Models
- AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions
- Unmasking Database Vulnerabilities: Zero-Knowledge Schema Inference Attacks in Text-to-SQL Systems
- TrustSQL: Benchmarking Text-to-SQL Reliability with Penalty-Based Scoring
- Tapilot-Crossing: Benchmarking and Evolving LLMs Towards Interactive Data Analysis Agents
- Are Your LLM-based Text-to-SQL Models Secure? Exploring SQL Injection via Backdoor Attacks
- SAFENLIDB: A Privacy-Preserving Safety Alignment Framework for LLM-based Natural Language Database Interfaces
- SpreadsheetBench: Towards Challenging Real World Spreadsheet Manipulation
- Confidence Estimation for Text-to-SQL in Large Language Models
- Interpretable LLM-based Table Question Answering
- Towards Robustness of Text-to-SQL Models Against Natural and Realistic Adversarial Table Perturbation
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering