Auditing and Repairing LLM-as-Judge Failures in a Production Text-to-SQL Pipeline
cs.CL, cs.LG
Submitted: 2026-09-09
Updated: 2026-09-09
Code: https://github.com/JamesL404/synca-audit
Terminology
Sources
- FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance
- To Err Is Human; To Annotate, SILICON? Toward Robust Reproducibility in LLM Annotation
- Recent Advances in Text-to-SQL: A Survey of What We Have and What We Expect
- Can LLM be a Personalized Judge?
- Effective Proxy for Human Labeling: Ensemble Disagreement Scores in Large Language Models for Industrial NLP
- Text-to-SQL Empowered by Large Language Models: A Benchmark Evaluation
- A Survey on LLM-as-a-Judge
- Judge's Verdict: A Comprehensive Analysis of LLM Judge Capability Through Human Agreement
- Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement
- No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding
- Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows
- Auto-Prompt Ensemble for LLM Judge
- Can LLM Already Serve as A Database Interface? A BIg Bench for Large-Scale Database Grounded Text-to-SQLs
- Evaluating the Evaluator: Measuring LLMs' Adherence to Task Evaluation Instructions
- DIN-SQL: Decomposed In-Context Learning of Text-to-SQL with Self-Correction
- Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
- Semantic Agreement Enables Efficient Open-Ended LLM Cascades
- Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models
- Large Language Models are not Fair Evaluators
- Self-Instruct: Aligning Language Models with Self-Generated Instructions
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering