Quizzing the Translation: A Prover-Grounded Evaluation Metric for NL to FOL
cs.CL, cs.AI, cs.LO
Submitted: 2026-09-27
Updated: 2026-09-27
Code: https://github.com/pu-suo/siv-metric
Terminology
Sources
- Do LLMs Really Struggle at NL-FOL Translation? Revealing their Strengths via a Novel Benchmarking Strategy
- A Theorem-Proving-Based Evaluation of Neural Semantic Parsing
- ASSESS: A Semantic and Structural Evaluation Framework for Statement Similarity
- Generalized Tree Edit Distance (GTED): A Faithful Evaluation Metric for Statement Autoformalization
- Qwen2.5 Technical Report
- Divide and Translate: Compositional First-Order Logic Translation and Verification for Complex Logical Reasoning
- Assessing the Sensitivity and Alignment of FOL Closeness Metrics
- Improving Symbolic Translation of Language Models for Logical Reasoning
- Tool-Assisted Agent on SQL Inspection and Refinement in Real-World Scenarios
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering