DualityCert: Verifier-Gated Language-Model Repair of Broken Duality Claims in Quantum Field Theory
cs.CR, cs.AI, cs.LG, hep-th
Submitted: 2026-07-26
Updated: 2026-09-28
Comments: 17 pages, 2 figures, 9 tables. v2: added reference and note on concurrent related work. Code, benchmark, and all per-attempt records: https://github.com/xingyang-yu/QFTCert
Code: https://github.com/xingyang-yu/QFTCert
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Grading the Unspoken: Evaluating Tacit Reasoning in Quantum Field Theory and String Theory with LLMs
- Theoretical Physics Benchmark (TPBench) -- a Dataset and Study of AI Reasoning Capabilities in Theoretical Physics
- Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark
- LeanDojo: Theorem Proving with Retrieval-Augmented Language Models
- Resolution of Erd\H{o}s Problem #728: a writeup of Aristotle's Lean proof
- Learning to Disprove: Formal Counterexample Generation with Large Language Models
- Electric-Magnetic Duality in Supersymmetric Non-Abelian Gauge Theories
- Lectures on supersymmetric gauge theories and electric-magnetic duality
- HepLean: Digitalising high energy physics
- Large Language Models Cannot Self-Correct Reasoning Yet
- Training Verifiers to Solve Math Word Problems
- Teaching Large Language Models to Self-Debug
- Self-Refine: Iterative Refinement with Self-Feedback
- Reflexion: Language Agents with Verbal Reinforcement Learning
- Let's Verify Step by Step
- Quiver Mutations, Seiberg Duality and Machine Learning
- Learning to Trace Seiberg Dualities
- Formalization of QFT
- Axioms for physical reasoning: codifying the Seiberg--Witten solution in Lean
- PhysProver: Advancing Automatic Theorem Proving for Physics
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs