LLM-Based FORM Code Generation with Verification-Driven Fine-Tuning
hep-ph, cs.CE, cs.CL
Submitted: 2026-09-20
Updated: 2026-09-20
Code: https://github.com/huggingface/trl
License: http://creativecommons.org/licenses/by/4.0/
The gist: FORM is a domain-specific symbolic manipulation language widely used in particle physics for processing the very large algebraic expressions arising from multi-loop Feynman diagram calculations.
Terminology
Abstract
FORM is a domain-specific symbolic manipulation language widely used in particle physics for processing the very large algebraic expressions arising from multi-loop Feynman diagram calculations. Despite its central role in precision theoretical physics, no artificial-intelligence tooling exists, to our knowledge, for assisting physicists in writing FORM code. We show that contemporary large language models (LLMs), including frontier models with hundreds of billions of parameters, achieve a zero-percent execution pass rate on our instruction-following and tutorial-style FORM tasks without documentation in a single attempt, establishing FORM as a genuine zero-shot language for LLMs at the time of writing. We then present a verification-driven data generation pipeline that uses the FORM binary itself as an execution oracle to produce and validate a corpus of 4,633 training examples spanning deterministic computations, open-ended programs, tutorial code, and knowledge question-answer pairs. Fine-tuning a compact open-weights model (Qwen3-8B) with quantized low-rank adaptation (QLoRA) yields a specialist that, evaluated on four complementary benchmarks (840 tasks, single attempt each), decisively outperforms frontier models with up to 756B parameters in execution rate and in strict, FORM-verified output matching on the larger benchmarks, and remains statistically indistinguishable from them on the smaller, harder ones. General reasoning and coding capabilities are preserved within 2.6 percentage points.
Sources
- New features of FORM
- FORM version 4.0
- FORM Version 5.0
- The FORM project
- Forcer, a FORM program for the parametric reduction of four-loop massless propagator diagrams
- The Multiple Zeta Value Data Mine
- Evaluating Large Language Models Trained on Code
- QLoRA: Efficient Finetuning of Quantized LLMs
- LoRA: Low-Rank Adaptation of Large Language Models
- Qwen3 Technical Report
- GLM-5: from Vibe Coding to Agentic Engineering
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- Measuring Massive Multitask Language Understanding
- Training Verifiers to Solve Math Word Problems
- Knowledge Transfer from High-Resource to Low-Resource Programming Languages for Code LLMs
- Synthetic Programming Elicitation for Text-to-Code in Very Low-Resource Programming and Formal Languages
- Supervising Ralph Wiggum: Exploring a Metacognitive Co-Regulation Agentic AI Loop for Engineering Design
- Gemma 4 Technical Report
- gpt-oss-120b & gpt-oss-20b Model Card
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Related papers
- Classification of g-modes for neutron stars with a strong transition: Novel universal relation including slow stable hybrid stars
- Higgsino Dark Matter Interpretation of the LUX-ZEPLIN 248 keV Nuclear-Recoil Event
- A Unified Bogoliubov Approach to Primordial Gravitational Waves: From Inflation to Reheating
- Probing Memory-Burdened Primordial Black Holes with High-Energy Neutrinos
- Enhanced Dark Matter Quantum Sensing via Phase-Space Geometric Interferometry
- Axions as Dark Matter, Dark Energy, and Dark Radiation