Constrained Decoding Eliminates Structural Failures in Small LLMs but Reveals a Scale-Dependent Semantic Gap
cs.CL
Submitted: 2026-09-20
Updated: 2026-09-20
Comments: 6 pages, ACL format Code and task suite: https://github.com/CruiseDevice/small-llm-structured-benchmark (tag v1.0)
Code: https://github.com/CruiseDevice/small-llm-structured-benchmark
License: http://creativecommons.org/licenses/by/4.0/
The gist: Small open-source large language models (LLMs) in the 0.6B-4B parameter range are increasingly deployed for structured output generation (JSON, function calling, data extraction), yet little is known
Terminology
Abstract
Small open-source large language models (LLMs) in the 0.6B-4B parameter range are increasingly deployed for structured output generation (JSON, function calling, data extraction), yet little is known about how constrained decoding (CD) interacts with model scale in this regime. We benchmark five models from three families across 14 structured-output tasks under three decoding conditions (native, Outlines, XGrammar). We introduce a two-axis evaluation that separates structural correctness (schema validity) from semantic correctness (content accuracy). We find that CD eliminates all structural failures across all models (schema validity: 78.6-92.9% to 100%), but content accuracy reveals a persistent semantic gap that is scale-dependent: type coercion failures are fully CD-rescuable, while instruction-semantic failures (e.g., multi-step function calling) remain CD-resistant. Schema conformance is necessary but not sufficient for semantic correctness; CD's reach ends exactly where schema conformance ends.
Sources
- XGrammar: Flexible and Efficient Structured Generation Engine for Large Language Models
- When Correct Isn't Usable: Improving Structured Output Reliability in Small Language Models
- JSONSchemaBench: A Rigorous Benchmark of Structured Outputs for Language Models
- Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
- Qwen3 Technical Report
- The Hidden Cost of Structured Generation in LLMs: Draft-Conditioned Constrained Decoding
- STED and Consistency Scoring: A Framework for Evaluating LLM Structured Output Reliability
- Efficient Guided Generation for Large Language Models
- HuggingFace's Transformers: State-of-the-art Natural Language Processing
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering