Learning an Interior Layout Policy in a Domain Specific Language Action Space
Yuhao Lu, Weichen Zhang, Wenyi Xiao, Haohui Chen, Yiyun Fei
Taobao & Tmall group of Alibaba
cs.CV, cs.AI
Submitted: 2026-07-31
Updated: 2026-08-11
License: http://creativecommons.org/licenses/by/4.0/
The gist: The paper proposes LayoutDSL, a novel LLM-based framework for indoor scene layout generation that learns an interior layout policy in a domain-specific language (DSL) action space.
Terminology
Summary
The paper proposes LayoutDSL, a novel LLM-based framework for indoor scene layout generation that learns an interior layout policy in a domain-specific language (DSL) action space. The authors argue that existing methods often oversimplify the task by reducing room conditions to coarse 3D bounding boxes and neglecting structural elements such as doors and windows, and more fundamentally, they cast layout generation as continuous parameter prediction, which hinders the model from learning the underlying reasoning logic of intelligent layout design. To address these problems, the paper introduces a layout DSL that symbolically encodes spatial relationships and geometric constraints, along with a generator–interpreter system for bidirectional conversion between layout parameters and DSL statements. The authors construct 3D-FrontDSL, a new layout dataset that incorporates key architectural elements—such as walls, doors, windows, and holes—together with corresponding layout DSL statements. Based on these paired annotations, they perform supervised fine-tuning (SFT) to align the LLM with the syntax and semantics of the layout DSL and to learn structured DSL action sequences for layout reasoning. To further enhance layout quality and foster a more robust exploration mechanism, they incorporate reinforcement learning (RL) with verifiable rewards grounded in geometric feasibility and design constraints. Specifically, they construct a comprehensive reward signal by combining established geometric metrics (Collision, Out-of-bounds, Reachability) with design-aware criteria (Forbidden-placement, Space Logicality), and update the policy via Group Relative Policy Optimization (GRPO). The paper reports that after DSL-based fine-tuning and RL, a 4B-parameter LLM achieves significant layout performance gains over the baseline and surpasses much larger industry-leading LLMs evaluated in an unfine-tuned, few-shot setting. The contributions are three-fold: designing a layout domain-specific language and developing a DSL generator–interpreter system, turning indoor scene layout generation into a policy-learning problem over a structured and interpretable action space; constructing 3D-FrontDSL, a new layout dataset with room-structure annotations and Layout DSL statements for room-structure-conditioned layout generation; and proposing comprehensive and verifiable layout rewards that incorporate design principles, showing that reinforcement learning with these rewards substantially improves layout generation performance.
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems:
Improvement: Replace direct coordinate/parameter prediction with a structured, symbolic action space for spatial tasks.
Implementation:
-
Design a DSL with 26 predefined alignment types (e.g.,
wall center left,obj front center) that encode spatial relationships relative to anchors (walls, doors, windows, or previously placed objects) -
Build a bidirectional generator–interpreter system that converts between raw coordinates and DSL statements
-
Train the model to output sequences of DSL actions rather than continuous geometric parameters
What the improved system can do:
-
Generate layouts with 28.4% higher mean score (0.878 vs 0.594) compared to state-of-the-art methods
-
Maintain stable positional descriptions across rooms of varying geometry (direct coordinates show high variance, DSL descriptions remain consistent)
-
Produce interpretable, human-readable placement decisions that can be audited and modified
Abstract
Indoor scene layout generation is a challenging task in interior design. Existing methods often oversimplify the task by reducing room conditions to coarse 3D bounding boxes and neglecting structural elements such as doors and windows. More fundamentally, many prior approaches formulate spatial reasoning as direct coordinate prediction, thereby casting interior layout design as continuous regression over raw geometric parameters, which hinders the model from learning the underlying reasoning logic of intelligent layout design. We propose LayoutDSL, a novel LLM-based framework for learning an interior layout policy in a domain-specific language (DSL) action space. The DSL provides an explicit symbolic representation of layout information and serves as a structured action space for layout reasoning, where each action corresponds to an interpretable design decision. Under this DSL-based policy learning paradigm, we construct 3D-FrontDSL, a dataset of room-structure annotations paired with synthetic DSL action sequences for supervised fine-tuning. To promote a more generalizable and scalable policy with verifiable feedback, we design rewards grounded in interior design principles and physical plausibility, and optimize the policy via reinforcement learning. Extensive experiments demonstrate that LayoutDSL substantially improves spatial plausibility and design logicality over strong baselines and existing methods.
Sources
- Open-Universe Indoor Scene Generation using LLM Program Synthesis and Uncurated Object Databases
- ReSpace: Text-Driven Autoregressive 3D Indoor Scene Synthesis and Editing
- SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- A Survey of Frontiers in LLM Reasoning: Inference Scaling, Learning to Reason, and Agentic Systems
- LLM Post-Training: A Deep Dive into Reasoning Large Language Models
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- InstructScene: Instruction-Driven 3D Indoor Scene Synthesis with Semantic Graph Prior
- DeepSeek-V3 Technical Report
- LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL
- HSM: Hierarchical Scene Motifs for Multi-Scale Indoor Scene Generation
- Direct Numerical Layout Generation for 3D Indoor Scene Synthesis via Spatial Reasoning
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
- SceneEval: Evaluating Semantic Coherence in Text-Conditioned 3D Indoor Scene Synthesis
- SpatialGrammar: A Domain-Specific Language for LLM-Based 3D Indoor Scene Generation
- Chat2Layout: Interactive 3D Furniture Layout with a Multimodal LLM
- Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models