Architect-Ant: Editable Automatic Furnishing of Architectural Floor Plans

arXiv:2606.10953 · cs.AI, cs.CV · Submitted 2026-06-09 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Architect-Ant: Editable Automatic Furnishing of Architectural Floor Plans".

Jane: Architect-Ant introduces a framework for furnishing residential floor plans by treating furniture layout synthesis as structured sequence generation over editable geometric objects,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: Let's talk about the title and who's behind this work. "Architect-Ant: Editable Automatic Furnishing of Architectural Floor Plans" tells us exactly what they did, and we see a team from KAUST in Saudi Arabia leading the charge on this research.

Jane: It’s interesting to see how they've framed it—not just as an automated placement tool, but something that maintains editability throughout the entire process, which is a key distinction. It's not about generating a final render; it's about generating the blueprint for the furniture arrangement.

Lu: The combination of geometric data and structured text generation suggests they are addressing a real bottleneck in existing automatic furniture arrangement tools, where datasets often lack those specific object-level annotations that make layouts truly editable.

Meng: They built AntPlan-two hundred seventy as a curated dataset, which sounds like they recognized that having high-quality, annotated data is the foundation for any reliable system in this domain.

Lalam: Having a curated dataset focused on ten residential room categories shows a very focused effort to build something practical and usable rather than just chasing broad visual appeal. It’s about building specific tools for specific needs.

The paper's summary: Tom: Now, let's look at what the paper actually summarizes regarding the methodology behind Architect-Ant. Essentially, they used a multi-stage pipeline to train a large vision language model to generate these structured layouts by adapting it using pseudo-labeled data first.

Jane: So, it starts with getting a prior understanding of furniture categories and spatial relations, then fine-tunes the model on those pseudo-labeled layouts to get it used to the specific text format they defined.

Lu: The core innovation here is using that structured output—the line-oriented grammar—to allow for direct computation of constraints like clearance and wall affinity, which means validity checks happen directly on the geometry, not after a visual rendering.

Meng: The training pipeline involves a rule-based evaluator that scores samples based on geometric and semantic criteria, such as checking for containment or door obstruction, which provides explicit design preferences to guide the learning process.

Lalam: That rule-based scoring system is really smart because it translates abstract human design rules into quantifiable signals that the AI can learn from directly, which is a powerful way to bake in functional plausibility early on.

The paper's improvements: Tom: The paper points out several key improvements they made over previous methods, and one big suggestion is moving away from just training the model to better incorporating a preference optimization step using Direct Preference Optimization, or DPO.

Jane: They emphasize that traditional model-pair DPO might lead to reward hacking if you're not careful; instead, they use a synthetic-pair approach where only one bounding box is perturbed at a time, which localizes the preference signal.

Lu: This targeted learning means the model learns the precise effect of changing a single placement decision while keeping other things constant, which is crucial for refining quality in complex arrangements like kitchens or bathrooms.

Meng: They also highlight that using VLM-as-judge can provide an independent signal on functional layout quality, catching subtle visual failures that the hard geometric rules might miss.

Lalam: The paper suggests a separation between the editable structured DSL and the final blueprint visualization, which means you get both a flexible design tool and a high-fidelity rendering, which is exactly what interior designers need.

Conclusion: Tom: We've gone through the title, the summary of their training process, and the specific improvements they introduced in this paper on Architect-Ant. Overall, it looks like they’ve successfully created a system that produces layouts that are both geometrically sound and functionally plausible.

Jane: It seems the main implication is making furniture layout synthesis a structured sequence generation task rather than just an image generation task, giving users true control over the output.

Lu: I think the future potential lies in how this DSL representation can be used to build more complex generative systems that handle not just furniture, but dynamic architectural elements too.

Meng: For practical impact, having a system where we can tweak a layout by editing the text description rather than re-running an entire training process is something I’m looking forward to seeing deployed in the industry.

Lalam: This work really sets a high bar for AI that needs to operate within strict physical constraints, and this paper on Architect-Ant shows a clear path toward that kind of reliable, usable generation.

King Abdullah University of Science and Technology (KAUST) · Miami University

cs.AI, cs.CV

Submitted: 2026-06-09

Updated: 2026-09-30

Comments: 26 pages

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 86/100

The gist: Architect-Ant introduces a framework for furnishing residential floor plans by treating furniture layout synthesis as structured sequence generation over editable geometric objects, combining

Key concepts

Structured Sequence Generation
This approach treats placing furniture not as creating a single image, but as generating an ordered list of editable geometric instructions. Each instruction specifies an object's type, class name, and precise coordinates. This allows the system to build complex layouts step-by-step using a compact Domain-Specific Language (DSL) that remains easily modifiable.
Deterministic Rule Scoring
This is a scoring mechanism that evaluates candidate furniture placements based on explicit geometric and semantic rules. For example, it checks for object overlap, clearance around doors, or wall affinity. These scores provide clear feedback on how well a layout adheres to predefined design constraints.
Direct Preference Optimization (DPO)
This is a training technique used to refine the model's output by optimizing it directly against human preferences. Instead of just maximizing a score, DPO trains the model to assign higher probabilities to layouts that have been explicitly scored as better by the rule-based evaluator, improving visual quality and functional plausibility.
Domain-Specific Language (DSL)
The DSL is a compact text format used to represent furniture arrangements. It looks like 'FURNITURE OBJ class=<snake_case> x=<m> y=<m> w=<m> h=<m>', which describes every piece of furniture and its exact location. This allows the layout to be both machine-readable for generation and editable by designers.

Terminology

Summary

Architect-Ant introduces a framework for furnishing residential floor plans by treating furniture layout synthesis as structured sequence generation over editable geometric objects, combining pseudo-labeled data with deterministic rule scoring and direct preference optimization to produce geometrically valid and functionally plausible layouts. This method addresses the limitation of existing automatic furniture arrangement by focusing on generating object-level representations in a compact Domain-Specific Language (DSL) that preserves editability, allowing for workflows in real estate visualization and interior design.

Data Preparation and Representation

The research utilizes AntPlan-270, a curated dataset of 270 architectural floor plans with per-room furniture bounding box annotations across ten residential room categories. This dataset is constructed by extracting structural primitives (walls, doors, windows) using an RT-DETR-X detector trained on CubiCasa5K and bootstrapping furniture bounding boxes from a hand-labeled subset. The resulting data is decomposed into room-level samples, where each sample carries the room geometry in metric coordinates and a furniture pseudo-label list with per-instance bounding boxes. The output layout is represented using a compact line-oriented grammar: FURNITURE OBJ class= x= y= w= h=, which allows for direct computation of geometric and combinatorial constraints like overlap, containment, clearance, door obstruction, reachability, wall affinity.

Model Training Pipeline

The training involves a multi-stage process to adapt a pretrained vision-language model (Qwen3.5-9B) to the target representation. The stages include:

  1. Prompting for an initial prior over furniture categories and coarse spatial relations.

  2. A lightweight fine-tuning stage on pseudo-labeled layouts to adapt the model to the target format and approximate room statistics using pseudo-labeled layouts.

  3. A rule-based evaluator that scores sampled layouts using geometric and semantic criteria, such as containment within the room, object overlap, door access, traversable paths, wall affinity, which supplies explicit design preferences.

  4. Preference optimization over candidate object placements to further refine layout quality by training the model to assign higher probability to better-scoring layouts.

Inference and Selection

At inference time, Architect-Ant operates through a two-step process: first, the model samples K DSL candidates per prompt using a temperature of 0.9 and top-p of 0.95; second, the deterministic rule scorer ranks these candidates and selects the highest-scoring one as the system output. This ensures that layouts are selected based on explicit geometric and semantic criteria, which are decomposed by rule family to provide a traceable preference signal for training.

Evaluation Metrics

The performance of Architect-Ant is evaluated using two complementary views. The headline view is the per-room mean ± standard deviation of scores across the K candidates per prompt, reflecting typical generation quality and consistency. The secondary view is best-of-K, the average score of the best candidate per prompt selected by the scorer, which reflects the inference-time protocol. Furthermore, an independent visual judge (Gemini 3 Flash Preview) is employed to provide a separate evaluation signal on functional layout quality, focusing on criteria such as wall intersections, door and passageway clearance, functional grouping and wall hugging.

Key Findings

Experiments on out-of-distribution datasets like CubiCasa5K show that Architect-Ant produces geometrically valid and functionally plausible layouts. The synthetic-pair Direct Preference Optimization (DPO) recipe is used as the main recipe, as it localizes each preference to a single perturbed bounding box, whereas broader model-pair DPO can lead to reward hacking by maximizing the rule score without improving visual quality. The VLM-as-judge study indicates that while hard rule scores are useful, softer spatial preferences—such as placing chairs more symmetrically around tables or tightening kitchen groupings—are often captured by the judge, suggesting that DPO improves visual-functional quality in specific categories like the bathroom and kitchen. The system successfully maintains a separation between the editable structured DSL and the downstream blueprint-style visualization, where a domain-specific diffusion model renders the symbolic layout.

The gist: Architect-Ant generates geometrically valid and functionally plausible layouts by training a vision-language model on pseudo-labeled data, using deterministic rule scoring to derive preference signals for direct preference optimization, and selecting the highest-scoring layout at inference time.

How it works

  1. The system formulates furnished room layout synthesis as structured sequence generation over editable geometric objects rather than image generation.

  2. It adapts a pretrained generator to this representation using pseudo-labeled layouts, providing a task-specific starting point for later preference optimization.

Improvements for AI systems

Here are the specific improvements that can be made to existing AI systems, based on the principles and methodology presented in this paper:


The core improvement is shifting from generating raw pixel outputs (images) to generating structured, editable data (a Domain-Specific Language or DSL) that is inherently verifiable and controllable.

  1. textbfFrom Image Generation to Structured Layout Synthesis (Architect-Ant Framework):

  2. A new AI system can take a simple floor plan geometry and a list of required furniture items, and instead of outputting a final rendered image, it outputs an editable text structure (DSL) describing the exact position, size, and type of every object. This DSL is the source of truth for all subsequent architectural workflows.

  3. textbfFrom Black-Box Generation to Constraint-Driven Learning (Rule-Based Scoring):

  4. The AI's learning process is augmented by a deterministic rule-based evaluator that assigns explicit penalties based on geometric and semantic constraints (e.g., wall clearance, door swing, object overlap). This moves the system beyond purely statistical accuracy to ensure functional validity.

  5. textbfFrom Post-Hoc Repair to Training Supervision (Preference Optimization via DPO):

  6. The AI is trained not just on correct outputs, but on a preference signal derived from the rule scorer. By using Direct Preference Optimization (DPO), the model learns to assign higher probability to layouts that satisfy these explicit geometric and functional constraints during training, rather than relying solely on external checks applied after generation.

  7. textbfFrom Single-Objective Learning to Multi-Signal Supervision (Procedural Reasoning Traces):

  8. The training data is enriched with procedural reasoning traces that show the step-by-step decision-making process for placing each object, including in some cases, on-the-fly correction of a placement error. This provides the model with a rich supervisory signal about architectural constraints (like wall affinity and circulation) that simple image labels cannot convey.

  9. textbfFrom Rule Scores Alone to Visual Plausibility Assessment (VLM-as-Judge):

  10. A secondary, independent evaluation loop can be established where a powerful Vision-Language Model (VLM) acts as a judge. This VLM compares two generated layouts based on high-level functional quality criteria (e.g., Is the arrangement plausible for this room type?), providing an orthogonal signal to the rule scorer to catch subtle, non-geometric visual failures that hard rules miss, especially in complex scenarios like kitchens.

  11. textbfFrom Simple Perturbations to Targeted Learning (Synthetic Pair DPO):

  12. To ensure the model learns the precise effect of a single placement decision, the DPO training uses synthetic pairs where only one bounding box is perturbed while keeping the procedural trace identical. This localizes the preference signal to a single placement change, preventing reward hacking where models exploit superficial differences in trace style or surface form.

In summary, these improvements result in an AI system that generates layouts that are:

  1. Functionally valid (passes hard geometric checks).

  2. Editor-friendly (output is structured text/DSL).

  3. Learned robustly (trained to follow explicit rules via DPO).

Abstract

Furnished floor plans support real-estate visualization, interior design, and architectural workflows, yet automatic furnishing remains challenged by limited real-world data and the need to satisfy interacting geometric and functional constraints. We ask whether professional furnishing knowledge can be learned from real floor plans using a pretrained model, enabling direct constraint-aware layout generation without relying on costly iterative agentic inference. We introduce AntPlan, a curated dataset of 505 real professional architectural floor plans with dense furniture annotations spanning 92 object classes and ten residential room categories, and Architect-Ant, a framework for generating furniture layouts. Architect-Ant represents layouts with an editable coordinate-based DSL and first learns professional furnishing patterns through supervised fine-tuning. It is then optimized with GRPO using a Layout Rule Score (LRS) that aggregates geometric and functional constraints derived from professional plans, providing outcome-level supervision without prescribed reasoning traces. Experiments against diverse state-of-the-art baselines show that Architect-Ant combines low geometric violation rates with high functional completeness, while qualitative results more closely reflect real-world residential furnishing patterns. The resulting layouts remain object-level editable and can be converted into 3D scenes.

Sources

Related papers