ESPLoRA: Enhanced Spatial Precision with Low-Rank Adaption in Text-to-Image Diffusion Models for High-Definition Synthesis

arXiv:2504.13745 · cs.CV, cs.AI · Submitted 2025-04-18 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "ESPLoRA: Enhanced Spatial Precision with Low-Rank Adaption in Text-to-Image Diffusion Models for High-Definition Synthesis".

Jane: Diffusion models have revolutionized text-to-image synthesis, yet they still struggle to properly render spatial relationships described in text prompts, which is particularly problematic for applications like urban planning.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: The title itself, ESPLoRA: Enhanced Spatial Precision with Low-Rank Adaptation in Text-to-Image Diffusion Models for High-Definition Synthesis, tells us exactly what the paper is about—improving spatial precision using low-rank adaptation to make high-definition synthesis better. It’s focused on a very specific technical mechanism.

Jane: That focus on low-rank adaptation is key because it shows they found a way to inject complex spatial understanding into these large models efficiently, avoiding the computational burdens associated with older external conditioning methods that the paper mentioned.

Lu: Andrea Rigo and her team at DISI, along with collaborators from Toretei S.r.l., MIT-IBM Watson AI Lab, and Fondazione Bruno Kessler are clearly drawing on a broad network of expertise to tackle this spatial problem from multiple angles, Lu.

Meng: I'm wondering how they balance the complexity of their data generation pipeline—filtering LAION-400M, using MLLMs like Molmo for context, and depth estimation models like Depth Anything—with keeping the fine-tuning process manageable on existing diffusion architectures.

Lalam: It’s fascinating to see this approach because it shows that the core issue isn't just about more data or bigger models; it’s about a more targeted, structurally informed way of fine-tuning to enhance a specific type of reasoning.

The paper's summary: Tom: So, in short, ESPLoRA builds on spatially explicit training data to create a novel T2I dataset, then they fine-tune diffusion models using LoRA on this synthetic set to significantly improve how well the models follow those spatial relationships described in the prompts.

Jane: That's the core idea—they generate millions of relationship-image pairs and then use that structured data, combined with a bias-aware transformation algorithm called TORE, to make existing T2I models much better at rendering things like "a balcony on the left of an external staircase."

Lu: The methodology involves a multi-stage pipeline where they extract objects using Grounding DINO and then develop geometric constraints based on those bounding boxes to derive relationships, which is a very systematic way to ground text in spatial reality.

Meng: I see the point about creating synthetic images that satisfy all the prompted relations; that validation step sounds crucial for ensuring the training data isn't just noisy object detection artifacts.

Lalam: The combination of strictly validated synthetic data and this geometric grounding seems like it provides a robust way to teach the model spatial rules rather than just letting it guess from raw text, which is really powerful for building more reliable visual understanding.

The paper's improvements: Tom: The authors highlight a few specific enhancements, one of which is the combination of their three elements: validated data, geometry-grounded extraction, and the bias-aware prompt transformation using TORE to exploit model biases towards certain relationships like "top" or "left."

Jane: That TORE algorithm is interesting because it doesn't change what the user asks for; instead, it flips a requested relationship to a higher-performing variant, like transforming "A bottom B" into "B top A," which boosts accuracy without altering the prompt’s meaning.

Lu: The paper also proposes refined evaluation metrics to assess three dee spatial relations using soft and strict accuracy scores based on geometric constraints, which gives us a clearer way to measure success than just looking at photorealism.

Meng: From an engineering impact view, this suggests that we can significantly improve performance on specific benchmarks like T2I-CompBench by leveraging these identified biases through TORE, leading to a measured improvement of up to thirteen point three three percent over previous methods.

Lalam: That systematic exploitation of model biases through TORE shows how we can guide the model's internal representations in a way that is highly effective for complex spatial tasks, which could really improve the fidelity of our visual AI systems.

Conclusion: Tom: So, to wrap up on ESPLoRA: they’ve shown that by combining curated spatial data with LoRA fine-tuning and the TORE transformation, they can significantly boost a diffusion model's ability to follow complex spatial instructions, achieving up to a thirteen point three three percent performance gain on benchmarks.

Jane: The main implication is that we can move towards more functional AI tools for urban planning and design where adhering to precise positional constraints is absolutely necessary for realistic modeling, which was a major limitation before this work.

Lu: It opens the door for us to think about how spatial reasoning can be encoded in generative models in a way that’s efficient, moving away from the high computational costs of external network conditioning.

Meng: Practically speaking, it means we can deploy more reliable generative systems for visualization tasks where strict adherence to design policies is required without needing massive re-training cycles on the base model itself.

Lalam: It really suggests that spatial understanding in AI isn't just an emergent property; it can be intentionally guided and enhanced through targeted fine-tuning methods like ESPLoRA, which has huge implications for how we build culturally aware visual agents.

Andrea Rigo, Luca Stornaiuolo, Mauro Martino, Bruno Lepri, Nicu Sebe

DISI, University of Trento · Toretei S.r.l. · Visual AI Lab, MIT-IBM Watson AI Lab, Cambridge, Massachusetts, U.S.A. · Fondazione Bruno Kessler

cs.CV, cs.AI

Submitted: 2025-04-18

Updated: 2026-01-20

Code: https://github.com/black-forest-labs/flux

Importance score: 88/100

The gist: Diffusion models have revolutionized text-to-image synthesis, yet they still struggle to properly render spatial relationships described in text prompts, which is particularly problematic for

Key concepts

ESPLoRA
Enhanced Spatial Precision with Low-Rank Adaptation in Text-to-Image Diffusion Models for High-Definition Synthesis. It is a method that improves how text-to-image models render spatial relationships by using low-rank adaptation and curated spatial data.
Low-Rank Adaptation (LoRA)
A technique used to efficiently inject complex spatial understanding into large diffusion models. It allows for fine-tuning the models without incurring the high computational costs associated with older external conditioning methods.
TORE
A bias-aware prompt transformation algorithm. Instead of changing what a user asks for, TORE flips a requested relationship to a higher-performing variant, such as transforming 'A bottom B' into 'B top A,' to boost accuracy.
Geometric Grounding
The process of extracting objects and developing geometric constraints based on their bounding boxes. This systematically grounds text in spatial reality by deriving relationships from the extracted object positions.

Terminology

Summary

Diffusion models have revolutionized text-to-image synthesis, yet they still struggle to properly render spatial relationships described in text prompts, which is particularly problematic for applications like urban planning. This paper introduces ESPLoRA (Enhanced Spatial Precision with Low-Rank Adaptation), a flexible fine-tuning framework that significantly enhances the spatial consistency of diffusion models by injecting specialized spatial knowledge without increasing generation time or compromising output quality. The authors propose a novel pipeline involving curated, spatially explicit training data and a bias-aware transformation algorithm to achieve superior performance on spatial benchmarks.

The Core Problem and Motivation

Existing text-to-image (T2I) diffusion models often fail to accurately represent the spatial relationships described in prompts, leading to incorrect generation of scenes with explicit positional constraints like on the left of or in front of. This limitation is pronounced in emerging AI-driven urban planning tools where adhering to specific design policies is crucial. Previous works typically rely on external network conditioning or predefined layouts, which incur higher computational costs and reduce flexibility. The authors address this by building upon a curated dataset of spatially explicit prompts, meticulously extracted from LAION-400M, to ensure precise alignment between textual descriptions and spatial layouts.

ESPLoRA Framework: Fine-Tuning with Low-Rank Adaptation

ESPLoRA is the proposed fine-tuning framework based on Low-Rank Adaptation (LoRA), specifically designed to enhance spatial consistency in generative models without increasing generation time or compromising output quality. The authors demonstrate that naively applying LoRA to spatial understanding yields suboptimal results. Their key insight is that the combination of three elements creates a synergistic framework:

  1. Strictly validated synthetic training data.

  2. Geometry-grounded relationship extraction from images.

  3. Bias-aware prompt transformation using the TORE algorithm to exploit model biases towards certain relationships (e.g., top, left, front).

Data Curation and Synthetic Dataset Generation

The methodology involves a comprehensive pipeline to create a novel urban-realistic database. This pipeline consists of several stages:

. Filtering the LAION-400M database keeping only image-text pairs where the caption contains a realistic urban scenario.

. Extracting all urban objects and the background urban context with the Multimodal LLM (MLLM) Molmo.

. Generating depth maps using the monocular depth estimation model Depth Anything.

. Using open vocabulary object detector Grounding DINO to detect objects in images, yielding bounding boxes.

. Developing a metric based on geometric constraints to extract relationships from bounding boxes.

The result is a dataset of approximately 5.6 million prompts derived from around 22,000 natural images. A second synthetic set is created by generating images with Flux.1 from these prompts and re-validating them, ensuring only images satisfying all prompted relations are kept, mitigating inaccuracies introduced by object detection models and spatial constraints.

Evaluation Metrics and Bias Exploitation

The authors propose refined evaluation metrics grounded in geometric constraints to capture 3D spatial relations such as in front of or behind. They introduce two primary metrics:

  1. Soft accuracy: Counting the sample as correct (score 1) if a specific relationship was correctly rendered, ignoring other relationships present in the prompt.

  2. Strict accuracy: Counting the sample correct if all relationships in the prompt are found correctly rendered by constraints.

Furthermore, they identify systematic biases in T2I models when rendering specific spatial relationships (e.g., top, left, front). They propose TORE (Transforming Original Relations Effectively), a pre-processing step that leverages these biases by flipping requested relationships to their corresponding higher-performing variants (e.g., transforming A bottom B into B top A), which increases generation accuracy without changing the prompt meaning.

Performance and Results

The authors evaluate ESPLoRA on two open-source state-of-the-art diffusion models, SDXL and Flux.1, using both simple (one relationship) and complex (two relationships) prompts. The results show that fine-tuning with ESPLoRA outperforms current baselines like CoMPaSS on spatial consistency benchmarks. Specifically, the combination of ESPLoRA and TORE leads to a 13.33% performance improvement over previous approaches. The method is robust to textual paraphrases of relations while retaining general prompt following capabilities, indicating genuine spatial understanding rather than overfitting to specific data distributions. The final comparison shows that Flux.1 + ESPLoRA + TORE achieves superior image quality compared to CoMPaSS, setting a new state-of-the-art on the evaluated spatial benchmark.

Limitations

The authors acknowledge two main limitations in their dataset construction pipeline:

  1. In some cases, they capture a relation such as a person next to a person while the corresponding image contains many people, which can lead the model to cheat by duplicating subjects.

Improvements for AI systems

Here are specific improvements that can be made to existing text-to-image (T2I) diffusion models by implementing the ESPLoRA framework, based on this research:

  1. Enhance Spatial Consistency in Generative Outputs: The core improvement is the ability to accurately render complex spatial relationships described in text prompts.

  2. Improve Performance on Urban Planning and Design Modeling: The system can generate realistic and functional urban scenarios that strictly adhere to specific design constraints (e.g., A balcony on the left of an external staircase, the balcony under a light).

  3. Achieve High-Definition Synthesis with Minimal Computational Overhead: Unlike methods relying on external network conditioning or LLM reasoning engines, ESPLoRA fine-tunes a lightweight Low-Rank Adaptation (LoRA) adapter, achieving spatial consistency improvement without increasing generation time or compromising output quality.

  4. Enable Flexible Spatial Conditioning: The framework allows for the injection of specialized spatial knowledge into pre-trained models without requiring extensive re-training of the base model, making it highly adaptable to different domains (e.g., applying the same pipeline to any domain with defined spatial relationships).

  5. Improve Generalization Across Textual Variations: The system becomes robust to alternative textual representations of spatial relations (e.g., transforming A right of B into B left of A), retaining general prompt-following capabilities while improving relational understanding, rather than overfitting to specific phrasing.

  6. Mitigate Systematic Spatial Biases: By incorporating the TORE algorithm, the system can exploit inherent biases in T2I models (e.g., a tendency to render certain relationships like top or left more accurately) to further improve performance on benchmarks like T2I-CompBench, leading to a measurable performance gain (up to 13.33% improvement over previous approaches).

  7. Establish Rigorous Spatial Evaluation: The system can be rigorously evaluated using novel, geometry-grounded metrics (like the Urban Benchmark) that assess both single and multi-relation prompts through strict geometric constraints on bounding boxes, providing a more precise quantification of spatial coherence than existing metrics.

In summary, the improved AI system will transition from merely generating photorealistic images to being a spatially intelligent generative agent capable of reliably interpreting complex positional language and producing geometrically accurate scenes in demanding applications like architectural visualization and urban simulation.

Sources

Related papers