CountLoop: Training-Free High-Instance Image Generation via Iterative Agent Guidance

arXiv:2508.16644 · cs.CV · Submitted 2025-08-18 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "CountLoop: Training-Free High-Instance Image Generation via Iterative Agent Guidance".

Jane: COUNTLOOP introduces a training-free framework that achieves precise instance control in high-instance image generation by alternating between VLM-based planning and iterative refinement,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: Now that we’ve covered the title and high-level concept, let’s look closer at what CountLoop actually does, as described in the summary of CountLoop: Training-Free High-Instance Image Generation via Iterative Agent Guidance.

Jane: Essentially, this paper outlines a complete methodology where an AI system doesn't generate an image directly from a prompt but instead goes through a structured design phase and then refines that design based on how well the resulting image actually meets the criteria.

Lu: The summary highlights that the system operates as an iterative design process rather than just one single operation, emphasizing this loop between VLM-based planning and iterative refinement.

Meng: So, it’s not just a fancy text prompt anymore; it’s a multi-step reasoning system where one model plans the structure and another model judges the result.

Lalam: It's about taking complex requests, like "one hundred forty oranges and thirty-one birds in Harry Potter theme," and systematically breaking down the creation into spatial planning, synthesis, critique, and correction steps.

Tom: Precisely; it moves beyond simple text-to-image generation by injecting explicit structural control via a planning graph that captures object attributes like position and size.

Jane: This planning graph is the blueprint that tells the system exactly where every single orange or bird needs to go before any pixels are actually rendered in the final image.

Lu: The paper details how this graph, defined by nodes with attributes like normalized position and depth, is converted into a textual template that guides the initial synthesis step.

Meng: So, the first stage produces an image based on that blueprint, but it acknowledges that this single-pass method has limitations regarding attribute leakage when objects overlap heavily.

Lalam: That’s where the second half of CountLoop steps in, using instance-driven attention masking to ensure each object maintains its identity even when they are densely packed together.

Tom: It’s about sequential injection of latent features into the diffusion space, where each object is added based on its position in depth to build a final composite feature map.

Jane: This cumulative composition ensures that the final image isn't just a blob of colors but accurately represents every specified instance in its correct spatial relationship.

Lu: And then the refinement phase comes back in, where the Critic VLM looks at the output for things like counting accuracy and aesthetic quality to generate feedback for further planning adjustments.

Meng: So we are essentially using AI agents to act as a design team that constantly revises the plan based on expert critique until everything is perfect.

Lalam: That’s a really powerful concept because it addresses the failures of prior methods where they either under-produce or over-produce objects, which is what we see with count saturation and semantic leakage.

The paper's summary: Tom: Let’s talk about the specific advantages this paper claims CountLoop offers over existing approaches when looking at the improvements section of CountLoop: Training-Free High-Instance Image Generation via Iterative Agent Guidance.

Jane: The authors focus on several key areas, such as achieving count-faithful high-density scenes where prior methods often suffer from saturation or grid artifacts.

Lu: They highlight that they can produce images with precise object counts, even in extreme cardinalities like exactly one hundred forty oranges and thirty-one birds, which is a significant improvement over previous methods.

Meng: That level of accuracy—getting specific numbers right even when density is high—is what makes this useful for applications where precision matters for gameplay or data augmentation.

Lalam: Furthermore, they address the issue of semantic leakage by using instance-driven attention masking to ensure clear separation between instances in densely occluded scenes.

Tom: That’s important because it tackles the problem of feature mixing that happens when objects are too close together, something that existing methods struggle with because they lack explicit per-instance control.

Jane: They also claim guaranteed spatial coherence by using a Design VLM to iteratively refine the scene graph guided by natural spatial relationships, rather than rigid, auto-regressive arrangements.

Lu: This means objects are placed with realistic minimum distances and directional constraints derived from the planning graph, which results in layouts that look much more natural and less like a grid.

Meng: From an engineering view, getting those realistic spatial relationships baked in upfront should reduce the need for heavy post-processing to fix layout errors later on.

Lalam: And they provide training-free, parameter-efficient refinement through the text-editing operator, meaning we can correct structural flaws without needing to retrain or fine-tune the underlying diffusion backbone.

Tom: So, in short, these improvements boil down to achieving count fidelity, preventing attribute leakage through instance masking, and ensuring spatial coherence through iterative agent guidance.

Jane: And they show that this entire process can be done using a VLM as a structured critic to guide the iteration until the quality score hits a target threshold.

Lu: It’s about creating a system that leverages structured reasoning to overcome the inherent limitations of current text-to-image models in handling complex, high-instance scenes.

Meng: I just wonder how robust this is when we move beyond standard benchmarks to really unpredictable, open-world scenarios where the scene structure isn't perfectly defined by the initial prompt.

Lalam: That’s a fair question; the framework's strength relies on a well-defined initial plan from the Design VLM, so its success might depend heavily on how well that planner interprets ambiguity.

The paper's improvements: Tom: So we’ve walked through the structure of CountLoop: Training-Free High-Instance Image Generation via Iterative Agent Guidance, and it seems this framework offers a structured solution to the issues plaguing high-instance image generation today.

Jane: It really is a system that uses an initial plan from a Design VLM, followed by synthesis guided by instance masks, and then a feedback loop driven by an AI Critic VLM to ensure the final output is both accurate in count and visually coherent.

Lu: The paper demonstrates how to use structured reasoning to overcome architectural constraints like cross-attention failing to preserve per-instance identity through explicit spatial control mechanisms.

Meng: From a practical standpoint, the ability for parameter-free refinement without needing massive retraining suggests a much faster path toward deploying reliable generative tools in production environments.

Lalam: For the broader AI culture, this means we can build synthetic data that is perfectly count-labeled, which is invaluable for training other models and creating high-quality reference materials.

Tom: Absolutely; CountLoop: Training-Free High-Instance Image Generation via Iterative Agent Guidance shows a clear path to achieving precise instance control through VLM planning and iterative refinement.

Jane: It’s a solid piece of work because it systematically addresses the failures of scaling up to high instance counts by providing mechanisms for layout generation, masking leakage, and structural correction.

Lu: The implications are that we can start seeing more realistic scenes in generative models where object placement is governed by explicit spatial reasoning instead of just implicit visual correlations.

Meng: I think this moves us closer to systems that can handle highly specific demands in complex environments without needing an endless cycle of manual debugging to fix count errors or layout issues.

Lalam: It’s a huge step forward because it enables the creation of perfectly count-labeled synthetic data, which will significantly help us train open-world counting models with much tighter error bounds.

Conclusion: Tom: So to wrap up our deep dive into "CountLoop: Training-Free High-Instance Image Generation via Iterative Agent Guidance," we’ve seen how this framework uses a VLM planner and a Critic VLM to iteratively refine scenes, which is really smart stuff for tackling high-density image generation.

Jane: It’s fascinating because it moves away from simple single passes and uses that design-planning graph to guide the image synthesis, which helps keep things organized during the creation process.

Lu: I think what really stands out about CountLoop is how it handles the semantic leakage by using instance-specific attention masks derived from that planning graph, which is a very sophisticated way to ensure objects stay distinct even when they overlap a lot.

Meng: From my side, the training-free refinement loop is what impresses me most; being able to update the plan based on structured feedback without needing to retrain the whole backbone feels like a big win for deployment speed.

Lalam: I think this work has huge cultural implications because if we can generate synthetic data that’s perfectly count-labeled, it means we can build much more robust and reliable training sets for future multimodal AI systems.

Tom: Exactly, Lalam; that synthetic data will let us train models with much tighter bounds on accuracy than what we get from current methods.

Jane: And the spatial coherence aspect is also vital; getting those natural layouts instead of rigid grids makes the generated content feel genuinely realistic.

Lu: The cumulative latent composition method, where you sequentially inject features based on depth, is a neat trick that effectively builds a complex scene layer by layer in the diffusion space.

Meng: I’m curious though, how scalable is this iterative refinement process when we move from a few objects to hundreds of instances in real-world applications?

Lalam: Well, the ability to use it for creating count-supervised synthetic data gives us a solid way to test and verify the system's performance across those scales before deploying it widely.

Tom: That’s a fair question, Meng; scaling is always a factor in engineering, but this paper shows we have the structural tools to manage that complexity through the agentic guidance.

Jane: Overall, CountLoop: Training-Free High-Instance Image Generation via Iterative Agent Guidance presents a very practical approach to solving some of the hardest problems in generative AI right now.

Lu: I’m really looking forward to seeing how researchers build upon this framework to introduce even more complex spatial logic into those initial planning graphs.

Meng: I’ll keep an eye on how the parameter-free refinement operator evolves; that's where the real efficiency gains for engineers will come from.

Lalam: I hope this work inspires a new generation of tools where AI doesn't just create pretty pictures but creates precisely what we need for scientific and creative applications.

Anindya Mondal, Ayan Banerjee, Sauradip Nag, Josep Lladós, Xiatian Zhu, Anjan Dutta

University of Surrey · Universitat Autònoma de Barcelona (UAB) · Simon Fraser University

cs.CV

Submitted: 2025-08-18

Updated: 2026-09-29

Comments: Published at TMLR 2026

Code: https://github.com/black-forest-labs/flux

Project page: https://mondalanindya.github.io/CountLoop

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 90/100

The gist: COUNTLOOP introduces a training-free framework that achieves precise instance control in high-instance image generation by alternating between VLM-based planning and iterative refinement, which

Key concepts

Design VLM
This AI model interprets text prompts to create detailed planning graphs. It determines where each object should go in space, including its size, position, and relationships with other objects. This structured plan guides the subsequent image generation process.
Layout Aligned Attention Masking
This technique ensures that the AI focuses its attention only on the specific spatial area assigned to an object during image creation. It uses binary masks derived from the planned layout to confine features, preventing different objects' attributes from mixing and causing visual confusion in dense scenes.
Cumulative Latent Composition
Instead of generating one image at once, this method builds the final picture piece by piece within the AI's latent space. It sequentially places each object's appearance into the map, ensuring that nearer objects correctly overwrite farther ones without blending, leading to accurate final compositions.

Terminology

Summary

COUNTLOOP introduces a training-free framework that achieves precise instance control in high-instance image generation by alternating between VLM-based planning and iterative refinement, which resolves critical failures like count saturation and semantic leakage in dense scenes.

How it works

COUNTLOOP operates as an iterative design process rather than a single-pass operation, following a structured loop:

  1. A Design VLM interprets the prompt to produce realistic, non-grid layouts with natural object placement.

  2. These layouts guide image synthesis via a cumulative attention mechanism that mitigates attribute leakage and preserves object clarity under overlap.

  3. A Critic VLM assesses the output for counting accuracy and aesthetic quality, providing structured feedback to refine both the layout and prompt iteratively until the output meets target criteria.

VLM-Guided Layout Generation

The Design VLM parses the input prompt into a planning graph that captures both object attributes and spatial relationships. This graph is defined as G = (V, E, Bbg), where V denotes object-instance nodes with attributes like category, normalized position [x, y], size [w, h], depth prior d, and color. Edges in E encode spatial relations via directional operators (e.g., “above,” “left-of”), normalized distances, and angular orientations. This structured representation is converted into a textual prompt template PG = ϕ([′Object′]), [′Relation′], [′Context′]) which is fed into the Design VLM to produce a JSON output detailing object positions, depth, and sizes.

Layout Aligned Image Generation

To generate images that faithfully follow the specified arrangement while preventing attribute leakage, COUNTLOOP employs an iterative strategy inspired by multi-turn image generation. It adopts Layout Aligned Attention Masking by projecting discrete layouts into a continuous space using a layout encoder (GLIGEN). For each object instance i, a binary spatial mask Mi ∈ [0, 1] is derived from the layout li ∈ L, which is then reshaped into Mˆ i using bilinear interpolation to match the latent dimension of Across. The masked layout feature Ai mask is computed as A i mask = A i cross ⊙ Mˆ i, confining the receptive field of attention to the corresponding object’s spatial region and preventing feature mixing.

Cumulative Latent Composition

The framework builds a global latent feature map F by sequentially placing each object’s latent features in the diffusion latent space. For each instance i, the update is Fi+1(x, y) = 1(x,y)∈li ⊙A i mask + (1−1(x,y)∈li)⊙Fi. This process ensures that for every pixel covered by layout li, the existing feature is replaced by the instance-specific attention feature Ai mask. By composing instances in order of increasing depth (Far → Near), nearer objects overwrite farther ones without any spurious blending, resulting in a composite latent feature map F that faithfully encodes each object’s appearance and position according to the input layouts.

Layout Refinement via Iterative Feedback

After generating an image I, an iterative refinement loop is initiated where a Critic VLM reconfigures itself to analyze the output. It assesses two key aspects: (a) spatial coherence and appearance fidelity using a pretrained image encoder, and (b) counting accuracy by employing an off-the-shelf object detector. The Critic VLM computes a composite score S = α· sc + (1−α)· sa, where sc is the normalized count score derived from the detector's raw count cˆ and target count cgt, and sa is the aesthetic score from an external estimator. The Critic produces structured feedback in the form of text (Pfeed) which is used to update the planning graph G by employing a parameter-free textual refinement operator Ψ, which interprets this feedback to generate an updated graph G′ = Ψ(G, Pfeed, Popt). This process repeats until the composite score S exceeds a quality threshold τ or after a fixed number of rounds K.

Key Contributions and Results

The authors present COUNTLOOP as a training-free framework that achieves precise instance control through:

  1. A cumulative attention mechanism that sequentially injects each object in the latent space using instance-specific attention masks, effectively mitigating semantic leakage.

  2. A VLM as a structured critic to evaluate generated images along two axes: count consistency and appearance fidelity, providing interpretable feedback to refine the layout and prompt iteratively.

  3. Parameter-free refinement via an LLM-based text-editing agent Ψ that updates the planning graph through structured natural-language reasoning, compatible with any frozen diffusion model.

Improvements for AI systems

Here are the specific improvements that COUNTLOOP enables, and what an improved AI system can achieve:

  1. Improved Generation of Count-Faithful High-Density Scenes: The system can now generate images with precise object counts (e.g., exactly 140 oranges and 31 birds) even in extreme cardinalities where prior methods suffer from count saturation, semantic leakage, or grid artifacts.

  2. Mitigation of Semantic Leakage through Instance-Driven Attention Masking: The system can ensure clear object separation and prevent feature mixing between instances in densely occluded scenes by applying instance-specific attention masks derived from the planning graph layout.

  3. Guaranteed Spatial Coherence via Iterative Agent Guidance: The system can produce natural, non-grid layouts by iteratively refining a scene graph guided by a Design VLM, ensuring objects are placed with realistic spatial relationships (e.g., minimum distances) rather than rigid, auto-regressive arrangements.

  4. Scalable High-Instance Control Across Diverse Class Regimes: The system maintains robust count accuracy and high spatial quality across varying object counts (from low to high instance regimes), demonstrating a gap reduction of up to 57% on standard benchmarks and 43-48% on high-instance scenes.

  5. Training-Free, Parameter-Efficient Refinement: The system can adapt its layout and prompt structure in response to evaluation feedback using a parameter-free textual refinement operator (Ψ), meaning it can correct structural flaws (like object overlap) without requiring the retraining or fine-tuning of the underlying frozen diffusion backbone.

  6. Enhanced Data Augmentation for Object Counting Models: The system can be used to generate photorealistic, self-labeled training data by embedding category and cardinality in the prompt, which can then be used to fine-tune open-world counting models (like CountGD) with guaranteed numerical labels, significantly reducing the MAE/RMSE tail error.

  7. Creation of Count-Supervised Synthetic Data for T2V Models: The system can invert the data generation pipeline by taking a specification (e.g., 100 boxes on shelf A, 20 boxes on shelf B), generating the scene, verifying it with an open-vocabulary detector, and iteratively correcting it until exact per-class counts are met. This provides perfectly count-labeled supervision pairs for training text-to-video models.

  8. Versatile Style Control: The system can maintain precise object identities while allowing each instance to adopt distinct visual styles (e.g., photorealistic, anime, oil painting) through LoRA fine-tuning on the diffusion U-Net, enabling style transfer without altering core content or count fidelity.

Sources

Related papers