iARCS: Iterative Agentic RL for Controllable 3D Scene Generation

arXiv:2608.06161 · cs.AI · Submitted 2026-08-13 · Read on arXiv

Saugat Adhikari, Ashok Prasad Neupane, Pramish Paudel, Ajad Chhatkuli, Danda Pani Paudel

Pulchowk Campus, Institute of Engineering, Tribhuvan University · NAAMII · INSAIT, Sofia University St. Kliment Ohridski

cs.AI

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: 15 pages, 9 figures, 4 tables. Includes appendix

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 59/100

The gist: The paper presents iARCS, an "iterative agentic reinforcement learning framework that adapts a pretrained scene generator to natural-language task requirements." The work addresses a core limitation

Terminology

Summary

The paper presents iARCS, an iterative agentic reinforcement learning framework that adapts a pretrained scene generator to natural-language task requirements. The work addresses a core limitation in synthetic 3D scene generation: "existing generators often optimize perceptual realism without reliably satisfying task-critical functional constraints. This mismatch limits the usefulness of synthetic data for downstream training, where accessibility, traversability, and spatial rule compliance are often essential."

The authors are Saugat Adhikari, Ashok Prasad Neupane, Pramish Paudel, Danda Pani Paudel, and Ajad Chhatkuli, affiliated with Pulchowk Campus (IOE, Tribhuvan University), NAAMII (Kathmandu), and INSAIT (Sofia University). The paper is arXiv:2608.06161v1.

The paper motivates synthetic 3D indoor scene generation as increasingly important for computer vision and embodied AI, where large-scale and diverse training data are essential but expensive to collect and annotate in the real world. While recent generative models have significantly improved perceptual realism and distribution matching in scene generation, the authors note that many downstream applications require scenes that satisfy functional constraints beyond semantic plausibility, citing examples such as accessibility-aware clearance, controllable traversability for robot testing, and explicit spatial rules over object size, count, and relative placement.

The paper identifies a practical data bottleneck because "benchmark datasets and unconstrained generators rarely provide enough task-critical edge cases. Consequently, generated training data can appear realistic while remaining weakly aligned with downstream functional utility."

Existing approaches are reviewed across three lines: procedural generation (provide scalable variation, but their handcrafted rules make fine-grained task semantics hard to specify and verify consistently), LLM-augmented construction (reliability still depends on LLM scene composition ability, with weak guarantees under complex multi-constraint settings), and data-driven generative models (ATISS, DiffuScene, MiDiffusion, PhyScene) which match training distributions and improve plausibility, but still struggle to satisfy user-defined functional constraints. The paper argues that Text conditioning alone is usually insufficient for strict functional scene requirements and that Reward-guided optimization improves controllability, yet existing practice still depends on manually engineered objectives and substantial reward-design effort, which limits clean scaling to arbitrary new constraints.

The paper defines a 3D scene as an unordered set of N objects, S = o1, o2,..., oN, where each object is parameterized as oi = (ti, si, cos(θi), sin(θi), ci), with ti ∈ R3 the centroid translation, si ∈ R3 the 3D dimensions, θi the orientation, and ci the one-hot encoded object category. Letting F denote the floor boundary and C a user-specified functional constraint expressed in natural language, the generation process is a floor-plan-conditioned policy πθ(xF). The objective is to find adapted parameters θC maximizing the expected composite reward: θC = arg max Ex∼πθ(xF) [Rtotal(S, C, F)].

iARCS operates in two distinct optimization stages:

  • Stage 1: Base Model Refinement with Universal Rewards. "Pretrained generative models often internalize dataset flaws such as object collisions. To rectify this, we first optimize the base policy πθ using a set of universal, rule-based rewards (Ru). These enforce basic physical laws and structural integrity."

  • Stage 2: Functional Adaptation via Agentic Feedback. "Once a physically plausible baseline is established, we introduce the user constraint C. An LLM agent translates the natural language prompt into an executable reward program (Rt). The model is then optimized using the composite reward Rtotal, supported by an iterative reward reflection loop to prevent reward hacking and ensure stable learning."

The total composite reward is a weighted sum: Rtotal = Σ wu Ru + Σ wt Rt, where The weights wu and wt are chosen by the LLM and adjusted through reward reflection. In experiments, uniform weights are used (Ru = Rt = 1).

The LLM agent (the paper uses Gemini[37] model as LLM function generator) translates user prompts into executable reward programs via a rigorous three-step procedure:

  1. Reasoning: The agent identifies critical, restrictive constraints and spatial dependencies implied by the prompt. For example, given narrow bedroom with 0.2m walking path, the agent reasons that Furniture must be pushed against the walls to ensure a continuous 0.2m wide aisle from the door to the bed.

  2. Decomposition: Abstract constraints are broken down into atomic, measurable geometric checks, such as Walking Space = 0.2m (Check), Furniture Alignment (against the walls), and Path Clearance (continuous aisle).

  3. Reward Function: The agent outputs executable Python code (e.g., def robot 3DFRONT walkability reward(scene):...) that computes a scalar reward based on the generated scene and the decomposed metrics.

To address reward hacking or objectives too sparse for the current model state, the Reward Reflection module operates after fixed RL iterations: the LLM inspects the Reward Statistics and top-down projection images of representative sampled scenes. It may either Redefine (Debug and simplify the revised reward functions if the underlying logic is flawed or too hard for current policy) or perform Curriculum Generation (Decompose the reward into a simpler, intermediate objective to 'warm up' the model before introducing the full constraint). Reflections are performed every 10 epochs of RL training loop.

The diffusion denoising process is treated as a multi-step Markov Decision Process (MDP) and optimized using Denoising Diffusion Policy Optimization (DDPO). By utilizing RL, we successfully optimize the generator for complex, non-differentiable geometric constraints that standard gradient-based guidance methods cannot handle. LoRA is used for efficient fine-tuning with a rank (r) of 16 and a scaling factor (α) of 16, integrated into all attention projection layers (q proj, k proj, v proj, out proj) within the Transformer blocks, with 20 DDIM steps (η = 1.0) for rollouts, the Adam optimizer at a learning rate of 1 × 10−5, and early stopping to find an optimal operating point that balances the trade-off between diversity, quality, and rewards.

The universal rewards are manually designed targeting collision avoidance, boundary adherence, object density regularization, and functional accessibility, implemented as four algorithms in the appendix: Collision Avoidance Reward (penalizing 3D bounding-box overlap), Boundary Adherence Reward (penalizing objects outside the floor polygon), Accessibility Reward (measuring reachable free-space coverage and clearance via flood fill with agent radius 0.3), and Object Count Diversity Reward (KL divergence against the training distribution of object counts).

The experiments use 3D-FRONT (a widely used synthetic dataset containing 6,813 professionally designed indoor scenes), with the base generator being a continuous domain-only MiDiffusion model pretrained on 3D-FRONT operating in object-parameter space with canonical CAD models from 3D-FUTURE. All experiments use bedroom scenes. Baselines are ATISS (an autoregressive transformer for indoor scene synthesis) and MiDiffusion (a continuous-only variant of a mixed discrete-continuous diffusion model).

Evaluation metrics include: FID (computed via CLIP embeddings on top-down projections), Scene Classification Accuracy (SCA), object-collision rate Colobj, scene-collision rate Colscene, out-of-bound placement rate Rout, reachable-object ratio Rreach, and walkability score Rwalkable (the ratio between the area of largest connected walkable region and the total walkable area). The test set is filtered to scenes satisfying task constraints with a non-penetration threshold of −0.25.

Compared with baselines, iARCS consistently improves physical plausibility metrics and attains higher Rwalkable than the original data distribution. Key results: iARCS achieves Colobj of 40.45% (vs. 54.12% for ATISS, 52.67% for MiDiffusion, and 42.00% for ground truth), Colscene of 64.63%, Rout of 3.04%, Rreach of 87.82%, and Rwalkable of 0.861. The paper notes that 3D-FRONT dataset itself contains poor-quality scenes with frequent collisions and functional issues, and that This observation highlights a broader limitation of relying only on FID and SCA when evaluating function-critical scene generation. iARCS's FID (1.60) is slightly lower than some baselines (ATISS 1.39, MiDiffusion 1.34), attributed to quality issues in the raw data distribution, while SCA (68.51%) is the highest.

Training MiDiffusion on 3D-FRONT augmented with 4,000 iARCS-generated synthetic scenes (using training-set floor layouts, fine-tuned for 200 epochs at lr 1e−5) yields clear improvements in both physical plausibility and functional utility while maintaining similar diversity: Colobj drops from 52.67% to 41.49%, Colscene from 81.67% to 63.61%, Rout from 5.89% to 3.12%, Rreach rises from 85.7% to 92.52%, Rwalkable rises from 0.806 to 0.8272, with FID matched at 1.34. The paper emphasizes these gains are achieved without any external data or additional guidance, showing that the method can improve the generator through self-augmentation alone.

Three task-specific settings are evaluated: (1) Robot grasping scene where all support surfaces are within a 1.0 m vertical reach limit, (2) Room with a TV stand positioned so a farsighted person can view a 4K TV from bed, and (3) A bedroom scene with a functional study zone. For each task, iARCS is compared against 3D-FRONT* (the subset of 3D-FRONT satisfying the constraint). Lower FID and competitive SCA relative to 3D-FRONT* indicate that each task-adapted iARCS policy recovers diversity beyond strict dataset filtering while respecting task-specific constraints. For example, in the farsighted TV task, iARCS achieves FID of 2.905 vs. 4.706 for 3D-FRONT*; in the study zone task, FID of 2.629 vs. 14.49. The paper states this demonstrates its utility for constraint-aware synthetic data generation and augmentation.

The two-stage strategy is ablated on the task Scene with TV and bed at a 3 m distance. Under the same reward budget, "Directly optimizing task-specific constraints with physical plausibility, functional utility, and task-specific rewards yields worse results than our two-stage strategy: first training with universal rewards (physical plausibility + functional utility), then jointly fine-tuning with task-specific rewards." Single-stage achieves Colobj of 62.17% and Rwalkable of 0.679, whereas two-stage achieves Colobj of 38.79% and Rwalkable of 0.744.

The paper lists two main limitations: "First, reward quality depends on LLM-generated task decomposition and reward-code synthesis; ambiguous prompts can produce suboptimal or incomplete constraints, which may affect optimization reliability. Second, although our two-stage training improves physical plausibility and functional utility, it introduces additional compute cost due to iterative RL fine-tuning and repeated reward evaluation."

The paper's contributions are: (1) iARCS, an iterative agentic RL framework for controllable 3D scene generation from natural-language constraints; (2) a two-stage training strategy (universal-reward pretraining + task-specific fine-tuning) that improves physical plausibility and functional utility; and (3) demonstration of effective task-specific constraint optimization and that iARCS-generated data improves a base generator while maintaining competitive diversity.

The conclusion states: "We presented iARCS, an iterative agentic RL framework for controllable 3D indoor scene generation that combines LLM-based task decomposition, executable reward synthesis, and staged policy refinement. Experiments show that iARCS improves physical plausibility and functional utility while maintaining competitive distribution quality, and that iARCS-generated data can further improve a base scene generator through self-augmentation. We also demonstrated strong task-conditioned generation, where iARCS achieves lower FID than constraint-satisfying dataset subsets, indicating improved diversity under user-specified constraints. Future work aims to extend generalization beyond 3D-FRONT-style indoor distributions to broader domains, including physically grounded 4D generation tasks."

The base denoising network uses a Transformer architecture with 8 layers, an embedding dimension of 512, 4 attention heads, a feed-forward dimension of 2048, and a dropout rate of 0.1, with GELU activation, adaptive layer normalization, and a PointNet-based architecture with layer dimensions [4, 64, 64, 512, 64] for floor plan feature extraction, following LEGO-Net. RL fine-tuning uses the Adam optimizer with a learning rate of 3 × 10−4, batch size of 32, executed on a single NVIDIA RTX 2070 Ti GPU using mixed-precision FP16. The reward modeling uses a 'gemini-2.5-pro' model as a high-level reward generator, augmented by the universal rewards.

Improvements for AI systems

Here are specific improvements to AI systems inspired by iARCS, and what the improved system would be able to do.

Instead of manually designing reward functions for each new constraint, the AI system uses an LLM agent to translate natural-language requirements into executable reward code. A reflection loop monitors reward statistics and sampled outputs, then automatically simplifies, debugs, or decomposes rewards into easier intermediate objectives.

What the improved AI system can do:

  • Accept novel, high-level user constraints such as “bedroom with a 0.2 m walking path from door to bed” and generate scenes that satisfy these constraints without new human reward engineering.

  • Detect and correct reward hacking by reviewing intermediate outputs and adjusting reward definitions during training.

  • Automatically create curricula when constraints are too difficult, first solving a simpler version and then gradually introducing the full requirement.

The generator first optimizes universal, rule-based rewards — collision avoidance, boundary adherence, accessibility, and object-count diversity. Only after physical plausibility is established does the system add task-specific rewards. This prevents task optimization from destroying basic physical realism.

Generated scenes that pass functional checks are added back to the training data and used to retrain the base generator. This closes the loop: the generator improves itself without any external data or additional guidance.

Instead of relying only on FID and scene classification accuracy, the system evaluates generated scenes using geometric and functional metrics — collision rates, reachable-object ratios, walkability, and boundary adherence. These metrics are used as rewards and as filters for deciding which generated data is useful.

Diffusion models are treated as multi-step MDPs and optimized with Denoising Diffusion Policy Optimization, using LoRA adapters to keep fine-tuning computationally light. This makes iterative, constraint-specific adaptation feasible on a single consumer GPU.

Overall, these improvements yield an AI system that can generate controllable, physically plausible, task-compliant 3D scenes from natural-language instructions, refine its own generator through self-augmentation, and adapt to new functional constraints with minimal human effort.

Sources

Related papers