Predictive Spatio-Temporal Scene Graphs for Semi-Static Scenes

arXiv:2605.00121 · cs.RO · Submitted 2026-04-30 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "Predictive Spatio-Temporal Scene Graphs for Semi-Static Scenes".

Dev: Predictive Spatio-Temporal Scene Graphs for Semi-Static Scenes addresses the challenge of enabling robots to perform complex reasoning across geometry and semantics in environments where objects exhibit semi-static changes over time.

Rosa: First, who's behind it and why it matters.

Paper summary: Rosa: We've discussed the concept of modeling semi-static dynamics over time, and this paper, "Predictive Spatio-Temporal Scene Graphs for Semi-Static Scenes," focuses on how to enable robots to reason across geometry and semantics in environments that exhibit structured temporal changes.

Dev: Essentially, the thesis is that existing spatio-semantic methods often lack the ability to reason about time at all, so this work proposes a way to incorporate temporal information into map updates.

Taro: I'm hearing that the key contribution isn't just tracking where things are now, but learning patterns of behavior—things like an object cycling through locations daily.

Rosa: That’s right, Taro; the paper claims they can learn these cyclic behaviors and use them to predict the future state of an object in a structured environment.

Dev: The method hinges on proposing Perpetua*, which is a persistence estimator that builds upon Perpetua but adds Bayesian model selection for better long-horizon predictions.

Taro: So, what does this Perpetua* estimator actually do in terms of modeling the dynamics? Does it just track presence or absence?

Rosa: It models the presence or absence of a specific feature over time using a mixture formulation to capture multiple persistence hypotheses simultaneously, including emergence filters alongside traditional persistence models.

Dev: And the switching mechanism for these hypotheses is governed by Bayesian model selection, which uses a new switching prior that can be informed by environment-specific observations or Large Language Models.

Taro: That external knowledge input from LLMs allows the estimator to make smarter choices about which model—persistence or emergence—is more likely given the current evidence.

Rosa: It means they leverage environmental context to select between models based on the marginal evidence, ensuring they keep the core strengths of Perpetua while overcoming its long-term forecasting limitations.

Dev: The resulting representation, called PredictiveGraphs, integrates this with an open-vocabulary scene graph structure to model these object-level semi-static changes over extended horizons.

Taro: So, instead of a static map, we get a dynamic structure where objects are connected to the receptacles they’ve been seen in over time.

Rosa: Exactly; each semi-static object is connected via edges to the set of receptacles it has been previously observed in, which forms the core of their edge set E j tN (<ref:2605.00121#pg2>).

Dev: The crucial part is that each edge (oj, ok) is associated with a corresponding binary persistence variable X j,k tN, which tells us if the object oj is present in receptacle ok at time tN.

Taro: And this allows them to define the goal: for any query time t greater than tN, they infer the posterior probabilities of those variables to find the most likely location or determine absence.

Rosa: That's right; so, given a scene graph G tN, they are trying to predict the environment’s state for a text query at time t > tN such as "where is my coffee mug?" (<ref:2605.00121#pg2>).

Dev: The entire goal is achieved by inferring those probabilities to identify the receptacle with the highest likelihood of containing the target object or determining that it's unlikely to be present at any receptacle, which could be an absence measurement.

Taro: It sounds like they’re essentially turning historical observations into a probabilistic model for future locations, which is really powerful for modeling routine behaviors.

Rosa: That's the essence of it; they are moving from simple spatial mapping to modeling temporal patterns driven by routine behaviors rather than pure randomness.

Dev: This framework directly addresses the fundamental trade-off between real-time tracking and long-horizon forecasting that plagued previous persistence estimators.

Conclusion: Rosa: So, wrapping up on this paper, "Predictive Spatio-Temporal Scene Graphs for Semi-Static Scenes," the main contribution is providing a representation called PredictiveGraphs that supports predictive tempo-spatio-semantic queries.

Dev: And the authors are proposing Perpetua*, which extends Perpetua with Bayesian model selection to handle persistence estimation more robustly across different scenarios.

Taro: The implication for autonomy is that robots can move from just reacting to what's there now, to proactively predicting where things will be in the future based on learned temporal patterns.

Rosa: Exactly, Taro; this means we’re not just seeing a static scene; we’re seeing a scene with learned temporal dynamics that allow for foresight.

Dev: The impact seems to be significant because it allows for more reliable planning in complex, structured environments where things are expected to change predictably over time.

Taro: If this works well outside the lab, the real-world implications could be huge for navigation systems operating in homes or even industrial settings where routines exist.

Rosa: I'm curious about how this translates into practical terms; it’s not just a theoretical framework, it’s meant to enable agents to actively scan and navigate based on these predictions.

Dev: The embodied LLM planning architecture that uses location prediction and active navigation tools shows they are aiming for an agent that can perform semantic search, predict locations, and then physically move toward the target.

Taro: And their validation results suggest this system maintains high performance under noisy perception conditions and even shows the ability to preemptively adapt navigation plans by anticipating blocked paths.

Rosa: So in simple terms, they've developed a method that models object-level semi-static dynamics and predicts future environment states for use in embodied planning.

Dev: The final point is that this work provides a way to move toward systems that can handle the temporal complexity of real-world scenes effectively by integrating persistence estimation with scene graphs.

Department of Computer Science and Operations Research, Universite de Montréal (University of Montreal) · Mila - Quebec AI Institute

cs.RO

Submitted: 2026-04-30

Updated: 2026-10-02

Comments: Accepted for publication in IEEE Robotics and Automation Letters (RA-L). Webpage at https://montrealrobotics.ca/predictive-graphs/

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 92/100

The gist: Predictive Spatio-Temporal Scene Graphs for Semi-Static Scenes addresses the challenge of enabling robots to perform complex reasoning across geometry and semantics in environments where objects

Key concepts

Perpetua*
This is a persistence estimator that uses Bayesian model selection to determine if an object's feature is present or absent over time. It models multiple hypotheses, including emergence and persistence, using environment-specific knowledge or LLMs to make informed switching decisions for better long-horizon predictions.
PredictiveGraphs
This is a spatio-temporal map representation that integrates Perpetua* with a scene graph. It attaches persistence estimators to objects to track their temporal patterns, allowing the system to model semi-static changes and accurately predict future object states, such as relocation or removal.
Temporal Edge Formulation
This method defines connectivity between semi-static objects and receptacles based on historical observations. The weight of these edges is determined by the Perpetua* belief computed via learned persistence dynamics, creating a time-varying scene graph that evolves as object patterns are learned.
Embodied LLM Planning
This architecture uses a LangChain framework to equip an agent with three tools: semantic search (using CLIP), location prediction (querying Perpetua* models), and active navigation. This enables the agent to perform semantic search, predict where objects will be, and navigate toward them actively.

Terminology

Summary

Predictive Spatio-Temporal Scene Graphs for Semi-Static Scenes addresses the challenge of enabling robots to perform complex reasoning across geometry and semantics in environments where objects exhibit semi-static changes over time. The core contribution is a method that integrates a persistence estimator with an open-vocabulary scene graph structure to produce PredictiveGraphs, allowing for predictive tempo-spatio-semantic queries in dynamic settings.

The gist: A new representation, PredictiveGraphs, integrates Perpetua∗ with an open-vocabulary scene graph to model object-level semi-static dynamics and predict future environment states.

Perpetua∗: The Persistence Estimator

Perpetua∗ is a persistence estimator that extends Perpetua by incorporating Bayesian model selection to improve long-horizon predictions while preserving real-time adaptability. It models the presence or absence of a specific feature over time using a mixture formulation to capture multiple persistence hypotheses, including emergence filters alongside persistence models. The switching mechanism is governed by Bayesian model selection, which uses a new switching prior that can be informed by environment-specific observations or Large Language Models (LLMs). This allows the estimator to leverage environment-specific knowledge or LLMs to select between the emergence and persistence models based on the marginal evidence, ensuring it maintains the core strengths of its predecessor while also overcoming its predictive limitations.

PredictiveGraphs: The Tempo-Spatio-Semantic Representation

PredictiveGraphs is an open-vocabulary map representation that integrates Perpetua∗ with a scene graph structure (ConceptGraphs) to model object-level semi-static changes over extended horizons. The method achieves this by attaching persistence estimators to objects in the map to track their temporal patterns. The key insight is that in structured operational settings, objects exhibit a temporal structure driven by routine behaviors rather than randomness. This integration allows the representation to (1) break the staticity assumption and accurately model semi-static objects, (2) predict future object states, including relocation or removal from the environment, and (3) maintain continuously updated based on learned temporal patterns.

Temporal Edge Formulation

To encode spatio-temporal object-receptacle relationships, a dense edge formulation is adopted. For each semi-static object, connectivity is established to the set of receptacles where it has been historically observed. Formally, for each semi-static object oj ∈ OS tN, we establish connectivity to the set of receptacles Rj = ⟨ok ∈ OR tN A(oj, ok) = 1⟩. The edge weight between an object and a receptacle at query time is given by the Perpetua∗ belief computed via (13), where this belief evolves according to learned persistence dynamics. This yields a time-varying scene graph G tN = (V tN, E tN), where the edge weights evolve according to learned persistence dynamics.

Online Updates and Querying

The system is designed for continuous operation through online updates and temporal querying. When new observations are acquired, the Perpetua∗ estimators are updated in real-time by constructing a local scene graph and applying classification procedures to generate presence measurement (yt = 1) for each semi-static object oj found on ok or an absence measurement. For queries at time t > tN, the system recomputes edge weights using the corresponding Perpetua∗ estimators. This enables querying "the map at arbitrary future times t > tN, even in the absence of new observations by recomputing the edge weights for all semi-static object-receptacle pairs."

Embodied LLM Planning

The final component is an embodied planning architecture that integrates this temporal object map with a LangChain framework. This architecture equips the agent with three tools: (1) Semantic Search using CLIP to identify objects matching open-vocabulary queries; (2) Location Prediction, which queries the underlying Perpetua∗ models to obtain a probability distribution over receptacles for a given object at some time t; and (3) Active Navigation, which navigates toward the predicted receptacle. This allows for semantic search, location prediction, and active navigation, enabling the agent to actively scan the environment and terminate early if the target object is detected or deemed absent. The system is validated in real-world tasks where it demonstrates superior performance over baselines under both privileged and noisy perception conditions.

Evaluation Summary

Validation across simulation and real-world dynamic navigation tasks demonstrates that PredictiveGraphs outperforms baselines in predicting future environment states, even in the presence of distributional shifts. The method shows robust performance in adaptive navigation, achieving high Success Rates (SR) and Success Weighted by Path Length (SPL) compared to methods like ConceptGraphs or DualMap. Furthermore, the system's ability to preemptively adapt its navigation plan by anticipating blocked paths is highlighted in real-world experiments.

Improvements for AI systems

As a fastidious researcher, I have analyzed the provided paper, Predictive Spatio-Temporal Scene Graphs for Semi-Static Scenes. The proposed system is a novel framework called PredictiveGraphs that integrates a Bayesian persistence estimator called Perpetua∗ with open-vocabulary scene graphs to enable tempo-spatio-semantic reasoning.

Here are the specific improvements and capabilities this AI system can provide:


) Improve Long-Term, Cyclic, and Contextual Reasoning in Dynamic Environments:

The system moves beyond traditional change detection (which only maintains a hypothesis about the present) to learn and predict semi-static dynamics driven by routine behaviors (e.g., a mug moving between a cupboard, countertop, and sink).

The improved AI system can:

  1. Maintain long-term memory of object states by leveraging the learned temporal patterns of objects across multiple observations.

  2. Predict the future state of an object (its relocation or removal) at arbitrary future times, even when no new observations are available (i.e., during periods between sensor readings).

  3. Reason contextually: Given a query like where is my coffee mug?, the system can use the historical temporal patterns to predict its most likely location across all known receptacles, rather than just relying on outdated spatial maps.

) Enable Robust and Adaptive Navigation in Semi-Static Settings:

The integration of Perpetua∗ allows for proactive, plan-aware navigation that anticipates environmental changes.

The improved AI system can:

  1. Generate a prioritized list of potential locations (receptacles) to visit based on the probability distribution predicted by Perpetua∗, rather than just the last known location.

  2. Proactively adapt its path planning: If a predicted receptacle is determined to be blocked or if an object is deemed absent from all likely receptacles, it can select a longer but feasible alternative route (as demonstrated in Fig. 5), avoiding unnecessary exploration and preventing task failure due to map obsolescence or dynamic shifts.

  3. Handle Long Weekend Scenarios: The system can adapt its predictions when the underlying environmental dynamics suddenly shift (e.g., Monday and Friday adopting weekend behaviors), demonstrating robustness against distributional shifts that plague standard prediction models.

) Achieve Open-Vocabulary, Semantic Scene Understanding with Temporal Awareness:

By attaching temporal persistence estimators to an open-vocabulary scene graph (using CLIP embeddings for object retrieval), the system achieves high semantic grounding across time.

The improved AI system can:

  1. Answer natural language queries about objects that are not explicitly named in the initial map (e.g., where is my healthy vegetable?). It uses semantic search (CLIP) to identify candidate objects and then uses Perpetua∗ to predict their location, effectively grounding a query in time and space simultaneously.

  2. Maintain a dense, temporally rich scene graph where edges represent learned spatio-temporal relationships (an object has been observed at a receptacle). This allows for complex reasoning like what is the typical behavior of this object? or where has this item been over the last three days?.

) Enhance Real-World Deployment Reliability Under Noise:

The system is specifically designed to maintain performance even when perception pipelines introduce noise or errors.

The improved AI system can:

  1. Maintain high success rates (e.g., 90.0% Success Rate in Table IV for the noised setting) during real-world navigation, even when trained on noisy data or operating under imperfect perception conditions.

  2. Correct outdated map information by recomputing edge weights dynamically using the current Perpetua∗ belief, ensuring that decisions are based on the most up-to-date temporal understanding of the environment rather than relying solely on a stale global map.

Related papers