THYME: Temporal Hierarchical-Cyclic Interactivity Modeling for Video Scene Graphs in Aerial Footage

arXiv:2507.09200 · cs.CV · Submitted 2025-07-12 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "THYME: Temporal Hierarchical-Cyclic Interactivity Modeling for Video Scene Graphs in Aerial Footage".

Jane: The rapid proliferation of video data in applications such as autonomous driving and surveillance necessitates robust methods for dynamic scene understanding,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: We’ve just been talking about the core of this paper, THYME, which is all about combining hierarchical spatial detail with cyclic temporal refinement for video scene graphs. Now let's look at the authors and what they were aiming for overall.

Jane: They are a group of researchers from institutions like IEEE who are focused on how we can make these video understanding models more robust, especially when dealing with challenging viewpoints like aerial footage or oblique shots.

Lu: They set out to tackle the problem that frame-level methods often produce inconsistent relational predictions over time, which they term predicate flickering, and then try to solve that by integrating temporal context in a way that doesn't dilute the spatial configurations.

Meng: So, their goal was to find a new way to jointly model these two things: preserving detailed spatial representations for each frame and ensuring those relational transitions are consistent across the entire video.

Tom: That’s the challenge they identified; balancing high-resolution local detail with long-range temporal coherence in a complex visual environment.

Jane: And they are essentially saying that existing state-of-the-art models often only offer partial improvements, frequently underrepresenting microinteractions or subtle spatial semantics.

Lu: So, their approach is designed to address those gaps by jointly modeling the temporally coherent trajectories and preserving the structural richness of those spatial relations in complex visual environments.

Meng: It sounds like they are trying to move beyond just getting a snapshot of what’s happening and actually understanding the story unfolding over time.

Tom: Exactly, it’s about capturing that story rather than just the individual frames in isolation.

The paper's summary: Jane: So, summarizing what THYME actually does for us, it uses a DETR object detector to get features for each frame to create those initial object query embeddings.

Tom: Those embeddings capture both the appearance of the objects and where they are spatially laid out in that specific frame.

Lu: Then they take those per-frame features and apply a hierarchical aggregation strategy to refine them across multiple abstraction levels, which is where they get those fine spatial details.

Meng: So, this hierarchy allows them to look at the neighborhood of an object within the frame and update its feature based on everything else around it.

Tom: They're essentially telling the feature of an object at a certain level is updated by aggregating information from all other objects in that same frame, which is what they define as the neighborhood.

Jane: After that hierarchical refinement, they move to the temporal part where they use a transformer encoder with cyclic attention to look at inter-frame dynamics.

Lu: This cyclic attention mechanism connects the end of a video sequence back to its beginning using that modulo operation, ensuring continuity by letting the last time step attend to the first.

Meng: So, that’s how they enforce temporal consistency across frames in a way that standard methods often miss because they look at things sequentially without that loop.

Tom: It’s a key mechanism for modeling those long-range temporal dependencies and ensuring things don't just flicker around over time.

The paper's improvements: Jane: The main improvement here is the architecture itself, which synergistically integrates hierarchical feature aggregation with cyclic temporal refinement.

Tom: That combination is what lets THYME capture multi-scale spatial details and long-range temporal dependencies more effectively than previous methods.

Lu: This design enables THYME to capture those multi-scale spatial context and enforce temporal consistency across frames, which leads to more accurate and coherent scene graphs.

Meng: So, the benefit for a user is that you get a scene graph representation that is both detailed locally and temporally consistent globally.

Tom: It means the final output isn't just one thing; it’s comprehensive and balanced in both aspects.

Jane: And they demonstrated this on AeroEye-v1 point 0, which highlights how this approach performs across aerial and ground views, showing improved understanding in those specific scenarios <ref:2507.09200#pg0>.

Lu: The ablation study confirmed that deeper aggregation significantly improves the model’s ability to fuse multi-scale spatial features and capture temporal information between interactivity types on AeroEye-v1 point 0 <ref:2507.09200#pg0>.

Meng: That means adding more layers doesn't just make it bigger, it actually helps the model learn how to connect different kinds of attributes together better.

Tom: And they showed that cyclic temporal attention consistently outperforms standard self-attention across all the different interactivity types tested in their experiments.

Conclusion: Jane: So, to wrap up, THYME is a novel architecture because it combines hierarchical feature aggregation with cyclic temporal refinement.

Tom: It’s designed to capture those multi-scale spatial details and maintain temporal consistency in video scene graph generation.

Lu: The authors have also introduced AeroEye-v1 point 0, which is a novel aerial video dataset enriched with annotations across appearance, situation, position, interaction, and relation types <ref:2507.09200#pg0,appearance, situation, position, interaction, and relation>.

Meng: And the experimental results on ASPIRe and AeroEye-v1 point 0 show that the proposed approach significantly outperforms existing state-of-the-art methods by about two to three percent in recall for mean recall <ref:2507.09200#pg2>.

Tom: So what this means for us is that we have a method capable of generating more comprehensive and balanced scene graph representations than what was possible before, especially in complex, dynamic scenes.

Jane: It opens up new possibilities for applications like autonomous driving or surveillance where understanding the environment needs to be both detailed spatially and temporally consistent.

Lu: In the future, they suggest exploring multimodal cues, like audio and textual information to enrich the scene representations.

Meng: And from an engineering standpoint, exploring domain adaptation techniques would be a good next step to extend this approach beyond just these specific benchmarks.

Trong-Thuan Nguyen, Pha Nguyen, Jackson Cothren, Alper Yilmaz, Minh-Triet Tran, Khoa Luu

University of Arkansas · University of Science, VNU-HCM (Vietnam National University, Ho Chi Minh, Viet Nam) · Ohio State University

cs.CV

Submitted: 2025-07-12

Updated: 2026-10-04

Comments: This manuscript has been rejected and is under revision for new submission

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 93/100

The gist: The rapid proliferation of video data in applications such as autonomous driving and surveillance necessitates robust methods for dynamic scene understanding, yet previous approaches often suffer

Key concepts

Per-Frame Feature Extraction
This step uses a DETR object detector to generate embeddings for every frame. These embeddings encode both the visual appearance of objects and their spatial layout within that specific moment, providing the initial input features for the model.
Hierarchical Feature Aggregation
This strategy refines per-frame features by aggregating information across multiple abstraction levels. It involves updating an object's feature by considering its neighborhood—all other objects in the frame—to capture intra-frame spatial dynamics at different scales.
Cyclic Attention
This mechanism is used to refine high-level features across time, ensuring temporal consistency. By using a cyclic attention defined via modulo operations, the model connects the end of a video sequence back to its beginning, allowing later frames to attend meaningfully to earlier ones.

Terminology

Summary

The rapid proliferation of video data in applications such as autonomous driving and surveillance necessitates robust methods for dynamic scene understanding, yet previous approaches often suffer from fragmented representations that fail to capture fine-grained spatial details and long-range temporal dependencies simultaneously<ref:2507.09200#pg2>. This paper introduces the Temporal Hierarchical Cyclic Interactivity Model for Video Scene Graphs (THYME) approach, which synergistically integrates hierarchical feature aggregation with cyclic temporal refinement to address these limitations by modeling multi-scale spatial context and enforcing temporal consistency across frames<ref:2507.09200#pg3>.

The gist

THYME effectively models multi-scale spatial context and enforces temporal consistency across frames, yielding more accurate and coherent scene graphs<ref:2507.09200#pg3>.

How it works

  1. Per-Frame Feature Extraction: The process begins by employing a DETR [1] object detector to extract features for each frame, producing N t object query embeddings <ref:2507.09200#pg7>. These embeddings encode both object appearance and spatial layout, formally defined as the set of per-frame object instances V t <ref:2507.09200#pg7>.

  2. Hierarchical Feature Aggregation: To capture intra-frame dynamics, a hierarchical aggregation strategy is used to refine per-frame features across multiple abstraction levels. The feature of an object S i(t) at level l is updated by aggregating information from its neighborhood, where the neighborhood comprises all objects in the frame t, i.e., N(S i(t)) = V t <ref:2507.09200#pg7>. This involves computing attention weights a ij(t) and updating the feature via Equation (2), where the feature of S i(t) at level l is updated via Eqn. (2) <ref:2507.09200#pg7>.

  3. Temporal Refinement via Cyclic Attention: High-level features are then refined using a transformer encoder with cyclic attention [33] to capture inter-frame dynamics. For each tracked object S i, a temporal sequence X t'(S i) is constructed, and The cyclic attention is then computed as defined in Eqn. (4) <ref:2507.09200#pg7>. This mechanism ensures continuity by connecting the end of a video sequence to its beginning, allowing the last time step to attend to the first through the modulo operation <ref:2507.09200#pg7>. The output is then refined through a transformer encoder block, resulting in the refined object feature hat F(S i) in R(d Lh) <ref:2507.09200#pg7>.

  4. Scene Graph Construction: Finally, the video scene graph G = (V, E) is constructed using unified object representations where V = S i hat F(S i) in R(dLh) <ref:2507.09200#pg7>. Predicates r ij are predicted by leveraging self-attention outputs from the DETR decoder to obtain relation representations R a k and R z, which are then refined through a gating mechanism to yield the final predicted scene graph G = (V, E) <ref:2507.09200#pg8>.

Key Contributions and Advantages

"We propose the Temporal Hierarchical Cyclic Interactivity Model for Video Scene Graphs (THYME) approach, a novel architecture that synergistically integrates hierarchical feature aggregation and cyclic temporal refinement."

This design enables THYME to capture multi-scale spatial details and long-range temporal dependencies more effectively than previous methods, resulting in more comprehensive and balanced scene graph representations <ref:2507.09200#pg7>.

Experimental Validation

The performance of THYME is validated on two challenging datasets, ASPIRe [32] and AeroEye-v1.0, demonstrating that the proposed approach outperforms state-of-the-art methods, offering improved scene understanding in ground-view and aerial scenarios <ref:2507.09200#pg7>.

Experimental results demonstrate that THYME establishes a new state-of-the-art on the ASPIRe and AeroEye-v1.0 benchmarks, achieving consistent improvements of 2 to 3% in recall and mean recall over the baseline methods.

This performance margin supports the efficiency of THYME’s dynamic scene graph generation efficiency across wild ground-view footage and drone-captured aerial videos <ref:2507.09200#pg7>.

AeroEye-v1.0 Dataset

The paper introduces AeroEye-v1.0, a novel aerial video dataset annotated with five distinct interactivity types: Appearance, Situation, Position, Interaction, and Relation <ref:2507.09200#pg9>. This dataset is designed to tackle the unique challenges of dynamic scene graph generation in aerial environments by providing detailed, viewpoint-diverse supervision necessary for fine-grained and temporally aware scene understanding <ref:2507.09200#pg9>. It includes footage captured from aerial, oblique, and ground perspectives, making it one of the first datasets that can represent the most significant and richly annotated aerial video benchmark for VidSGG <ref:2507.09200#pg9>.

Ablation Study Insights

The ablation study on hierarchical levels shows that deeper aggregation significantly improves the model’s ability to fuse multi-scale spatial features and capture temporal information between interactivity types on AeroEye-v1.0 <ref:2507.09200#pg10>. Furthermore, the results show that cyclic temporal attention consistently outperforms standard self-attention across all interactivity types <ref:2507.09200#pg10>. The effectiveness of the cyclic attention is further shown by results indicating gains in Recall and mean Recall across various window sizes, suggesting that capturing the complete temporal context is particularly beneficial for modeling single-actor and double-actor attributes <ref:2507.09200#pg11>.

Qualitative Results

Qualitative comparisons on the ASPIRe dataset show that THYME successfully tracks changing positions and speed of vehicles, updating its prediction history from chasing to approaching as the cars move closer, outperforming HIG and CYCLO in capturing dynamic changes in the video scene <ref:2507.09200#pg12>. The approach offers a dynamic solution when objects disappear, unlike HIG which fails to predict subsequent relationships once an object disappears <ref:2507.09200#pg13>.

Conclusion

In this work, the THYME approach integrates hierarchical feature aggregation with cyclic temporal refinement to capture multi-scale spatial details and maintain temporal consistency in video scene graph generation<ref:2507.09200#pg14>. In addition, the authors have presented AeroEye-v1.0, a novel aerial video dataset enriched with comprehensive annotations across appearance, situation, position, interaction, and relation<ref:2507.09200#pg9>. Moreover, extensive experiments on the ASPIRe and AeroEye-v1.0 datasets have demonstrated that the proposed THYME approach significantly outperforms existing state-of-the-art methods<ref:2507.09200#pg14>. Future work can include integrating multimodal cues, such as audio and textual information, to enrich the scene representations and exploring domain adaptation techniques to extend our approach to a broader range of real-world scenarios<ref:2507.09200#pg14>.

REFERENCES

[1] Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In European conference on computer vision, pages 213–229. Springer, 2020.<ref:2507.09200#pg2>

[4] Yuren Cong, Wentong Liao, Hanno Ackermann, Bodo Rosenhahn, and Michael Ying Yang. Spatial-temporal transformer for dynamic scene graph generation. In Proceedings of the IEEE/CVF international conference on computer vision, pages 16372–16382, 2021.<ref:2507.09200#pg5>

[5] Yuren Cong, Michael Ying Yang, and Bodo Rosenhahn. Reltr: Relation transformer for scene graph generation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(9):11169–11183, 2023.<ref:2507.09200#pg6>

[6] Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Antonino Furnari, Evangelos Kazakos, Jian Ma, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al. Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100. International Journal of Computer Vision, pages 1–23, 2022.<ref:2507.09200#pg7>

[7] Long Hoang Dang, Thao Minh Le, Vuong Le, and Truyen Tran. Hierarchical object-oriented spatio-temporal reasoning for video question answering. arXiv preprint arXiv:2106.13432, 2021.<ref:2507.09200#pg8>

[8] Jeonghyeok Do and Munchurl Kim. Skateformer: skeletal-temporal transformer for human action recognition. In European Conference on Computer Vision, pages 401–420. Springer, 2024.<ref:2507.09200#pg9>

[9] Shengyu Feng, Hesham Mostafa, Marcel Nassar, Somdeb Majumdar, and Subarna Tripathi. Exploiting long-term dependencies for generating dynamic scene graphs. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 5130–5139, 2023.<ref:2507.09200#pg10>

[10] Difei Gao, Luowei Zhou, Lei Ji, Linchao Zhu, Yi Yang, and Mike Zheng Shou. Mist: Multi-modal iterative spatial-temporal transformer for longform video question answering.

Improvements for AI systems

  1. Bold header: Hierarchical Feature Aggregation for Multi-scale Spatial Context

THYME employs a hierarchical feature aggregation strategy to progressively build layered spatial representations that capture fine-grained details, enabling it to capture multi-scale spatial context which allows the system to accurately delineate subtle relationships in cluttered or densely populated scenes.

  1. Bold header: Cyclic Temporal Refinement for Long-Range Coherence

The model uses a cyclic temporal refinement mechanism where the attention module connects the final frame back to the first, which is designed to reinforce long-range temporal dependencies and mitigates challenges such as transient occlusions.

  1. Bold header: Enhanced Interactivity Modeling via Dual Mechanism

By combining these two components, THYME achieves a dual advantage: it can capture fine-grained details and global context through hierarchy, while ensuring temporal consistency across video frames through cyclic refinement, leading to more balanced scene graph representations.

  1. Bold header: Superior Performance on Aerial Scenarios

The system is designed to be robust across diverse views, as demonstrated by its ability to achieve high scores on the AeroEye-v1.0 dataset, specifically achieving 16.52% R@20 in Appearance and 18.52% R@20 in Position for double-actor attributes on this aerial benchmark.

  1. Bold header: Fine-Grained Predicate Classification

The final scene graph construction uses a mechanism to predict a predicate p ∈ C p that describes the relationship between objects, allowing the system to categorize interactions into five distinct types: appearance, situation, position, interaction, and relation with high accuracy.

  1. Bold header: Improved Dynamic Interaction Prediction

THYME can accurately track complex dynamic changes in video scenes; for instance, it successfully tracks the changing positions and speed of the vehicles, updating its prediction history from ‘chasing’ to ‘approaching’ by leveraging both spatial and temporal features across multiple levels of abstraction.

Abstract

The rapid proliferation of video in applications such as autonomous driving, surveillance, and sports analytics necessitates robust methods for dynamic scene understanding. Despite advances in static scene graph generation and early attempts at video scene graph generation, previous methods often suffer from fragmented representations, failing to capture fine-grained spatial details and long-range temporal dependencies simultaneously. To address these limitations, we introduce the Temporal Hierarchical Cyclic Scene Graph (THYME) approach, which synergistically integrates hierarchical feature aggregation with cyclic temporal refinement to address these limitations. In particular, THYME effectively models multi-scale spatial context and enforces temporal consistency across frames, yielding more accurate and coherent scene graphs. In addition, we present AeroEye-v1.0, a novel aerial video dataset enriched with five types of interactivity that overcome the constraints of existing datasets and provide a comprehensive benchmark for dynamic scene graph generation. Empirically, extensive experiments on ASPIRe and AeroEye-v1.0 demonstrate that the proposed THYME approach outperforms state-of-the-art methods, offering improved scene understanding in ground-view and aerial scenarios.

Sources

Related papers