HSTGFormer: Hyper Spatial-Temporal Graph Transformer for 3D Human Pose Estimation

arXiv:2608.12187 · cs.CV, cs.AI · Submitted 2026-08-12 · Read on arXiv

Ruochen Li, Shuang Chen, Wenke E, Farshad Arvin, Amir Atapour-Abarghouei

Durham University

cs.CV, cs.AI

Submitted: 2026-08-12

Updated: 2026-08-13

Comments: Accepted to BMVC 2026, full paper

License: http://creativecommons.org/publicdomain/zero/1.0/

Importance score: 100/100

The gist: HSTGFormer is a graph-enhanced Transformer framework proposed for efficient monocular 3D human pose estimation.

Terminology

Summary

HSTGFormer is a graph-enhanced Transformer framework proposed for efficient monocular 3D human pose estimation. The paper states: "Transformer-based methods have achieved strong performance in monocular 3D human pose estimation, but most existing approaches organise spatial and temporal reasoning as separate stages, which may weaken unified spatial-temporal interdependencies inherent in human motion and compress frame-level structural information before temporal modelling. To address this, the authors propose HSTGFormer, a graph-enhanced Transformer framework that reformulates spatial-temporal reasoning as localised coupled graph aggregation over joint-time nodes."

The method introduces two complementary graph structures. First, the "Hyper Spatial-Temporal Graph (HSTG), which decomposes global spatial-temporal reasoning into local spatial-temporal receptive fields around individual joint-time nodes by extending per-frame skeleton graphs into temporal neighbourhoods, thereby enabling structure-aware coupled reasoning while preserving local structural motion information. Second, it further incorporates an Adaptive Dual-Scale Temporal Graph (ADSTG) to capture joint-specific temporal dependencies over complementary short- and long-range windows. A lightweight node-wise fusion module further adaptively integrates the two graph representations for each joint-time node."

The paper's contributions are listed as: "We propose HSTGFormer, a graph-based Transformer framework for efficient 3D human pose estimation, which reformulates spatial-temporal modelling from a graph perspective to better capture motion continuity and structural dependencies. We introduce a Hyper Spatial-Temporal Graph (HSTG) structure (Sec. 3.2) that extends per-frame skeleton graphs to localised temporal neighbourhoods, enabling structure-aware modelling of local motion continuity while preserving human skeletal priors. We design an Adaptive Dual-Scale Temporal Graph (ADSTG) module (Sec. 3.3) to construct and fuse dynamic temporal graphs over complementary temporal ranges, allowing joint-specific and adaptive temporal dependency modelling."

For HSTG, the paper explains: "Given the embedded pose representation Xh ∈ RT ×J×D, we construct a Hyper Spatial-Temporal Graph Ghyper = (V, Ehyper), where each node vt, j ∈ V represents joint j at timestep t, and V = T J. From the perspective of each ego node vt, j, HSTG defines a local spatial-temporal neighbourhood by jointly considering anatomical connectivity and temporal proximity. The adjacency is Kronecker-factorised as Ahyper ≈ Atem ⊗ Aspa, where Aspa ∈ 0, 1 J×J denote the skeleton adjacency matrix with self-loops and Atem ∈ 0, 1 T ×T denote a temporal band adjacency with entries 1, if t − t ′ ≤ w and 0, otherwise. The reasoning is implemented via two factorised masked graph attention operations: skeleton-constrained spatial attention followed by local temporal attention. The output is Xhyper = Xh + σ LN H + X̄h Wu".

For ADSTG, the paper states: "ADSTG constructs dynamic temporal graphs along the temporal dimension... we define two centred temporal windows: Nm (t) = t ′ t − t ′ ≤ rm, m ∈ short, long... ADSTG computes dot-product similarities between node (t, j) and its temporal candidates (t ′, j), and retains the top-km most similar candidates to form a sparse dynamic temporal adjacency. The propagation uses a normalised temporal graph GCN defined as TempGCN(Xh, A) = ÃXh Wv with à = D− 1/2 (A + I)D− 1/2. The two scales are combined via a lightweight topology-aware gate based on the skeleton degree producing Xada = ∑ m∈ short,long Gm ⊙ Z m".

The fusion module is described: "we introduce a context-aware fusion module that predicts node-wise weights for each joint-time node (t, j)... we concatenate the token feature with Cspa and Ctem, and feed the result into a lightweight MLP followed by softmax: ωhyper, ωada = Softmax MLP Xh, Cspa, Ctem. The final graph representation is then computed as Xfuse = ωhyper ⊙ X̂hyper + ωada ⊙ X̂ada."

The training objective is L = Lmpjpe + λs Lnmpjpe + λv Lvel + λd Ldiff + λlb Llb, with λs = 0.5, λv = 20, λd = 0.5 and a load-balancing regularisation Llb = (1/K) ∑ r=1 K ((1/omega) ∑ i∈omega ωri − 1/K) 2.

Experiments are conducted on Human3.6M [9] and MPI-INF-3DHP [24]. On Human3.6M, our method achieves highly competitive performance... obtaining the best P-MPJPE of 31.5 mm and matching the best MPJPE of 37.9 mm among all compared methods. Compared to KTPFormer, our model reduces MPJPE from 40.1 mm to 37.9 mm and P-MPJPE from 31.9 mm to 31.5 mm, corresponding to relative improvements of 5.5% and 1.3%, respectively. Compared to MotionAGFormer-L, our method improves MPJPE from 38.4 mm to 37.9 mm and P-MPJPE from 32.5 mm to 31.5 mm. In terms of efficiency, "compared with TCPFormer... our method obtains comparable MPJPE and better P-MPJPE with 59.5% fewer parameters and 46.5% lower MACs/frame. It also reduces parameters and MACs/frame by 25.3% and 25.5% compared with MotionAGFormer-L."

On MPI-INF-3DHP, "our method achieves the best overall performance. Compared with TCPFormer... our method improves AUC from 87.7 to 89.3 and reduces MPJPE from 15.0 mm to 14.0 mm (-6.7%). Compared with MotionAGFormer-L, our method improves AUC by 4.0 points and reduces MPJPE by 13.6%."

Ablation studies show: Compared with the baseline without HSTG and ADSTG, adding either component brings clear improvements, reducing MPJPE from 41.1 mm to 39.2 mm and 38.8 mm, respectively. Replacing node-wise adaptive fusion with fixed 0.5/0.5 weights degrades MPJPE from 37.9 mm to 38.3 mm, and Removing the load-balance loss also weakens the performance. For HSTG design, factorised MLPs lead to a clear performance drop, with our design reducing MPJPE by 4.5% and P-MPJPE by 4.2%, and Factorised GCNs perform better than MLPs but still underperform our design. For ADSTG, replacing adaptive temporal graph modelling with fixed dual-scale temporal convolutions or fixed temporal attention also degrades performance.

Per-action analysis shows "our method achieves the best results on several representative actions, including Direction, Eating, Purchases, Sitting, Sitting Down, Walking Dog, and Walking Together, while remaining competitive on Discussion, Phone, Photo, and Pose. Qualitative visualisations confirm our method produces more structurally consistent and anatomically plausible 3D poses under challenging body articulations and self-occlusions."

The paper concludes: "In this paper, we propose HSTGFormer, a graph-enhanced Transformer framework for efficient monocular 3D human pose estimation... Extensive experiments on Human3.6M and MPI-INF-3DHP demonstrate that HSTGFormer achieves strong accuracy while maintaining high computational efficiency and low memory cost. Further ablation studies and qualitative analyses verify the effectiveness of the proposed graph reasoning design. Preliminary zero-shot results on humanoid robots and bees also suggest its potential applicability beyond human pose estimation."

Improvements for AI systems

Improvements to AI systems:

  1. Unified spatial-temporal reasoning via localized graph aggregation: Replace separate spatial and temporal processing stages in Transformer-based models with a hyper graph that couples spatial and temporal dimensions around each joint-time node. This preserves local skeletal structure while capturing motion continuity, reducing information loss before temporal modeling.

  2. Adaptive dual-scale temporal dependency modeling: Implement dynamic temporal graphs with short- and long-range windows per joint, using dot-product similarity to select top-k temporal neighbors. This allows joint-specific, adaptive temporal receptive fields instead of fixed or global temporal attention, improving accuracy on actions with varying speeds (e.g., Sitting Down vs. Walking).

  3. Kronecker-factorized graph attention for efficiency: Decompose the spatial-temporal adjacency into a Kronecker product of skeleton and temporal band matrices, then apply factorized masked graph attention. This reduces computational cost while maintaining structure-aware reasoning, enabling deployment on resource-constrained devices.

  4. Node-wise adaptive fusion with topology-aware gating: Fuse hyper graph and dual-scale temporal graph outputs using a lightweight MLP that predicts per-node weights, conditioned on skeleton degree and temporal context. This avoids fixed weighting and improves robustness across diverse poses, with a load-balancing loss to prevent mode collapse.

  5. Graph-based inductive bias for 3D pose estimation: Use skeletal priors (anatomical connectivity) and temporal locality as inductive biases, reducing reliance on large training data and improving generalization to unseen domains (e.g., zero-shot on humanoid robots and bees).

What the improved AI system can do:

  • Achieve state-of-the-art 3D human pose estimation accuracy (MPJPE 37.9 mm, P-MPJPE 31.5 mm on Human3.6M) while being 59.5% more parameter-efficient and 46.5% lower in MACs/frame than comparable Transformer methods.

  • Maintain high accuracy on challenging actions with complex articulations and self-occlusions, producing anatomically plausible poses.

  • Adapt temporal reasoning per joint and per action speed, improving performance on both fast and slow movements.

  • Operate efficiently on edge devices due to factorized graph operations and low memory footprint.

  • Transfer to non-human pose estimation tasks (e.g., animal or robot pose) with minimal fine-tuning, thanks to graph-based structural priors.

Abstract

Transformer-based methods have achieved strong performance in monocular 3D human pose estimation, but most existing approaches organise spatial and temporal reasoning as separate stages, which may weaken unified spatial-temporal interdependencies inherent in human motion and compress frame-level structural information before temporal modelling. In this paper, we propose HSTGFormer, a graph-enhanced Transformer framework that reformulates spatial-temporal reasoning as localised coupled graph aggregation over joint-time nodes. Specifically, HSTGFormer introduces a Hyper Spatial-Temporal Graph (HSTG), which decomposes global spatial-temporal reasoning into local spatial-temporal receptive fields around individual joint-time nodes by extending per-frame skeleton graphs into temporal neighbourhoods, thereby enabling structure-aware coupled reasoning while preserving local structural motion information. It further incorporates an Adaptive Dual-Scale Temporal Graph (ADSTG) to capture joint-specific temporal dependencies over complementary short- and long-range windows. A lightweight node-wise fusion module further adaptively integrates the two graph representations for each joint-time node. Experiments on Human3.6M and MPI-INF-3DHP show that HSTGFormer achieves strong accuracy with high computational efficiency.

Related papers