FUSE: Frame-Unified Stress Estimation from Facial Video
Stefanos Gkikas, Thomas Kassiotis, Yang Guo, Guangliang Li, Giorgos Giannakakis
Honda Research Institute Japan · Hellenic Mediterranean University · Ocean University of China
cs.CV, cs.AI
Submitted: 2026-08-18
Updated: 2026-08-19
Comments: The paper has been accepted at: IEEE | 2026 9th International Conference on Pattern Recognition and Artificial Intelligence (PRAI 2026)
Code: https://github.com/ggian/stress
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: FUSE (Frame-Unified Stress Estimation) is a facial-video stress detection framework that processes complete recordings as a single input without temporal windowing or external segmentation.
Terminology
Summary
FUSE (Frame-Unified Stress Estimation) is a facial-video stress detection framework that processes complete recordings as a single input without temporal windowing or external segmentation. The name reflects the defining operation of the method: rather than dividing a recording into short clips, all frames are fused into one unified two-dimensional representation from which the stress state is estimated.
This unification is realized by folding the temporal dimension into the channel dimension of the spatial representation, and the resulting high-dimensional input is processed using a unified asymmetric-attention architecture.
The methodology works as follows: For a video of T frames captured at 30 fps and subsampled at temporal stride τ, the retained sequence contains L = ⌊T/τ⌋ frames, each a 224 × 224 RGB image. At a stride τ = 1 applied to a 120-second recording, this yields L = 3,600 frames, corresponding to the complete recording without any external windowing or temporal segmentation. The temporal dimension is folded into the channel dimension—a step referred to as axis folding—creating a tensor of shape H × W × 3L. Geometric information is incorporated by encoding each spatial position using Fourier features with K = 6 frequency bands and a maximum frequency fmax = 10, adding 26 positional features. The spatial axes are flattened into a sequence of N = 50,176 tokens, with the token dimension becoming C′ = 3L + 26. The token sequence is partitioned into S = 4 contiguous spatial groups of length 12,544 tokens each, which carry no correspondence to temporal windows of the input video.
The model processes the segmented token sequence through four layers, each comprising a cross-attention block followed by Rl self-attention blocks. A single latent state is associated with each spatial segment, instantiated by replicating a shared initialization vector derived from M0 = 32 learnable global parameters. In cross-attention, each segment state aggregates information exclusively from its corresponding token subset, with the query side consisting of a single vector while the key-value side spans many tokens—an asymmetric operation. After cross-attention, self-attention is applied across all S segment states to enable global information exchange. Across the four layers, the segment-state dimensionality decreases progressively from 128 to 80. After the final layer, the segment states are averaged and passed through a linear classification head. Both attention operations use pre-layer normalization and residual connections, with attention and feedforward dropout of 0.10.
Experiments were conducted on a 58-subject stress dataset (24 men, 34 women, aged 26.9 ± 4.8 years). The experimental protocol comprised four stress-induction phases: social exposure, emotional recall, mental workload, and stressful video stimuli, with 11 tasks total (4 neutral, 6 stress-inducing, and 1 relaxation task). Stress induction was verified through heart rate monitoring showing statistically significant increase during stress tasks (p < 0.05), and subjective validation was confirmed using Self-Assessment Manikin scales. Facial video was recorded at 60 fps and subsampled to 30 fps with a resolution of 608 × 800 pixels. Binary classification distinguishes neutral from stress conditions. Subjects were partitioned into training, validation, and testing sets at the subject level using a stratified split protocol, with 38 training, 8 validation, and 12 testing subjects, ensuring no participant appears in more than one set.
Seven temporal-stride configurations were evaluated, ranging from full-frame input (τ = 1) to sparse subsampling (τ = 30). At the densest setting (τ = 1), the full 120-second recording is retained at 30 fps, yielding 3,600 frames and a token channel dimension of 10,826. This configuration has the highest parameter count (9.16M), computational cost (348.78 GFLOPs), and inference latency (133.02 ms), with a throughput of 7.52 samples per second, yet achieves a test accuracy of 69.03%. At τ = 30, only 120 frames are retained, reducing the parameter count to 5.82M, GFLOPs to 12.48, and latency to 14.77 ms, corresponding to an approximately 28-fold reduction in compute relative to τ = 1.
The highest test accuracy is achieved at τ = 15 (69.44%), using 24.07 GFLOPs, while the full-frame configuration remains competitive at 69.03%. Validation accuracy peaks at τ = 20 (70.42%). The configuration τ = 5 produces the lowest test accuracy (60.56%) despite its second-highest validation accuracy (68.80%), reflecting per-stride generalization variability over the 12-subject test partition. The remaining configurations fall between 63.75% and 66.25% on the test set. Test accuracy does not decrease monotonically with stride, and temporal density beyond a moderate threshold does not yield consistent discriminative benefit.
The central capability demonstrated is that FUSE processes the complete facial recording as a single input, without temporal windowing or external segmentation.
At τ = 1, this corresponds to 3,600 frames from a 120-second recording ingested in one model pass. The internal spatial token segmentation does not correspond to temporal windows. The non-monotonic relationship between stride and test accuracy indicates that denser temporal sampling does not necessarily improve generalization
and that moderate subsampling can preserve stress-relevant facial information while substantially reducing computational cost.
The higher strides likely perform comparably because consecutive frames at 30 fps are largely redundant, so retaining every frame mainly enlarges the input projection without adding useful information.
A limitation noted is that the evaluation uses a single dataset and a binary neutral-versus-stress classification setting, and the study does not include a direct comparison against windowed baselines under the same subject-level split, which is left for future work. The results support the feasibility of full-recording inference but should not be interpreted as a complete replacement for all window-based video stress-recognition strategies.
Overall, FUSE shifts the unit of analysis from short temporal clips to complete facial recordings,
providing a simpler evaluation pipeline, avoiding choices about window length and aggregation, and preserving the ability to model stress-related facial information over the full duration of the recording.
Improvements for AI systems
Improvements to AI Systems:
-
Unified Temporal Encoding via Axis Folding: Replace windowing or frame-sampling modules with a mechanism that folds the entire temporal dimension into the channel dimension of a spatial tensor. This allows the model to ingest complete recordings (e.g., 3,600 frames) in a single forward pass, eliminating the need for temporal segmentation, window-length selection, or aggregation heuristics.
-
Asymmetric Cross-Attention with Latent Segment States: Introduce a small set of learnable latent vectors (e.g., 32 initial parameters) that are replicated per spatial segment. Use asymmetric cross-attention where a single query vector per segment aggregates information from many tokens (key-value side), followed by self-attention across segment states. This reduces computational overhead while preserving global context, enabling efficient processing of ultra-high-dimensional inputs (e.g., 50,176 tokens with 10,826 channels).
-
Fourier Feature Positional Encoding for Spatial Geometry: Incorporate Fourier features with multiple frequency bands (e.g., K=6, fmax=10) to encode spatial positions. This provides a rich, continuous geometric prior that helps the model capture subtle facial muscle movements and spatial relationships without relying on learned positional embeddings that may not generalize across resolutions or recording conditions.
-
Stride-Aware Adaptive Sampling for Efficiency: Implement a dynamic stride selection mechanism that learns or infers the optimal temporal subsampling rate per input or task. The paper shows that moderate subsampling (τ=15) matches or exceeds full-frame accuracy while reducing compute 28-fold. An AI system can use this to automatically trade off latency and accuracy based on available resources or real-time constraints.
-
Non-Monotonic Generalization-Aware Training: Use the finding that denser temporal sampling does not consistently improve accuracy to design training curricula that emphasize moderate temporal density. This can prevent overfitting to redundant consecutive frames and improve generalization across subjects, as evidenced by the performance gap between validation and test at τ=5.
What the Improved AI System Can Do:
-
Process entire video recordings (e.g., 2 minutes at 30 fps) as a single input without any pre-processing into clips, preserving full temporal context for stress detection or similar fine-grained facial analysis tasks.
-
Achieve state-of-the-art accuracy (≈69.4%) on stress classification while reducing computational cost by up to 28× compared to full-frame processing, enabling deployment on edge devices or real-time systems with latency as low as 15 ms.
-
Automatically adapt its temporal sampling density based on input characteristics or hardware constraints, maintaining robust performance across varying frame rates and recording lengths without manual tuning.
-
Generalize better across subjects by avoiding over-reliance on redundant consecutive frames, as demonstrated by consistent performance across different stride configurations.
-
Provide a simpler evaluation pipeline for video-based affective computing, removing the need to choose window lengths, overlap ratios, or aggregation methods, thus reducing human bias and engineering effort.
-
Scale to longer recordings (e.g., 10+ minutes) by leveraging the asymmetric attention and latent state design, which keeps memory and compute manageable even with tens of thousands of frames.
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models