GenStream: Semantic Streaming Framework for Generative Reconstruction of Human-centric Media
cs.CV, cs.AI, cs.MM
Submitted: 2026-09-16
Updated: 2026-09-16
Comments: 9 pages. Published at ACM MM 2025. Code: https://github.com/emanuele-artioli/genstream
Journal ref: In Proceedings of the 33rd ACM International Conference on Multimedia 2025 (MM '25). ACM, New York, NY, USA, 12276-12284
Code: https://github.com/emanuele-artioli/genstream
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: Video streaming dominates global internet traffic, yet conventional pipelines remain inefficient for structured, human-centric content such as sports, performance, or interactive media.
Terminology
Abstract
Video streaming dominates global internet traffic, yet conventional pipelines remain inefficient for structured, human-centric content such as sports, performance, or interactive media. Standard codecs re-encode entire frames, foreground and background alike, treating all pixels uniformly and ignoring the semantic structure of the scene. This leads to significant bandwidth waste, particularly in scenarios where backgrounds are static and motion is constrained to a few salient actors. We introduce GenStream, a semantic streaming framework that replaces dense video frames with compact, structured metadata. Instead of transmitting pixels, GenStream encodes each scene as a combination of skeletal keypoints, camera viewpoint parameters, and a static 3D background model. These elements are transmitted to the client, where a generative model reconstructs photorealistic human figures and composites them into the 3D scene from the original viewpoint. This paradigm enables extreme compression, achieving over 99.9% bandwidth reduction compared to HEVC for the continuous data stream. We partially validate GenStream on Olympic figure skating footage and demonstrate potential for high perceptual fidelity under minimal data. While acknowledging the significant computational costs shifted to the client and challenges in generalization, GenStream opens new directions in volumetric avatar synthesis, canonical 3D actor fusion across views, and personalized viewing experiences, laying the groundwork for scalable, intelligent streaming in the post-codec era.
Sources
- BlazePose: On-device Real-time Body Pose tracking
- AI-Driven Innovations in Volumetric Video Streaming: A Review
- Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network
- DVC: An End-to-end Deep Video Compression Framework
- SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
- SAM 2: Segment Anything in Images and Videos
- SfM-Net: Learning of Structure and Motion from Video
- FAST: A Framework to Accelerate Super-Resolution Processing on Compressed Videos
- Implicit-explicit Integrated Representations for Multi-view Video Compression
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models