Broaden Your Views for Self-Supervised Video Learning
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Broaden Your Views for Self-Supervised Video Learning".
Jane: This paper introduces BraVe, a self-supervised learning framework for video that breaks temporal symmetry by training networks to predict a broad view from a narrow temporal window,
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So, wrapping up our look at "Broaden Your Views for Self-Supervised Video Learning," we're talking about this work by Recasens and the team. They introduce BraVe as a way to learn video representations by training models to predict a broad view from a narrow temporal window zero.
Jane: The core idea boils down to forcing the model to generalize from short clips to the entire video context, which helps it understand event structure better one <ref:2103.16559#pg0>. It’s about using that temporal dimension explicitly instead of ignoring it.
Lu: The authors suggest this framework is promising because they tackle a hard problem: extrapolating to the general context of an event one <ref:2103.16559#pg0>. They show how this can be done by defining those two regression losses, Ln→b and Lb→n zero.
Meng: From a practical view, the fact that it works across different modalities like audio or optical flow in the broad view means we don't have to lock ourselves into just one type of data for context learning zero. That adaptability is what makes it useful for diverse datasets.
Tom: It really comes down to how you structure the training task. This paper shows that by setting up this specific narrow-to-broad prediction, you get representations that are better at understanding the overall video scene than traditional methods zero.
Jane: So, if we simplify it for someone just listening in, it's about teaching the AI to look beyond a few seconds and figure out what the whole story is one <ref:2103.16559#pg0>. It’s not just looking at individual moments; it’s building a timeline.
Lu: The authors are showing us that this simple regression across views can yield good results on standard benchmarks like Kinetics and UCF101 zero. They're proving that this time-aware approach is effective for representation learning.
Meng: And while the paper shows how to use different augmentations, they also note a limitation—they have to be careful about how they enforce those constraints when using just visual inputs zero. That means we can’t just throw in any random transformation without considering its effect on the learned representations.
Tom: So that's the scope of this paper: proposing this specific time-based view prediction method and showing it works with various inputs and modalities zero. It's a solid step forward for how we teach AI to see video.
Conclusion: Tom: So we’re wrapping up our look at BraVe today, and we're talking about how it redefines what self-supervised video learning means zero.
Jane: It basically takes this idea of a short clip and forces the model to predict the whole video context so it learns general stuff one.
Tom: And the authors, Recasens and their team, they’ve really shown how you can use time in a new way for these representations zero.
Lu: What’s cool is they aren't just looking at one view; they're actually making two views—a narrow one and a broad one—and forcing them to talk to each other zero.
Meng: From an engineering standpoint, that back-and-forth prediction structure is clever because it keeps the representations from collapsing into something useless zero.
Lalam: And for culture, this means we can train models that understand long-term video narratives instead of just catching quick actions one.
Tom: So what does this mean for a regular person listening? It means the AI will get much better at understanding the full scene, not just one frozen moment zero.
Jane: It's about teaching the computer to look further back and forward in time to get a complete picture of what’s happening one.
Tom: The authors show they can use audio or optical flow in that broad view, which is really interesting for making video understanding more flexible zero.
Lu: They tested it on some standard benchmarks like Kinetics and AudioSet, and the results they got were actually quite strong zero.
Meng: But they also flagged that when you’re just using visual inputs alone, you have to be careful about the augmentations you use to keep things stable zero.
Tom: So for anyone who works with these models, it means you have this new setup where the model is constantly checking its understanding against a much bigger context zero.
Jane: And that’s what we need to keep in mind as we look at how these kinds of video representations start appearing more often zero.
Adria Recasens, Pauline Luc, Jean-Baptiste Alayrac, Luyu Wang, Ross Hemsley, Florian Strub, Corentin Tallec, Mateusz Malinowski, Viorica Patrăuceanu
DeepMind
cs.CV, cs.LG
Submitted: 2021-03-30
Updated: 2021-10-19
Code: https://github.com/eepmind/brave
Importance score: 83/100
The gist: This paper introduces BraVe, a self-supervised learning framework for video that breaks temporal symmetry by training networks to predict a broad view from a narrow temporal window, enabling richer
Key concepts
- Temporal Symmetry Breaking
- Traditional methods often treat video frames independently or use short windows. BraVe breaks this symmetry by forcing the network to predict a wide view using only a small segment of time. This teaches the model to understand how events unfold over longer temporal contexts, moving beyond simple frame-by-frame analysis.
- Narrow-to-Broad Prediction Loss
- The core learning mechanism involves two losses: one predicting the broad view from the narrow temporal window, and a complementary loss regressing the narrow view from the broad context. This dual loss structure ensures that both views are learned effectively and prevents representation collapse.
- Multimodal Broad Views
- BraVe allows for diverse inputs for the 'broad view,' including different modalities like optical flow, randomly convolved RGB frames, or audio. This flexibility enables richer representations by incorporating information from various data types simultaneously into the context prediction task.
Terminology
Summary
This paper introduces BraVe, a self-supervised learning framework for video that breaks temporal symmetry by training networks to predict a broad view from a narrow temporal window, enabling richer representations across different modalities like RGB, optical flow, and audio. This approach matters because it moves beyond traditional image-based self-supervised methods by explicitly leveraging the time dimension in videos to learn generalizable context understanding.
The gist: BraVe learns a representation by predicting a broad view that spans the longer temporal context of the full video clip as illustrated in Figure 1, solving such a task requires extrapolating to the general context in which a given event occurs.<ref:2103.16559#pg1>
How it works
BraVe is designed around a core task where one view has access to a narrow temporal window of the video while the other view has broad access to the video content, and Our models learn to generalise from the narrow view to the general content of the video
<ref:2103.16559#pg1> The framework minimizes a training loss defined as:
L(x) = Ln→b(x) Narrow→Broad + Lb→n(x) Broad→Narrow <ref:2103.16559#pg1> This loss is composed of two terms: (i) a prediction loss from the narrow to the broad view, and (ii) a complementary loss to regress the narrow view from the broad view
<ref:2103.16559#pg1>.
BraVe: losses and architectures
To avoid collapse of learned representations, BraVe draws inspiration from prior work by introducing three stages of processing as previously done in BYOL [30]: backbone networks (fn for the narrow view and fb acting on the broad view, f1b, f2b for the broad views), projector network g and predictor network h
<ref:2103.16559#pg1>. The two regression losses are defined as:
Ln→b(x) = h(zn) / kh(zn) squared - sg[zb] / zb k squared <ref:2103.16559#pg1> and Lb→n(x) = h(zb) / kh(zb) squared - sg[zn] / zn k squared <ref:2103.16559#pg1>. The predictor network is crucial to avoid collapse, and we do not use exponential moving averages (EMA) on the weights of the network that process the view being regressed
<ref:2103.16559#pg1>.
Broad views from visual and audio modalities
BraVe can process broad views with different backbones, enabling the use of alternative augmentations or modalities into the broad view such as optical flow, randomly convolved RGB frames, audio or their combinations
<ref:2103.16559#pg1>. When using visual inputs alone, the framework employs stochastic transformations to enforce invariance or equivariance constraints on the learned representations
<ref:2103.16559#pg1>. For instance, we also employ a recently introduced augmentation procedure relying on random convolutions [90], by which we augment only the broad view
<ref:2103.16559#pg1>.
Evaluation and Contributions
BraVe achieves state-of-the-art results on standard video and audio classification benchmarks, including UCF101, HMDB51, Kinetics, ESC-50 and AudioSet
<ref:2103.16559#pg1>. The contributions include: (i) We propose a novel framework for representation learning, called BraVe, which generates views at different time scales and learns representations via simple regression across views,
(ii) We explore using different augmentations and modalities in the broad view such as audio, flow or randomly convolved RGB frames,
and (iii) We evaluate this framework in the video domain, both with and without audio as an auxiliary supervisory signal
<ref:2103.16559#pg1>.
REFERENCES
[1] Sami Abu-El-Haija, Nisarg Kothari, Joonseok Lee, Paul Natsev, George Toderici, Balakrishnan Varadarajan, and Sudheendra Vijayanarasimhan. Youtube-8m: A large-scale video classification benchmark. arXiv preprint arXiv:1609.08675, 2016.<ref:2103.16559#pg2>
[2] Jean-Baptiste Alayrac, Piotr Bojanowski, Nishant Agrawal, Ivan Laptev, Josef Sivic, and Simon Lacoste-Julien. Unsupervised learning from narrated instruction videos. In CVPR, 2016.<ref:2103.16559#pg3>
[3] Jean-Baptiste Alayrac, Adria Recasens, Rosalia Schneider, Relja Arandjelovic, Jason Ramapuram, Jeffrey De Fauw, Lucas Smaira, Sander Dieleman, and Andrew Zisserman. Selfsupervised multimodal versatile networks. In NeurIPS, 2020.<ref:2103.16559#pg4>
[4] Humam Alwassel, Dhruv Mahajan, Lorenzo Torresani, Bernard Ghanem, and Du Tran. Self-supervised learning by cross-modal audio-video clustering. In NeurIPS, 2020.<ref:2103.16559#pg5>
[5] Relja Arandjelovic and Andrew Zisserman. Look, listen and learn. In ICCV, 2017.<ref:2103.16559#pg6>
[6] Relja Arandjelovic and Andrew Zisserman. Objects that sound. In ECCV, 2018.<ref:2103.16559#pg7>
[7] Yuki Markus Asano, Christian Rupprecht, and Andrea Vedaldi. Self-labelling via simultaneous clustering and representation learning. In ICLR, 2020.<ref:2103.16559#pg8>
[8] Philip Bachman, R Devon Hjelm, and William Buchwalter. Learning representations by maximizing mutual information across views. In Neural Information Processing Systems, 2019.<ref:2103.16559#pg9>
[9] Philip Bachman, R Devon Hjelm, and William Buchwalter. Learning representations by maximizing mutual information across views. In NeurIPS, 2019.<ref:2103.16559#pg9>
[10] Miguel A. Bautista, Artsiom Sanakoyeu, Ekaterina Sutter, and Bjorn Ommer. Cliquecnn: Deep unsupervised exemplar learning. In NeurIPS, 2016.<ref:2103.16559#pg10>
[11] Gedas Bertasius, Heng Wang, and Lorenzo Torresani. Is space-time attention all you need for video understanding? arXiv preprint arXiv:2102.05095, 2021.
[12] Mathilde Caron, Piotr Bojanowski, Armand Joulin, and Matthijs Douze. Deep clustering for unsupervised learning of visual features. In ECCV, 2018.
[13] Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin. Unsupervised learning of visual features by contrasting cluster assignments. In NeurIPS, 2020.
[14] Joao Carreira, Eric Noland, Andras Banki-Horvath, Chloe Hillier, and Andrew Zisserman. A short note about kinetics600. arXiv preprint arXiv:1808.01340, 2018.
[15] Joao Carreira and Andrew Zisserman. Quo vadis, action recognition? A new model and the Kinetics dataset. In CVPR, 2017.
[16] Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. In ICML, 2020.
[17] R Devon et al. Representation learning with video deep infomax. arXiv preprint arXiv:2007.13278, 2020.
[18] Carl Doersch, Abhinav Gupta, and Alexei A Efros. Unsupervised visual representation learning by context prediction. In ICCV, 2015.
[19] Carl Doersch and Andrew Zisserman. Multi-task selfsupervised visual learning. In ICCV, 2017.<ref:2103.
Improvements for AI systems
-
The BraVe framework can learn representations by
predicting a broad view that spans a longer temporal context of the full video clip
from anarrow view corresponding to a video clip of a few seconds,
enabling generalization from short clips to long-term context, which is crucial for tasks requiring understanding before, during, and after an event. -
The system can utilize different modalities for the broad view such as
optical flow or audio or their combinations,
allowing the model to extract complementary information; for instance, using optical flow in the broad view can provide supervision thatemphasize motion in the learned representations extracted from the source.
-
The framework supports processing views with different backbones, which allows for applying
different set of preprocessing and augmentation functions to any of the views,
leading to improved performance by enablingalternative augmentations or modalities into the broad view such as optical flow, randomly convolved RGB frames, audio or their combinations.
-
The model can be trained using a regression-based approach inspired by BYOL where
the views are processed by dedicated backbones and regress each other,
avoiding the need forcumbersome creation of explicit negatives
and achieving state-of-the-art results under a fixed computational budget. -
The system can be extended to handle multiple broad views from different modalities, where the loss is aggregated as
L(x) = X K k=1 L k n→b(x) + L k b→n(x),
which allows for leveragingthe complementary information provided by RGB and flow together with the larger overall capacity of the broad model
to improve performance on benchmarks like HMDB51 and UCF101. -
The system can be optimized for faster training by avoiding Exponential Moving Averages (EMA) on weights, as noted in
unlike [22, 30], who required the moving average for improved performance,
thereby reducing computational overhead while maintaining state-of-the-art performance. -
The framework benefits from temporal alignment when using audio, as
syncing helps slightly as previously observed in visual-audio work [46],
and the best option found issyncing at the beginning,
which supports the hypothesis thatwhen both views are in sync, the broad network can simply focus its prediction only on the narrow view.
Sources
- YouTube-8M: A Large-Scale Video Classification Benchmark
- Is Space-Time Attention All You Need for Video Understanding?
- A Short Note about Kinetics-600
- Representation Learning with Video Deep InfoMax
- AST: Audio Spectrogram Transformer
- SMART Frame Selection for Action Recognition
- Learning deep representations by mutual information estimation and maximization
- Improved Baselines with Momentum Contrastive Learning
- The Kinetics Human Action Video Dataset
- Decoupled Weight Decay Regularization
- Audio-Visual Instance Discrimination with Cross-Modal Agreement
- Representation Learning with Contrastive Predictive Coding
- On Compositions of Transformations in Contrastive Self-Supervised Learning
- Spatiotemporal Contrastive Video Representation Learning
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Contrastive Multiview Coding
- Robust and Generalizable Visual Representation Learning via Random Convolutions
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models