Attention clusters: Purely attention based local feature integration for video classification
Xiang Long, Chuang Gan, Gerard De Melo, Jiajun Wu, Xiao Liu, and Shilei Wen · 2018
Later among the works it cites.
Seeing voices and hearing faces: Cross-modal biometric matching
Arsha Nagrani, Samuel Albanie, and Andrew Zisserman · 2018
Later among the works it cites.
Audio-visual scene analysis with self-supervised multisensory features
Andrew Owens and Alexei A Efros · 2018
Later among the works it cites.
Learning to localize sound source in visual scenes
Arda Senocak, Tae-Hyun Oh, Junsik Kim, Ming-Hsuan Yang, and In So Kweon · 2018
Later among the works it cites.
Actor-centric relation network
Chen Sun, Abhinav Shrivastava, Carl Vondrick, Kevin Murphy, Rahul Sukthankar, and Cordelia Schmid · 2018
Later among the works it cites.
A closer look at spatiotemporal convolutions for action recognition
Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri · 2018
Later among the works it cites.
Non-local neural networks
Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He · 2018
Later among the works it cites.
Videos as space-time region graphs
Xiaolong Wang and Abhinav Gupta · 2018
Later among the works it cites.
Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification
Saining Xie, Chen Sun, Jonathan Huang, Zhuowen Tu, and Kevin Murphy · 2018
Later among the works it cites.
Pixel objectness: Learning to segment generic objects automatically in images and videos
B. Xiong, S. Jain, and K. Grauman · 2018
Later among the works it cites.
The sound of pixels
Hang Zhao, Chuang Gan, Andrew Rouditchenko, Carl Vondrick, Josh McDermott, and Antonio Torralba · 2018
Later among the works it cites.
ECO: Efficient convolutional network for online video understanding
Mohammadreza Zolfaghari, Kamaljeet Singh, and Thomas Brox · 2018
Later among the works it cites.
SlowFast Networks for Video Recognition
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He · 2019
Later among the works it cites.
What would you expect? anticipating egocentric actions with rolling-unrolling LSTMs and modality attention
Antonino Furnari and Giovanni Maria Farinella · 2019
Later among the works it cites.
Video Action Transformer Network
Rohit Girdhar, João Carreira, Carl Doersch, and Andrew Zisserman · 2019
Later among the works it cites.
Video representation learning by dense predictive coding
Tengda Han, Weidi Xie, and Andrew Zisserman · 2019
Later among the works it cites.
Momentum contrast for unsupervised visual representation learning
Original
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick · 2019
Later among the works it cites.
Deep multimodal clustering for unsupervised audiovisual learning
Di Hu, Feiping Nie, and Xuelong Li · 2019
Later among the works it cites.
Timeception for complex action recognition
Noureldien Hussein, Efstratios Gavves, and Arnold WM Smeulders · 2019
Later among the works it cites.
EPIC-Fusion: Audio-Visual Temporal Binding for Egocentric Action Recognition
Evangelos Kazakos, Arsha Nagrani, Andrew Zisserman, and Dima Damen · 2019
Later among the works it cites.
Self-supervised video representation learning with space-time cubic puzzles
Dahun Kim, Donghyeon Cho, and In So Kweon · 2019
Later among the works it cites.
Revisiting self-supervised visual representation learning
Original
Alexander Kolesnikov, Xiaohua Zhai, and Lucas Beyer · 2019
Later among the works it cites.
SCSampler: Sampling salient clips from video for efficient action recognition
Bruno Korbar, Du Tran, and Lorenzo Torresani · 2019
Later among the works it cites.
Multisensory Interactions in the Real World
Salvador Soto-Faraco, Daria Kvasova, Emmanuel Biau, Nara Ikumi, Manuela Ruzzoli, Luis Morís-Fernández, and Mireia Torralba · 2019
Later among the works it cites.
FBK-HUPBA Submission to the EPIC-Kitchens 2019 Action Recognition Challenge
Original
Swathikiran Sudhakaran, Sergio Escalera, and Oswald Lanz · 2019
Later among the works it cites.
Hierarchical feature aggregation networks for video action recognition
Original
Swathikiran Sudhakaran, Sergio Escalera, and Oswald Lanz · 2019
Later among the works it cites.
Contrastive bidirectional transformer for temporal representation learning
Original
Chen Sun, Fabien Baradel, Kevin Murphy, and Cordelia Schmid · 2019
Later among the works it cites.
Video classification with channel-separated convolutional networks
Original
Du Tran, Heng Wang, Lorenzo Torresani, and Matt Feiszli · 2019
Later among the works it cites.
Long-term feature banks for detailed video understanding
Chao-Yuan Wu, Christoph Feichtenhofer, Haoqi Fan, Kaiming He, Philipp Krahenbuhl, and Ross Girshick · 2019
Later among the works it cites.