Fetching the paper…
Reading the bibliography…
In this paper, we introduce Coarse-Fine Networks, a two-stream architecture which benefits from different abstractions of temporal resolution to learn better video representations for long-term motion.
The graph neural network model
Franco Scarselli, Marco Gori, Ah Chung Tsoi, Markus Hagenbuchner, and Gabriele Monfardini · 2008
Earlier work this paper cites.
Thumos challenge: Action recognition with a large number of classes, 2014
Yu-Gang Jiang, Jingen Liu, A Roshan Zamir, George Toderici, Ivan Laptev, Mubarak Shah, and Rahul Sukthankar · 2014
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas Leung, Rahul Sukthankar, and Li Fei-Fei · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
C3d: generic features for video analysis
Du Tran, Lubomir D Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri · 2014
Earlier work this paper cites.
Spatial transformer networks
Max Jaderberg, Karen Simonyan, Andrew Zisserman, et al · 2015
Earlier work this paper cites.
Beyond short snippets: Deep networks for video classification
Joe Yue-Hei Ng, Matthew Hausknecht, Sudheendra Vijayanarasimhan, Oriol Vinyals, Rajat Monga, and George Toderici · 2015
Earlier work this paper cites.
Daps: Deep action proposals for action understanding
Victor Escorcia, Fabian Caba Heilbron, Juan Carlos Niebles, and Bernard Ghanem · 2016
Earlier work this paper cites.
Convolutional two-stream network fusion for video action recognition
Christoph Feichtenhofer, Axel Pinz, and Andrew Zisserman · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Temporal action localization in untrimmed videos via multi-stage cnns
Zheng Shou, Dongang Wang, and Shih-Fu Chang · 2016
Earlier work this paper cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding
Gunnar A Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta · 2016
Earlier work this paper cites.
End-to-end learning of action detection from frame glimpses in videos
Serena Yeung, Olga Russakovsky, Greg Mori, and Li Fei-Fei · 2016
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
Deformable convolutional networks
Jifeng Dai, Haozhi Qi, Yuwen Xiong, Yi Li, Guodong Zhang, Han Hu, and Yichen Wei · 2017
Earlier work this paper cites.
Learning spatio-temporal features with 3d residual networks for action recognition
Kensho Hara, Hirokatsu Kataoka, and Yutaka Satoh · 2017
Cited alongside, same era.
The kinetics human action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al · 2017
Cited alongside, same era.
Temporal convolutional networks for action segmentation and detection
Colin Lea, Michael D Flynn, Rene Vidal, Austin Reiter, and Gregory D Hager · 2017
Cited alongside, same era.
Cdc: Convolutional-de-convolutional networks for precise temporal action localization in untrimmed videos
Zheng Shou, Jonathan Chan, Alireza Zareian, Kazuyuki Miyazawa, and Shih-Fu Chang · 2017
Cited alongside, same era.
Convnet architecture search for spatiotemporal feature learning
Du Tran, Jamie Ray, Zheng Shou, Shih-Fu Chang, and Manohar Paluri · 2017
Eco: Efficient convolutional network for online video understanding
Mohammadreza Zolfaghari, Kamaljeet Singh, and Thomas Brox · 2018
Later among the works it cites.
Slowfast networks for video recognition
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He · 2019
Later among the works it cites.
Scsampler: Sampling salient clips from video for efficient action recognition
Bruno Korbar, Du Tran, and Lorenzo Torresani · 2019
Later among the works it cites.
Temporal gaussian mixture layer for videos
AJ Piergiovanni and Michael S. Ryoo · 2019
Later among the works it cites.
Liteeval: A coarse-to-fine framework for resource efficient video recognition
Zuxuan Wu, Caiming Xiong, Yu-Gang Jiang, and Larry S Davis · 2019
Later among the works it cites.
Adaframe: Adaptive frame selection for fast video recognition
Zuxuan Wu, Caiming Xiong, Chih-Yao Ma, Richard Socher, and Larry S Davis · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Long-term temporal convolutions for action recognition
Gül Varol, Ivan Laptev, and Cordelia Schmid · 2017
Cited alongside, same era.
R-c3d: Region convolutional 3d network for temporal activity detection
Huijuan Xu, Abir Das, and Kate Saenko · 2017
Cited alongside, same era.
Temporal action detection with structured segment networks
Yue Zhao, Yuanjun Xiong, Limin Wang, Zhirong Wu, Xiaoou Tang, and Dahua Lin · 2017
Cited alongside, same era.
Object level visual reasoning in videos
Fabien Baradel, Natalia Neverova, Christian Wolf, Julien Mille, and Greg Mori · 2018
Cited alongside, same era.
Attend and interact: Higher-order object interactions for video understanding
Chih-Yao Ma, Asim Kadav, Iain Melvin, Zsolt Kira, Ghassan AlRegib, and Hans Peter Graf · 2018
Cited alongside, same era.
Learning latent super-events to detect multiple activities in videos
AJ Piergiovanni and Michael S Ryoo · 2018
Cited alongside, same era.
Learning to zoom: a saliency-based sampling layer for neural networks
Adria Recasens, Petr Kellnhofer, Simon Stent, Wojciech Matusik, and Antonio Torralba · 2018
Cited alongside, same era.
Later among the works it cites.
Grounded video description
Luowei Zhou, Yannis Kalantidis, Xinlei Chen, Jason J Corso, and Marcus Rohrbach · 2019
Later among the works it cites.
X3D: Expanding architectures for efficient video recognition
Christoph Feichtenhofer · 2020
Later among the works it cites.
Beyond fixed grid: Learning geometric image representation with a deformable grid
Jun Gao, Zian Wang, Jinchen Xuan, and Sanja Fidler · 2020
Later among the works it cites.
Stacked spatio-temporal graph convolutional networks for action segmentation
Pallabi Ghosh, Yi Yao, Larry Davis, and Ajay Divakaran · 2020
Later among the works it cites.
Action genome: Actions as compositions of spatio-temporal scene graphs
Jingwei Ji, Ranjay Krishna, Li Fei-Fei, and Juan Carlos Niebles · 2020
Later among the works it cites.
Representation learning on visual-symbolic graphs for video understanding
Effrosyni Mavroudi, Benjamín Béjar Haro, and René Vidal · 2020
Later among the works it cites.
Ar-net: Adaptive frame resolution for efficient action recognition
Yue Meng, Chung-Ching Lin, Rameswar Panda, Prasanna Sattigeri, Leonid Karlinsky, Aude Oliva, Kate Saenko, and Rogerio Feris · 2020
Later among the works it cites.
AssembleNet: Searching for multi-stream neural connectivity in video architectures
Michael Ryoo, AJ Piergiovanni, Mingxing Tan, and Anelia Angelova · 2020
Later among the works it cites.