Fetching the paper…
Reading the bibliography…
While many action recognition datasets consist of collections of brief, trimmed videos each containing a relevant action, videos in the real-world (e.g., on YouTube) exhibit very different properties: they are often several minutes long, where brief relevant clips are often interleaved with segments of extended duration containing little change.
Temporal localization of actions with actoms
Adrien Gaidon, Zaïd Harchaoui, and Cordelia Schmid · 2013
Earlier work this paper cites.
Diverse sequential subset selection for supervised video summarization
Boqing Gong, Wei-Lun Chao, Kristen Grauman, and Fei Sha · 2014
Earlier work this paper cites.
Action localization with tubelets from motion
Mihir Jain, Jan C. van Gemert, Hervé Jégou, Patrick Bouthemy, and Cees G. M. Snoek · 2014
Earlier work this paper cites.
Efficient feature extraction, encoding, and classification for action recognition
Vadim Kantorov and Ivan Laptev · 2014
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas Leung, Rahul Sukthankar, and Fei-Fei Li · 2014
Earlier work this paper cites.
Parsing videos of actions with segmental grammars
Hamed Pirsiavash and Deva Ramanan · 2014
Earlier work this paper cites.
Category-specific video summarization
Danila Potapov, Matthijs Douze, Zaïd Harchaoui, and Cordelia Schmid · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Video summarization by learning submodular mixtures of objectives
Michael Gygli, Helmut Grabner, and Luc J. Van Gool · 2015
Earlier work this paper cites.
Beyond short snippets: Deep networks for video classification
Joe Yue-Hei Ng, Matthew J. Hausknecht, Sudheendra Vijayanarasimhan, Oriol Vinyals, Rajat Monga, and George Toderici · 2015
Earlier work this paper cites.
Beyond short snippets: Deep networks for video classification
Joe Yue-Hei Ng, Matthew J. Hausknecht, Sudheendra Vijayanarasimhan, Oriol Vinyals, Rajat Monga, and George Toderici · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri · 2015
Earlier work this paper cites.
Soundnet: Learning sound representations from unlabeled video
Yusuf Aytar, Carl Vondrick, and Antonio Torralba · 2016
Earlier work this paper cites.
Out of time: Automated lip sync in the wild
Joon Son Chung and Andrew Zisserman · 2016
Earlier work this paper cites.
Daps: Deep action proposals for action understanding
Victor Escorcia, Fabian Caba Heilbron, Juan Carlos Niebles, and Bernard Ghanem · 2016
Earlier work this paper cites.
Spatiotemporal residual networks for video action recognition
Christoph Feichtenhofer, Axel Pinz, and Richard Wildes · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Fast temporal activity proposals for efficient detection of human actions in untrimmed videos
Fabian Caba Heilbron, Juan Carlos Niebles, and Bernard Ghanem · 2016
Earlier work this paper cites.
Temporal action localization in untrimmed videos via multi-stage cnns
Zheng Shou, Dongang Wang, and Shih-Fu Chang · 2016
Earlier work this paper cites.
Mofap: A multi-level representation for action recognition
Limin Wang, Yu Qiao, and Xiaoou Tang · 2016
Cited alongside, same era.
Temporal segment networks: Towards good practices for deep action recognition
Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool · 2016
Cited alongside, same era.
Actions ~ transformations
Xiaolong Wang, Ali Farhadi, and Abhinav Gupta · 2016
Cited alongside, same era.
End-to-end learning of action detection from frame glimpses in videos
Serena Yeung, Olga Russakovsky, Greg Mori, and Li Fei-Fei · 2016
Cited alongside, same era.
Real-time action recognition with enhanced motion vector cnns
Bowen Zhang, Limin Wang, Zhe Wang, Yu Qiao, and Hanli Wang · 2016
Cited alongside, same era.
Summary transfer: Exemplar-based subset selection for video summarization
Ke Zhang, Wei-Lun Chao, Fei Sha, and Kristen Grauman · 2016
Objects that sound
Relja Arandjelovic and Andrew Zisserman · 2018
Later among the works it cites.
Video model zoo
Facebook · 2018
Later among the works it cites.
Watching a small portion could be as good as watching all: Towards efficient video classification
Hehe Fan, Zhongwen Xu, Linchao Zhu, Chenggang Yan, Jianjun Ge, and Yi Yang · 2018
Later among the works it cites.
Slowfast networks for video recognition
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He · 2018
Later among the works it cites.
Learning to separate object sounds by watching unlabeled video
Ruohan Gao, Rogério Schmidt Feris, and Kristen Grauman · 2018
Later among the works it cites.
Cooperative learning of audio and video models from self-supervised synchronization
Bruno Korbar, Du Tran, and Lorenzo Torresani · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Video summarization with long short-term memory
Ke Zhang, Wei-Lun Chao, Fei Sha, and Kristen Grauman · 2016
Cited alongside, same era.
Look, listen and learn
Relja Arandjelovic and Andrew Zisserman · 2017
Cited alongside, same era.
SST: single-stream temporal action proposals
Shyamal Buch, Victor Escorcia, Chuanqi Shen, Bernard Ghanem, and Juan Carlos Niebles · 2017
Cited alongside, same era.
Quo vadis, action recognition? A new model and the kinetics dataset
João Carreira and Andrew Zisserman · 2017
Cited alongside, same era.
TURN TAP: temporal unit regression network for temporal action proposals
Jiyang Gao, Zhenheng Yang, Chen Sun, Kan Chen, and Ram Nevatia · 2017
Cited alongside, same era.
Audio set: An ontology and human-labeled dataset for audio events
Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R. Channing Moore, Manoj Plakal, and Marvin Ritter · 2017
Cited alongside, same era.
Later among the works it cites.
BSN: boundary sensitive network for temporal action proposal generation
Tianwei Lin, Xu Zhao, Haisheng Su, Chongjing Wang, and Ming Yang · 2018
Later among the works it cites.
The excitement of sports: Automatic highlights using audio/visual cues
Michele Merler, Dhiraj Joshi, Khoi-Nguyen C. Mac, Quoc-Bao Nguyen, Stephen Hammer, John Kent, Jinjun Xiong, Minh N. Do, John R. Smith, and Rogerio S. Feris · 2018
Later among the works it cites.
Automatic curation of sports highlights using multimodal excitement features
M. Merler, K. C. Mac, D. Joshi, Q. Nguyen, S. Hammer, J. Kent, J. Xiong, M. N. Do, J. R. Smith, and R. Feris · 2018
Later among the works it cites.
Audio-visual scene analysis with self-supervised multisensory features
Andrew Owens and Alexei A. Efros · 2018
Later among the works it cites.
A closer look at spatiotemporal convolutions for action recognition
Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri · 2018
Later among the works it cites.
Long-term temporal convolutions for action recognition
Gül Varol, Ivan Laptev, and Cordelia Schmid · 2018
Later among the works it cites.
Learning discriminative video representations using adversarial perturbations
Jue Wang and Anoop Cherian · 2018
Later among the works it cites.
Non-local neural networks
Xiaolong Wang, Ross B. Girshick, Abhinav Gupta, and Kaiming He · 2018
Later among the works it cites.
Long-term feature banks for detailed video understanding
Chao-Yuan Wu, Christoph Feichtenhofer, Haoqi Fan, Kaiming He, Philipp Krähenbühl, and Ross B. Girshick · 2018
Later among the works it cites.
Compressed video action recognition
Chao-Yuan Wu, Manzil Zaheer, Hexiang Hu, R. Manmatha, Alexander J. Smola, and Philipp Krähenbühl · 2018
Later among the works it cites.
Shufflenet: An extremely efficient convolutional neural network for mobile devices
Xiangyu Zhang, Xinyu Zhou, Mengxiao Lin, and Jian Sun · 2018
Later among the works it cites.
The sound of pixels
Hang Zhao, Chuang Gan, Andrew Rouditchenko, Carl Vondrick, Josh H. McDermott, and Antonio Torralba · 2018
Later among the works it cites.
Temporal relational reasoning in videos
Bolei Zhou, Alex Andonian, Aude Oliva, and Antonio Torralba · 2018
Later among the works it cites.
Classification with channel-separated convolutional networks
Du Tran, Heng Wang, Lorenzo Torresani, and Matt Feiszli · 2019
Closest in time.
Adaframe: Adaptive frame selection for fast video recognition
Zuxuan Wu, Caiming Xiong, Chih-Yao Ma, Richard Socher, and Larry S. Davis · 2019
Closest in time.