Fetching the paper…
Reading the bibliography…
We present the Moments in Time Dataset, a large-scale human-annotated collection of one million short videos corresponding to dynamic events unfolding within three seconds.
Working memory
Baddeley A · 1992
Earlier work this paper cites.
The berkeley framenet project
Collin F. Baker, Charles J. Fillmore, and John B. Lowe · 1998
Earlier work this paper cites.
Space-time interest points
Ivan Laptev and Tony Lindeberg · 2003
Earlier work this paper cites.
Time constraints and resource sharing in adults’ working memory spans
P. Barrouillet, S. Bernardin, and V. Camos · 2004
Earlier work this paper cites.
Recognizing human actions: a local svm approach
Christian Schuldt, Ivan Laptev, and Barbara Caputo · 2004
Earlier work this paper cites.
Actions as space-time shapes
Moshe Blank, Lena Gorelick, Eli Shechtman, Michal Irani, and Ronen Basri · 2005
Earlier work this paper cites.
Verbnet: A Broad-coverage, Comprehensive Verb Lexicon
Karin Kipper Schuler · 2005
Earlier work this paper cites.
Ontonotes: The 90 % \% solution
Eduard Hovy, Mitchell Marcus, Martha Palmer, Lance Ramshaw, and Ralph Weischedel · 2006
Earlier work this paper cites.
A duality based approach for realtime tv-l 1 optical flow
Christopher Zach, Thomas Pock, and Horst Bischof · 2007
Earlier work this paper cites.
A spatio-temporal descriptor based on 3d-gradients
Alexander Klaser, Marcin Marszałek, and Cordelia Schmid · 2008
Earlier work this paper cites.
Propbank
Paul Kingsbury and Martha Palmer · 2009
Earlier work this paper cites.
Action recognition by dense trajectories
Heng Wang, Alexander Kläser, Cordelia Schmid, and Cheng-Lin Liu · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton · 2012
Earlier work this paper cites.
Action bank: A high-level representation of activity in video
Sreemanananth Sadanand and Jason J Corso · 2012
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah · 2012
Earlier work this paper cites.
Hmdb51: A large video database for human motion recognition
Hilde Kuehne, Hueihan Jhuang, Rainer Stiefelhagen, and Thomas Serre · 2013
Earlier work this paper cites.
Thumos challenge: Action recognition with a large number of classes, 2014
YG Jiang, J Liu, A Roshan Zamir, G Toderici, I Laptev, M Shah, and R Sukthankar · 2014
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas Leung, Rahul Sukthankar, and Li Fei-Fei · 2014
Cited alongside, same era.
Parsing videos of actions with segmental grammars
Hamed Pirsiavash and Deva Ramanan · 2014
Cited alongside, same era.
A dataset and taxonomy for urban sound research
Justin Salamon, Christopher Jacoby, and Juan Pablo Bello · 2014
Cited alongside, same era.
Two-stream convolutional networks for action recognition in videos
Karen Simonyan and Andrew Zisserman · 2014
Cited alongside, same era.
Learning deep features for scene recognition using places database
Bolei Zhou, Agata Lapedriza, Jianxiong Xiao, Antonio Torralba, and Aude Oliva · 2014
Cited alongside, same era.
Activitynet: A large-scale video benchmark for human activity understanding
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles · 2015
Identity mappings in deep residual networks
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Later among the works it cites.
Cnn architectures for large-scale audio classification
Shawn Hershey, Sourish Chaudhuri, Daniel PW Ellis, Jort F Gemmeke, Aren Jansen, R Channing Moore, Manoj Plakal, Devin Platt, Rif A Saurous, Bryan Seybold, et al · 2016
Later among the works it cites.
The open world of micro-videos
Phuc Xuan Nguyen, Gregory Rogez, Charless Fowlkes, and Deva Ramanan · 2016
Later among the works it cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding
Gunnar A. Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta · 2016
Later among the works it cites.
Temporal segment networks: Towards good practices for deep action recognition
Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Long-term recurrent convolutional networks for visual recognition and description
Jeffrey Donahue, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Darrell · 2015
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
Open source computer vision library
Itseez · 2015
Cited alongside, same era.
Environmental sound classification with convolutional neural networks
Karol J Piczak · 2015
Cited alongside, same era.
Detection and classification of acoustic scenes and events
Dan Stowell, Dimitrios Giannoulis, Emmanouil Benetos, Mathieu Lagrange, and Mark D Plumbley · 2015
Cited alongside, same era.
Later among the works it cites.
Learning deep features for discriminative localization
Bolei Zhou, Aditya Khosla, Agata Lapedriza, Aude Oliva, and Antonio Torralba · 2016
Later among the works it cites.
Relja Arandjelović and Andrew Zisserman · 2017
Later among the works it cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Later among the works it cites.
From lifestyle vlogs to everyday interactions
David F Fouhey, Wei-cheng Kuo, Alexei A Efros, and Jitendra Malik · 2017
Later among the works it cites.
Audio set: An ontology and human-labeled dartaset for audio events
Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter · 2017
Later among the works it cites.
The" something something" video database for learning and evaluating visual common sense
Raghav Goyal, Samira Kahou, Vincent Michalski, Joanna Materzyńska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, et al · 2017
Later among the works it cites.
Ava: A video dataset of spatio-temporally localized atomic visual actions
Chunhui Gu, Chen Sun, Sudheendra Vijayanarasimhan, Caroline Pantofaru, David A Ross, George Toderici, Yeqing Li, Susanna Ricco, Rahul Sukthankar, Cordelia Schmid, et al · 2017
Later among the works it cites.
The kinetics human action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al · 2017
Later among the works it cites.
Places: A 10 million image database for scene recognition
Bolei Zhou, Agata Lapedriza, Aditya Khosla, Aude Oliva, and Antonio Torralba · 2017
Later among the works it cites.
Pulling actions out of context: Explicit separation for effective combination
Yang Wang and Minh Hoai · 2018
Closest in time.
Temporal relational reasoning in videos
Bolei Zhou, Alex Andonian, , Aude Oliva, and Antonio Torralba · 2018
Closest in time.