Fetching the paper…
Reading the bibliography…
We consider the task of identifying human actions visible in online videos.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014 · 1958
Earlier work this paper cites.
The equivalence of weighted kappa and the intraclass correlation coefficient as measures of reliability
Joseph L Fleiss and Jacob Cohen. 1973 · 1973
Earlier work this paper cites.
Verbs semantics and lexical selection
Zhibiao Wu and Martha Palmer. 1994 · 1994
Earlier work this paper cites.
Support-vector networks
Corinna Cortes and Vladimir Vapnik. 1995 · 1995
Earlier work this paper cites.
Learning and recognizing human dynamics in video sequences
Christoph Bregler. 1997 · 1997
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Feature-rich part-of-speech tagging with a cyclic dependency network
Kristina Toutanova, Dan Klein, Christopher D Manning, and Yoram Singer. 2003 · 2003
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009 · 2009
Earlier work this paper cites.
3d convolutional neural networks for human action recognition
Shuiwang Ji, Wei Xu, Ming Yang, and Kai Yu. 2012 · 2012
Earlier work this paper cites.
A database for fine grained activity detection of cooking activities
Marcus Rohrbach, Sikandar Amin, Mykhaylo Andriluka, and Bernt Schiele. 2012 · 2012
Earlier work this paper cites.
Action bank: A high-level representation of activity in video
Sreemanananth Sadanand and Jason J Corso. 2012 · 2012
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah. 2012 · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton. 2012 · 2012
Earlier work this paper cites.
Two-person interaction detection using body-pose features and multiple instance learning
Kiwon Yun, Jean Honorio, Debaleena Chattopadhyay, Tamara L Berg, and Dimitris Samaras. 2012 · 2012
Earlier work this paper cites.
A thousand frames in just a few words: Lingual description of videos through latent topics and sparse object stitching
Pradipto Das, Chenliang Xu, Richard F Doell, and Jason J Corso. 2013 · 2013
Earlier work this paper cites.
Realistic human action recognition with multimodal feature selection and fusion
Qiuxia Wu, Zhiyong Wang, Feiqi Deng, Zheru Chi, and David Dagan Feng. 2013 · 2013
Earlier work this paper cites.
Concreteness ratings for 40 thousand generally known english word lemmas
Marc Brysbaert, Amy Beth Warriner, and Victor Kuperman. 2014 · 2014
Earlier work this paper cites.
A fast and accurate dependency parser using neural networks
Danqi Chen and Christopher Manning. 2014 · 2014
Earlier work this paper cites.
Creating summaries from user videos
Michael Gygli, Helmut Grabner, Hayko Riemenschneider, and Luc Van Gool. 2014 · 2014
Cited alongside, same era.
Large-scale video classification with convolutional neural networks
Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas Leung, Rahul Sukthankar, and Li Fei-Fei. 2014 · 2014
Cited alongside, same era.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014 · 2014
Cited alongside, same era.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014 · 2014
Cited alongside, same era.
Coherent multi-sentence video description with variable level of detail
Anna Rohrbach, Marcus Rohrbach, Wei Qiu, Annemarie Friedrich, Manfred Pinkal, and Bernt Schiele. 2014 · 2014
Cited alongside, same era.
Two-stream convolutional networks for action recognition in videos
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. 2016 · 2016
Later among the works it cites.
Temporal segment networks: Towards good practices for deep action recognition
Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool. 2016 · 2016
Later among the works it cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman. 2017 · 2017
Later among the works it cites.
Hockey action recognition via integrated stacked hourglass network
Mehrnaz Fani, Helmut Neher, David A Clausi, Alexander Wong, and John Zelek. 2017 · 2017
Later among the works it cites.
Going deeper into action recognition: A survey
Samitha Herath, Mehrtash Harandi, and Fatih Porikli. 2017 · 2017
Later among the works it cites.
Tgif-qa: Toward spatio-temporal reasoning in visual question answering
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Karen Simonyan and Andrew Zisserman. 2014 · 2014
Cited alongside, same era.
Human action recognition by representing 3d skeletons as points in a lie group
Raviteja Vemulapalli, Felipe Arrate, and Rama Chellappa. 2014 · 2014
Cited alongside, same era.
Long-term recurrent convolutional networks for visual recognition and description
Jeffrey Donahue, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Darrell. 2015 · 2015
Cited alongside, same era.
Hierarchical recurrent neural network for skeleton based action recognition
Yong Du, Wei Wang, and Liang Wang. 2015 · 2015
Cited alongside, same era.
Activitynet: A large-scale video benchmark for human activity understanding
Bernard Ghanem Fabian Caba Heilbron, Victor Escorcia and Juan Carlos Niebles. 2015 · 2015
Cited alongside, same era.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Li Fei-Fei. 2015 · 2015
Cited alongside, same era.
Vision based hand gesture recognition for human computer interaction: a survey
Siddharth S Rautaray and Anupam Agrawal. 2015 · 2015
Cited alongside, same era.
Yunseok Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gunhee Kim. 2017 · 2017
Later among the works it cites.
The kinetics human action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al. 2017 · 2017
Later among the works it cites.
Pku-mmd: A large scale benchmark for continuous multi-modal human action understanding
Chunhui Liu, Yueyu Hu, Yanghao Li, Sijie Song, and Jiaying Liu. 2017 · 2017
Later among the works it cites.
Youtube-boundingboxes: A large high-precision human-annotated data set for object detection in video
Esteban Real, Jonathon Shlens, Stefano Mazzocchi, Xin Pan, and Vincent Vanhoucke. 2017 · 2017
Later among the works it cites.
Yolo9000: better, faster, stronger
Joseph Redmon and Ali Farhadi. 2017 · 2017
Later among the works it cites.
Visual question answering: A survey of methods and datasets
Qi Wu, Damien Teney, Peng Wang, Chunhua Shen, Anthony Dick, and Anton van den Hengel. 2017 · 2017
Later among the works it cites.
View adaptive recurrent neural networks for high performance human action recognition from skeleton data
Pengfei Zhang, Cuiling Lan, Junliang Xing, Wenjun Zeng, Jianru Xue, and Nanning Zheng. 2017 · 2017
Later among the works it cites.
Scaling egocentric vision: The epic-kitchens dataset
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al. 2018 · 2018
Later among the works it cites.
From lifestyle vlogs to everyday interactions
David F Fouhey, Wei-cheng Kuo, Alexei A Efros, and Jitendra Malik. 2018 · 2018
Later among the works it cites.
Ava: A video dataset of spatio-temporally localized atomic visual actions
Chunhui Gu, Chen Sun, David A Ross, Carl Vondrick, Caroline Pantofaru, Yeqing Li, Sudheendra Vijayanarasimhan, George Toderici, Susanna Ricco, Rahul Sukthankar, et al. 2018 · 2018
Later among the works it cites.
Tvqa: Localized, compositional video question answering
Jie Lei, Licheng Yu, Mohit Bansal, and Tamara Berg. 2018 · 2018
Later among the works it cites.
Deep contextualized word representations
Matthew E Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Later among the works it cites.
Moments in time dataset: one million videos for event understanding
Mathew Monfort, Alex Andonian, Bolei Zhou, Kandan Ramakrishnan, Sarah Adel Bargal, Yan Yan, Lisa Brown, Quanfu Fan, Dan Gutfreund, Carl Vondrick, et al. 2019 · 2019
Closest in time.