Fetching the paper…
Reading the bibliography…
Neural networks trained on datasets such as ImageNet have led to major advances in visual object classification.
Metaphors we live by
G. Lakoff and M. Johnson · 1981
Earlier work this paper cites.
Multiple view geometry in computer vision
R. Hartley and A. Zisserman · 2003
Earlier work this paper cites.
Recognizing human actions: a local svm approach
C. Schuldt, I. Laptev, and B. Caputo · 2004
Earlier work this paper cites.
Action mach a spatio-temporal maximum average correlation height filter for action recognition
M. D. Rodriguez, J. Ahmed, and M. Shah · 2008
Earlier work this paper cites.
Curriculum learning
Y. Bengio, J. Louradour, R. Collobert, and J. Weston · 2009
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Actions in context
M. Marszalek, I. Laptev, and C. Schmid · 2009
Earlier work this paper cites.
The winograd schema challenge
H. J. Levesque, E. Davis, and L. Morgenstern · 2011
Earlier work this paper cites.
Unbiased look at dataset bias
A. Torralba and A. A. Efros · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
A database for fine grained activity detection of cooking activities
M. Rohrbach, S. Amin, M. Andriluka, and B. Schiele · 2012
Earlier work this paper cites.
Decaf: A deep convolutional activation feature for generic visual recognition
J. Donahue, Y. Jia, O. Vinyals, J. Hoffman, N. Zhang, E. Tzeng, and T. Darrell · 2013
Earlier work this paper cites.
Vision meets robotics: The kitti dataset
A. Geiger, P. Lenz, C. Stiller, and R. Urtasun · 2013
Earlier work this paper cites.
Surfaces and Essences
D. Hofstadter, D. R. Hofstadter, and E. Sander · 2013
Earlier work this paper cites.
Learning to relate images
R. Memisevic · 2013
Cited alongside, same era.
Grounding action descriptions in videos
M. Regneri, M. Rohrbach, D. Wetzel, S. Thater, B. Schiele, and M. Pinkal · 2013
Cited alongside, same era.
Large-scale video classification with convolutional neural networks
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei · 2014
Cited alongside, same era.
Modeling deep temporal dependencies with recurrent grammar cells””
V. Michalski, R. Memisevic, and K. Konda · 2014
Cited alongside, same era.
Video (language) modeling: a baseline for generative models of natural videos
M. Ranzato, A. Szlam, J. Bruna, M. Mathieu, R. Collobert, and S. Chopra · 2014
Cited alongside, same era.
Cnn features off-the-shelf: an astounding baseline for recognition
A. Sharif Razavian, H. Azizpour, J. Sullivan, and S. Carlsson · 2014
Learning spatiotemporal features with 3d convolutional networks
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri · 2015
Later among the works it cites.
Anticipating the future by watching unlabeled video
C. Vondrick, H. Pirsiavash, and A. Torralba · 2015
Later among the works it cites.
Learning to poke by poking: Experiential learning of intuitive physics
P. Agrawal, A. Nair, P. Abbeel, J. Malik, and S. Levine · 2016
Later among the works it cites.
Learning physical intuition of block towers by example
A. Lerer, S. Gross, and R. Fergus · 2016
Later among the works it cites.
Synthesizing the preferred inputs for neurons in neural networks via deep generator networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Activitynet: A large-scale video benchmark for human activity understanding
B. G. Fabian Caba Heilbron, Victor Escorcia and J. C. Niebles · 2015
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Cited alongside, same era.
Y.-G. Jiang, Z. Wu, J. Wang, X. Xue, and S.-F. Chang · 2015
Cited alongside, same era.
A dataset for movie description
A. Rohrbach, M. Rohrbach, N. Tandon, and B. Schiele · 2015
Cited alongside, same era.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Cited alongside, same era.
A. Nguyen, A. Dosovitskiy, J. Yosinski, T. Brox, and J. Clune · 2016
Later among the works it cites.
The curious robot: Learning visual representations via physical interactions
L. Pinto, D. Gandhi, Y. Han, Y.-L. Park, and A. Gupta · 2016
Later among the works it cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding
G. A. Sigurdsson, G. Varol, X. Wang, A. Farhadi, I. Laptev, and A. Gupta · 2016
Later among the works it cites.
Affordance — wikipedia, the free encyclopedia, 2016
Wikipedia · 2016
Later among the works it cites.
Computational perception of physical object properties
J. Wu · 2016
Later among the works it cites.
Physics 101: Learning physical object properties from unlabeled videos
J. Wu, J. J. Lim, H. Zhang, J. B. Tenenbaum, and W. T. Freeman · 2016
Later among the works it cites.
Harnessing object and scene semantics for large-scale video understanding
Z. Wu, Y. Fu, Y.-G. Jiang, and L. Sigal · 2016
Later among the works it cites.
Situation recognition: Visual semantic role labeling for image understanding
M. Yatskar, L. Zettlemoyer, and A. Farhadi · 2016
Later among the works it cites.
Dense-captioning events in videos
R. Krishna, K. Hata, F. Ren, L. Fei-Fei, and J. C. Niebles · 2017
Closest in time.