The development of embodied cognition: Six lessons from babies
L. Smith and M. Gasser · 2005
Earlier work this paper cites.
The PASCAL visual object classes (VOC) challenge
M. Everingham, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman · 2010
Earlier work this paper cites.
HMDB: A large video database for human motion recognition
H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre · 2011
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay · 2011
Earlier work this paper cites.
WSABIE: Scaling up to large vocabulary image annotation
J. Weston, S. Bengio, and N. Usunier · 2011
Earlier work this paper cites.
UCF101: A dataset of 101 human actions classes from videos in the wild
Original
K. Soomro, A. R. Zamir, and M. Shah · 2012
Earlier work this paper cites.
DeViSE: A Deep Visual-Semantic Embedding Model
A. Frome, G. S. Corrado, J. Shlens, S. Bengio, J. Dean, M. Ranzato, and T. Mikolov · 2013
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Original
T. Mikolov, K. Chen, G. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Discriminative unsupervised feature learning with convolutional neural networks
A. Dosovitskiy, J. T. Springenberg, M. Riedmiller, and T. Brox · 2014
Earlier work this paper cites.
A multi-view embedding space for modeling internet images, tags, and their semantics
Y. Gong, Q. Ke, M. Isard, and S. Lazebnik · 2014
Earlier work this paper cites.
Improving image-sentence embeddings using large weakly annotated photo collections
Y. Gong, L. Wang, M. Hodosh, J. Hockenmaier, and S. Lazebnik · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Original
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Instructional videos for unsupervised harvesting and learning of action examples
S.-I. Yu, L. Jiang, and A. Hauptmann · 2014
Earlier work this paper cites.
Unsupervised visual representation learning by context prediction
C. Doersch, A. Gupta, and A. A. Efros · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Associating neural word embeddings with deep image representations using Fisher vectors
B. Klein, G. Lev, G. Sadeh, and L. Wolf · 2015
Earlier work this paper cites.
What’s cookin’? Interpreting cooking videos using text, speech and vision
J. Malmaud, J. Huang, V. Rathod, N. Johnston, A. Rabinovich, and K. Murphy · 2015
Earlier work this paper cites.
ESC: Dataset for Environmental Sound Classification
K. J. Piczak · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hockenmaier, and S. Lazebnik · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei · 2015
Earlier work this paper cites.
Unsupervised semantic parsing of video collections
O. Sener, A. R. Zamir, S. Savarese, and A. Saxena · 2015
Earlier work this paper cites.
Jointly modeling deep video and compositional text to bridge vision and language in a unified framework
R. Xu, C. Xiong, W. Chen, and J. J. Corso · 2015
Earlier work this paper cites.
YouTube-8M: A large-scale video classification benchmark
Original
S. Abu-El-Haija, N. Kothari, J. Lee, P. Natsev, G. Toderici, B. Varadarajan, and S. Vijayanarasimhan · 2016
Earlier work this paper cites.
Unsupervised learning from narrated instruction videos
J.-B. Alayrac, P. Bojanowski, N. Agrawal, I. Laptev, J. Sivic, and S. Lacoste-Julien · 2016
Earlier work this paper cites.
SoundNet: Learning sound representations from unlabeled video
Y. Aytar, C. Vondrick, and A. Torralba · 2016
Earlier work this paper cites.
Shuffle and learn: Unsupervised learning using temporal order verification
I. Misra, C. L. Zitnick, and M. Hebert · 2016
Earlier work this paper cites.
Unsupervised learning of visual representations by solving jigsaw puzzles
M. Noroozi and P. Favaro · 2016
Earlier work this paper cites.
Ambient sound provides supervision for visual learning
A. Owens, J. Wu, J. H. McDermott, W. T. Freeman, and A. Torralba · 2016
Earlier work this paper cites.
Jointly modeling embedding and translation to bridge video and language
Y. Pan, T. Mei, T. Yao, H. Li, and Y. Rui · 2016
Earlier work this paper cites.
Learning deep structure-preserving image-text embeddings
L. Wang, Y. Li, and S. Lazebnik · 2016
Earlier work this paper cites.