Boris T Polyak, “Some methods of speeding up the convergence of iteration methods,”
1964
Earlier work this paper cites.
Raia Hadsell, Sumit Chopra, and Yann LeCun, “Dimensionality reduction by learning an invariant mapping,” in
2006
Earlier work this paper cites.
Rong-En Fan, Kai-Wei Chang, Cho-Jui Hsieh, Xiang-Rui Wang, and Chih-Jen Lin, “LIBLINEAR: A library for large linear classification,”
2008
Earlier work this paper cites.
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in
2009
Earlier work this paper cites.
Mark Everingham, Luc Van Gool, Christopher K. I. Williams, John M. Winn, and Andrew Zisserman, “The pascal visual object classes (VOC) challenge,”
2009
Earlier work this paper cites.
Michael Gutmann and Aapo Hyvärinen, “Noise-contrastive estimation: A new estimation principle for unnormalized statistical models,” in
2010
Earlier work this paper cites.
Vicente Ordonez, Girish Kulkarni, and Tamara L. Berg, “Im2Text: Describing images using 1 million captioned photographs,” in
2011
Earlier work this paper cites.
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay, “Scikit-learn: Machine learning in Python,”
2011
Earlier work this paper cites.
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton, “Imagenet classification with deep convolutional neural networks,” in
2012
Earlier work this paper cites.
James Martens and Ilya Sutskever, “Training deep and recurrent networks with hessian-free optimization,” in
2012
Earlier work this paper cites.
Micah Hodosh, Peter Young, and Julia Hockenmaier, “Framing image description as a ranking task: Data, models and evaluation metrics,”
2013
Earlier work this paper cites.
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton, “On the importance of initialization and momentum in deep learning,” in
2013
Earlier work this paper cites.
Jeff Donahue, Yangqing Jia, Oriol Vinyals, Judy Hoffman, Ning Zhang, Eric Tzeng, and Trevor Darrell, “Decaf: A deep convolutional activation feature for generic visual recognition,” in
2014
Earlier work this paper cites.
Ali Sharif Razavian, Hossein Azizpour, Josephine Sullivan, and Stefan Carlsson, “Cnn features off-the-shelf: an astounding baseline for recognition,” in
2014
Earlier work this paper cites.
Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik, “Rich feature hierarchies for accurate object detection and semantic segmentation,” in
2014
Earlier work this paper cites.
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick, “Microsoft COCO: Common objects in context,” in
2014
Earlier work this paper cites.
Jia Deng, Olga Russakovsky, Jonathan Krause, Michael S Bernstein, Alex Berg, and Li Fei-Fei, “Scalable multi-label annotation,” in
2014
Earlier work this paper cites.
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier, “From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions,”
2014
Earlier work this paper cites.
Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, and Tamara Berg, “Referitgame: Referring to objects in photographs of natural scenes,” in
2014
Earlier work this paper cites.
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov, “Dropout: a simple way to prevent neural networks from overfitting,”
2014
Earlier work this paper cites.
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein,
2015
Earlier work this paper cites.
Jonathan Long, Evan Shelhamer, and Trevor Darrell, “Fully convolutional networks for semantic segmentation,” in
2015
Earlier work this paper cites.
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan, “Show and tell: A neural image caption generator,” in
2015
Earlier work this paper cites.
Andrej Karpathy and Li Fei-Fei, “Deep visual-semantic alignments for generating image descriptions,” in
2015
Earlier work this paper cites.
Jeffrey Donahue, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Darrell, “Long-term recurrent convolutional networks for visual recognition and description,” in
2015
Earlier work this paper cites.
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh, “VQA: Visual question answering,” in
2015
Earlier work this paper cites.
Carl Doersch, Abhinav Gupta, and Alexei A Efros, “Unsupervised visual representation learning by context prediction,” in
2015
Earlier work this paper cites.
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C Lawrence Zitnick, “Microsoft COCO captions: Data collection and evaluation server,”
Original
2015
Earlier work this paper cites.
Sergey Ioffe and Christian Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in
2015
Earlier work this paper cites.
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh, “CIDEr: Consensus-based image description evaluation,” in
2015
Earlier work this paper cites.
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun, “Faster R-CNN: Towards real-time object detection with region proposal networks,” in
2015
Earlier work this paper cites.
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, “Deep residual learning for image recognition,” in
2016
Earlier work this paper cites.
Yuke Zhu, Oliver Groth, Michael Bernstein, and Li Fei-Fei, “Visual7w: Grounded question answering in images,” in
2016
Earlier work this paper cites.
Deepak Pathak, Philipp Krahenbuhl, Jeff Donahue, Trevor Darrell, and Alexei A Efros, “Context encoders: Feature learning by inpainting,” in
2016
Earlier work this paper cites.
Richard Zhang, Phillip Isola, and Alexei A Efros, “Colorful image colorization,” in
2016
Earlier work this paper cites.
Ranjay A Krishna, Kenji Hata, Stephanie Chen, Joshua Kravitz, David A Shamma, Li Fei-Fei, and Michael S Bernstein, “Embracing error to enable rapid crowdsourcing,” in
2016
Earlier work this paper cites.
Junhua Mao, Jiajing Xu, Kevin Jing, and Alan L Yuille, “Training and evaluating multimodal word embeddings with large-scale web annotated images,” in
2016
Earlier work this paper cites.
Bart Thomee, David A Shamma, Gerald Friedland, Benjamin Elizalde, Karl Ni, Douglas Poland, Damian Borth, and Li-Jia Li, “YFCC100M: The new data in multimedia research,”
2016
Earlier work this paper cites.
Mehdi Noroozi and Paolo Favaro, “Unsupervised learning of visual representations by solving jigsaw puzzles,” in
2016
Earlier work this paper cites.