A simple approximation for unbiased estimation of the standard deviation
John Gurland and Ram C Tripathi · 1971
Earlier work this paper cites.
Efficient pattern recognition using a new transformation distance
Patrice Simard, Yann LeCun, and John S Denker · 1993
Earlier work this paper cites.
Incorporating invariances in support vector learning machines
Bernhard Schölkopf, Chris Burges, and Vladimir Vapnik · 1996
Earlier work this paper cites.
Efficient backprop
Y. LeCun, L. Bottou, G. Orr, and K. Muller · 1998
Earlier work this paper cites.
Sphering and its properties
Guoying Li and Jian Zhang · 1998
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Best practices for convolutional neural networks applied to visual document analysis
Patrice Y Simard, Dave Steinkraus, and John C Platt · 2003
Earlier work this paper cites.
Beyond accuracy, f-score and roc: a family of discriminant measures for performance evaluation
Marina Sokolova, Nathalie Japkowicz, and Stan Szpakowicz · 2006
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Original
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Normalized online learning
Stephane Ross, Paul Mineiro, and John Langford · 2013
Earlier work this paper cites.
Learning with marginalized corrupted features
Laurens van der Maaten, Minmin Chen, Stephen Tyree, and Kilian Weinberger · 2013
Earlier work this paper cites.
Report on the 11th iwslt evaluation campaign, iwslt 2014
Mauro Cettolo, Jan Niehues, Sebastian Stüker, Luisa Bentivogli, and Marcello Federico · 2014
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Original
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
Original
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Original
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Deep unordered composition rivals syntactic methods for text classification
Mohit Iyyer, Varun Manjunatha, Jordan Boyd-Graber, and Hal Daumé III · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Improving neural machine translation models with monolingual data
Original
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2015
Earlier work this paper cites.
3d shapenets: A deep representation for volumetric shapes
Zhirong Wu, Shuran Song, Aditya Khosla, Fisher Yu, Linguang Zhang, Xiaoou Tang, and Jianxiong Xiao · 2015
Earlier work this paper cites.
Layer normalization
Original
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Deep networks with stochastic depth
Gao Huang, Yu Sun, Zhuang Liu, Daniel Sedra, and Kilian Q Weinberger · 2016
Earlier work this paper cites.