Fetching the paper…
Reading the bibliography…
Modern deep learning is primarily an experimental science, in which empirical advances occasionally come at the expense of probabilistic rigor.
The statistical analysis of compositional data
John Aitchison · 1982
Earlier work this paper cites.
Multivariate methods for proportional shape
G Campbell and J Mosimann · 1987
Earlier work this paper cites.
A simple weight decay can improve generalization
Anders Krogh and John A Hertz · 1992
Earlier work this paper cites.
Principles of compositional data analysis
John Aitchison · 1994
Earlier work this paper cites.
Training with noise is equivalent to tikhonov regularization
Chris M Bishop · 1995
Earlier work this paper cites.
Logratios and natural laws in compositional data analysis
John Aitchison · 1999
Earlier work this paper cites.
Isometric logratio transformations for compositional data analysis
Juan José Egozcue, Vera Pawlowsky-Glahn, Glòria Mateu-Figueras, and Carles Barcelo-Vidal · 2003
Earlier work this paper cites.
Modelling compositional data using dirichlet regression models
Rafiq H Hijazi and Robert W Jernigan · 2009
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeffrey Dean · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Actor-mimic: Deep multitask and transfer reinforcement learning
Emilio Parisotto, Jimmy Lei Ba, and Ruslan Salakhutdinov · 2015
Cited alongside, same era.
Simultaneous deep transfer across domains and tasks
Eric Tzeng, Judy Hoffman, Trevor Darrell, and Kate Saenko · 2015
Cited alongside, same era.
Learning transferable architectures for scalable image recognition
Barret Zoph, Vijay Vasudevan, Jonathon Shlens, and Quoc V Le · 2018
Later among the works it cites.
Rethinking the usage of batch normalization and dropout in the training of deep neural networks
Guangyong Chen, Pengfei Chen, Yujun Shi, Chang-Yu Hsieh, Benben Liao, and Shengyu Zhang · 2019
Later among the works it cites.
Batch normalization is a cause of adversarial vulnerability
Angus Galloway, Anna Golubeva, Thomas Tanay, Medhat Moussa, and Graham W Taylor · 2019
Later among the works it cites.
Understanding the disharmony between dropout and batch normalization by variance shift
Xiang Li, Shuo Chen, Xiaolin Hu, and Jian Yang · 2019
Later among the works it cites.
The continuous bernoulli: fixing a pervasive error in variational autoencoders
Gabriel Loaiza-Ganem and John P Cunningham · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jan Chorowski and Navdeep Jaitly · 2016
Cited alongside, same era.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2016
Cited alongside, same era.
L2 regularization versus batch and weight normalization
Twan Van Laarhoven · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Do deep nets really need weight decay and dropout?
Alex Hernández-García and Peter König · 2018
Cited alongside, same era.
When does label smoothing help?
Rafael Müller, Simon Kornblith, and Geoffrey E Hinton · 2019
Later among the works it cites.
Regularized evolution for image classifier architecture search
Esteban Real, Alok Aggarwal, Yanping Huang, and Quoc V Le · 2019
Later among the works it cites.
A survey on image data augmentation for deep learning
Connor Shorten and Taghi M Khoshgoftaar · 2019
Later among the works it cites.
Dropout vs. batch normalization: an empirical study of their impact to deep learning
Christian Garbin, Xingquan Zhu, and Oge Marques · 2020
Closest in time.
The continuous categorical: a novel simplex-valued exponential family
Elliott Gordon-Rodriguez, Gabriel Loaiza-Ganem, and John P Cunningham · 2020
Closest in time.