Fetching the paper…
Reading the bibliography…
Dropout has been demonstrated as a simple and effective module to not only regularize the training process of deep neural networks, but also provide the uncertainty estimation for prediction.
A method for solving the convex programming problem with convergence rate o (1/kˆ 2)
Yurii E Nesterov · 1983
Earlier work this paper cites.
A practical bayesian framework for backpropagation networks
David JC MacKay · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
An empirical evaluation of deep architectures on problems with many factors of variation
Hugo Larochelle, Dumitru Erhan, Aaron Courville, James Bergstra, and Yoshua Bengio · 2007
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky et al · 2009
Earlier work this paper cites.
MNIST handwritten digit database
Yann LeCun, Corinna Cortes, and CJ Burges · 2010
Earlier work this paper cites.
Practical variational inference for neural networks
Alex Graves · 2011
Earlier work this paper cites.
Bayesian learning via stochastic gradient langevin dynamics
Max Welling and Yee W Teh · 2011
Earlier work this paper cites.
Normality tests for statistical analysis: A guide for non-statisticians
Asghar Ghasemi and Saleh Zahediasl · 2012
Earlier work this paper cites.
Improving neural networks by preventing co-adaptation of feature detectors
Geoffrey E Hinton, Nitish Srivastava, Alex Krizhevsky, Ilya Sutskever, and Ruslan R Salakhutdinov · 2012
Earlier work this paper cites.
Bayesian learning for neural networks , volume 118
Radford M Neal · 2012
Earlier work this paper cites.
Adaptive dropout for training deep neural networks
Jimmy Ba and Brendan Frey · 2013
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Yoshua Bengio, Nicholas Léonard, and Aaron Courville · 2013
Earlier work this paper cites.
Stochastic variational inference
Matthew D Hoffman, David M Blei, Chong Wang, and John Paisley · 2013
Earlier work this paper cites.
Auto-encoding variational Bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Regularization of neural networks using dropconnect
Li Wan, Matthew Zeiler, Sixin Zhang, Yann Le Cun, and Rob Fergus · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning · 2014
Earlier work this paper cites.
Stochastic backpropagation and approximate inference in deep generative models
Danilo Jimenez Rezende, Shakir Mohamed, and Daan Wierstra · 2014
Cited alongside, same era.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Cited alongside, same era.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Ian Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dan Mané, Rajat Monga, Sherry Moore, Derek Murray, Chris Olah, Mike Schuster, Jonathon Shlens, Benoit Steiner, Ilya Sutskever, Kunal Talwar, Paul Tucker, Vincent Vanhoucke, Vijay Vasudevan, Fernanda Viégas, Oriol Vinyals, Pete Warden, Martin Wattenberg, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng · 2015
Cited alongside, same era.
Conditional computation in neural networks for faster models
Emmanuel Bengio, Pierre-Luc Bacon, Joelle Pineau, and Doina Precup · 2015
Cited alongside, same era.
Variational dropout sparsifies deep neural networks
Dmitry Molchanov, Arsenii Ashukha, and Dmitry Vetrov · 2017
Later among the works it cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean · 2017
Later among the works it cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Later among the works it cites.
Latent alignment and variational attention
Yuntian Deng, Yoon Kim, Justin Chiu, Demi Guo, and Alexander Rush · 2018
Later among the works it cites.
Squeeze-and-excitation networks
Jie Hu, Li Shen, and Gang Sun · 2018
Later among the works it cites.
Accurate uncertainties for deep learning using calibrated regression
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Charles Blundell, Julien Cornebise, Koray Kavukcuoglu, and Daan Wierstra · 2015
Cited alongside, same era.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Cited alongside, same era.
Variational dropout and the local reparameterization trick
Durk P Kingma, Tim Salimans, and Max Welling · 2015
Cited alongside, same era.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Cited alongside, same era.
Obtaining well calibrated probabilities using Bayesian binning
Mahdi Pakdaman Naeini, Gregory Cooper, and Milos Hauskrecht · 2015
Cited alongside, same era.
Faster R-CNN: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Cited alongside, same era.
Efficient object localization using convolutional networks
Jonathan Tompson, Ross Goroshin, Arjun Jain, Yann LeCun, and Christoph Bregler · 2015
Cited alongside, same era.
Dropout as a Bayesian approximation: Representing model uncertainty in deep learning
Yarin Gal and Zoubin Ghahramani · 2016
Cited alongside, same era.
Volodymyr Kuleshov, Nathan Fenner, and Stefano Ermon · 2018
Later among the works it cites.
Evaluating bayesian deep learning methods for semantic segmentation
Jishnu Mukhoti and Yarin Gal · 2018
Later among the works it cites.
Kernel implicit variational inference
Jiaxin Shi, Shengyang Sun, and Jun Zhu · 2018
Later among the works it cites.
Hydranets: Specialized dynamic architectures for efficient inference
Ravi Teja Mullapudi, William R Mark, Noam Shazeer, and Kayvon Fatahalian · 2018
Later among the works it cites.
Tips and tricks for visual question answering: Learnings from the 2017 challenge
Damien Teney, Peter Anderson, Xiaodong He, and Anton Van Den Hengel · 2018
Later among the works it cites.
ARM: Augment-REINFORCE-merge gradient for discrete latent variable models
Mingzhang Yin and Mingyuan Zhou · 2018
Later among the works it cites.
L0-ARM: Network sparsification via stochastic binary optimization
Yang Li and Shihao Ji · 2019
Later among the works it cites.
Thompson sampling via local uncertainty
Zhendong Wang and Mingyuan Zhou · 2019
Later among the works it cites.
Deep modular co-attention networks for visual question answering
Zhou Yu, Jun Yu, Yuhao Cui, Dacheng Tao, and Qi Tian · 2019
Later among the works it cites.
Learnable Bernoulli dropout for Bayesian deep learning
Shahin Boluki, Randy Ardywibowo, Siamak Zamani Dadaneh, Mingyuan Zhou, and Xiaoning Qian · 2020
Later among the works it cites.
DisARM: An antithetic gradient estimator for binary latent variables
Zhe Dong, Andriy Mnih, and George Tucker · 2020
Later among the works it cites.
Bayesian attention modules
Xinjie Fan, Shujian Zhang, Bo Chen, and Mingyuan Zhou · 2020
Later among the works it cites.
Probabilistic Best Subset Selection by Gradient-Based Optimization
Mingzhang Yin, Nhat Ho, Bowei Yan, Xiaoning Qian, and Mingyuan Zhou · 2020
Later among the works it cites.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio · 2057
Closest in time.