Fetching the paper…
Reading the bibliography…
In this work we explore a straightforward variational Bayes scheme for Recurrent Neural Networks.
Bayesian back-propagation
Wray L Buntine and Andreas S Weigend · 1991
Earlier work this paper cites.
Transforming neural-net output levels to probability distributions
John S Denker and Yann Lecun · 1991
Earlier work this paper cites.
Keeping the neural networks simple by minimizing the description length of the weights
Geoffrey E Hinton and Drew Van Camp · 1993
Earlier work this paper cites.
Building a large annotated corpus of English: The Penn Treebank
Mitchell P Marcus, Mary Ann Marcinkiewicz, and Beatrice Santorini · 1993
Earlier work this paper cites.
Simplifying neural nets by discovering flat minima
Sepp Hochreiter, Jürgen Schmidhuber, et al · 1995
Earlier work this paper cites.
Probable networks and plausible predictions—a review of practical Bayesian methods for supervised neural networks
David JC MacKay · 1995
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Variational algorithms for approximate Bayesian inference
Matthew James Beal · 2003
Earlier work this paper cites.
Graphical models, exponential families, and variational inference
Martin J Wainwright, Michael I Jordan, et al · 2008
Earlier work this paper cites.
Recurrent neural network based language model
Tomas Mikolov, Martin Karafiát, Lukas Burget, Jan Cernockỳ, and Sanjeev Khudanpur · 2010
Earlier work this paper cites.
Recursive bayesian recurrent neural networks for time-series modeling
Derrick T Mirikitani and Nikolay Nikolaev · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Practical variational inference for neural networks
Alex Graves · 2011
Earlier work this paper cites.
Two problems with variational expectation maximisation for time-series models
Richard E Turner and Maneesh Sahani · 2011
Earlier work this paper cites.
Bayesian learning via stochastic gradient Langevin dynamics
Max Welling and Yee W Teh · 2011
Earlier work this paper cites.
Bayesian learning for neural networks , volume 118
Radford M Neal · 2012
Earlier work this paper cites.
Auto-encoding variational Bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Fast dropout training
Sida I Wang and Christopher D Manning · 2013
Earlier work this paper cites.
Learning stochastic recurrent networks
Justin Bayer and Christian Osendorfer · 2014
Cited alongside, same era.
Variational recurrent auto-encoders
Otto Fabius and Joost R van Amersfoort · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Cited alongside, same era.
Stochastic backpropagation and approximate inference in deep generative models
Danilo Jimenez Rezende, Shakir Mohamed, and Daan Wierstra · 2014
Cited alongside, same era.
Probabilistic line searches for stochastic optimization
Maren Mahsereci and Philipp Hennig · 2015
Later among the works it cites.
Learning to learn by gradient descent by gradient descent
Marcin Andrychowicz, Misha Denil, Sergio Gomez, Matthew W Hoffman, David Pfau, Tom Schaul, and Nando de Freitas · 2016
Later among the works it cites.
Bayesian recurrent neural network for language modeling
Jen-Tzung Chien and Yuan-Chu Ku · 2016
Later among the works it cites.
A theoretically grounded application of dropout in recurrent neural networks
Yarin Gal and Zoubin Ghahramani · 2016
Later among the works it cites.
Scalable Bayesian learning of recurrent neural networks for language modeling
Zhe Gan, Chunyuan Li, Changyou Chen, Yunchen Pu, Qinliang Su, and Lawrence Carin · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Expectation backpropagation: Parameter-free training of multilayer neural networks with continuous or discrete weights
Daniel Soudry, Itay Hubara, and Ron Meir · 2014
Cited alongside, same era.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey E Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Cited alongside, same era.
Recurrent neural network regularization
Wojciech Zaremba, Ilya Sutskever, and Oriol Vinyals · 2014
Cited alongside, same era.
Deep speech 2: End-to-end speech recognition in English and Mandarin
Dario Amodei, Rishita Anubhai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Jingdong Chen, Mike Chrzanowski, Adam Coates, Greg Diamos, et al · 2015
Cited alongside, same era.
Weight uncertainty in neural networks
Charles Blundell, Julien Cornebise, Koray Kavukcuoglu, and Daan Wierstra · 2015
Cited alongside, same era.
A recurrent latent variable model for sequential data
Junyoung Chung, Kyle Kastner, Laurent Dinh, Kratarth Goel, Aaron C Courville, and Yoshua Bengio · 2015
Cited alongside, same era.
Dropout as a Bayesian approximation: Representing model uncertainty in deep learning
Yarin Gal and Zoubin Ghahramani · 2015
Cited alongside, same era.
Alex Graves · 2016
Later among the works it cites.
Curiositydriven exploration in deep reinforcement learning via Bayesian neural networks
Rein Houthooft, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel · 2016
Later among the works it cites.
Exploring the limits of language modeling
Rafal Jozefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu · 2016
Later among the works it cites.
Ke Li and Jitendra Malik · 2016
Later among the works it cites.
Efficient exploration for dialogue policy learning with BBQ networks & replay buffer spiking
Zachary C Lipton, Jianfeng Gao, Lihong Li, Xiujun Li, Faisal Ahmed, and Li Deng · 2016
Later among the works it cites.
Optimization of image description metrics using policy gradient methods
Siqi Liu, Zhenhai Zhu, Ning Ye, Sergio Guadarrama, and Kevin Murphy · 2016
Later among the works it cites.
Knowing when to look: Adaptive attention via a visual sentinel for image captioning
Jiasen Lu, Caiming Xiong, Devi Parikh, and Richard Socher · 2016
Later among the works it cites.
Pointer Sentinel Mixture Models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Later among the works it cites.
Hierarchical variational models
Rajesh Ranganath, Dustin Tran, and David Blei · 2016
Later among the works it cites.
Show and tell: Lessons learned from the 2015 mscoco image captioning challenge
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan · 2016
Later among the works it cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al · 2016
Later among the works it cites.
Julian Georg Zilly, Rupesh Kumar Srivastava, Jan Koutník, and Jürgen Schmidhuber · 2016
Later among the works it cites.