Fetching the paper…
Reading the bibliography…
Recent advances in deep reinforcement learning have made significant strides in performance on applications such as Go and Atari games.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
Thompson, William R · 1933
Earlier work this paper cites.
Mushroom records drawn from the audubon society field guide to north american mushrooms
Schlimmer, Jeff · 1981
Earlier work this paper cites.
The jackknife, the bootstrap and other resampling plans
Efron, Bradley · 1982
Earlier work this paper cites.
Fusion, propagation, and structuring in belief networks
Pearl, Judea · 1986
Earlier work this paper cites.
Keeping the neural networks simple by minimizing the description length of the weights
Hinton, Geoffrey E and Van Camp, Drew · 1993
Earlier work this paper cites.
Bayesian learning for neural networks
Neal, Radford M · 1994
Earlier work this paper cites.
Scaling up the accuracy of naive-bayes classifiers: A decision-tree hybrid
Kohavi, Ron · 1996
Earlier work this paper cites.
An introduction to variational methods for graphical models
Jordan, Michael I, Ghahramani, Zoubin, Jaakkola, Tommi S, and Saul, Lawrence K · 1999
Earlier work this paper cites.
Gaussian processes for classification: Mean-field algorithms
Opper, Manfred and Winther, Ole · 2000
Earlier work this paper cites.
Eigentaste: A constant time collaborative filtering algorithm
Goldberg, Ken, Roeder, Theresa, Gupta, Dhruv, and Perkins, Chris · 2001
Earlier work this paper cites.
Power ep
Minka, Thomas · 2004
Earlier work this paper cites.
Gaussian Processes for Machine Learning (Adaptive Computation and Machine Learning)
Rasmussen, Carl Edward and Williams, Christopher K. I · 2005
Earlier work this paper cites.
Pattern recognition and machine learning
Bishop, Christopher M · 2006
Earlier work this paper cites.
Sparse gaussian processes using pseudo-inputs
Snelson, Edward and Ghahramani, Zoubin · 2006
Earlier work this paper cites.
UCI machine learning repository, 2007
Asuncion, Arthur and Newman, David · 2007
Earlier work this paper cites.
Multi-task gaussian process prediction
Bonilla, Edwin V, Chai, Kian M., and Williams, Christopher · 2008
Earlier work this paper cites.
Graphical models, exponential families, and variational inference
Wainwright, Martin J, Jordan, Michael I, et al · 2008
Earlier work this paper cites.
Pure exploration in multi-armed bandits problems
Bubeck, Sébastien, Munos, Rémi, and Stoltz, Gilles · 2009
Cited alongside, same era.
Variational learning of inducing variables in sparse gaussian processes
Titsias, Michalis K · 2009
Cited alongside, same era.
Solving twoarmed Bernoulli bandit problems using a Bayesian learning automaton
Granmo, OleChristoffer · 2010
Cited alongside, same era.
The million song dataset
Bertin-Mahieux, Thierry, Ellis, Daniel P.W., Whitman, Brian, and Lamere, Paul · 2011
Cited alongside, same era.
An empirical evaluation of Thompson sampling
Chapelle, Olivier and Li, Lihong · 2011
Cited alongside, same era.
Practical variational inference for neural networks
Graves, Alex · 2011
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Rusu, Andrei A., Veness, Joel, Bellemare, Marc G., Graves, Alex, Riedmiller, Martin, Fidjeland, Andreas K., Ostrovski, Georg, Petersen, Stig, Beattie, Charles, Sadik, Amir, Antonoglou, Ioannis, King, Helen, Kumaran, Dharshan, Wierstra, Daan, Legg, Shane, and Hassabis, Demis · 2015
Later among the works it cites.
Scalable bayesian optimization using deep neural networks
Snoek, Jasper, Rippel, Oren, Swersky, Kevin, Kiros, Ryan, Satish, Nadathur, Sundaram, Narayanan, Patwary, Mostofa, Prabhat, Mr, and Adams, Ryan · 2015
Later among the works it cites.
Ba, Jimmy Lei, Kiros, Jamie Ryan, and Hinton, Geoffrey E · 2016
Later among the works it cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Gal, Yarin and Ghahramani, Zoubin · 2016
Later among the works it cites.
Black-box alpha divergence minimization
Hernández-Lobato, José Miguel, Li, Yingzhen, Rowland, Mark, Bui, Thang D., Hernández-Lobato, Daniel, and Turner, Richard E · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bayesian learning via stochastic gradient langevin dynamics
Welling, Max and Teh, Yee Whye · 2011
Cited alongside, same era.
Analysis of thompson sampling for the multi-armed bandit problem
Agrawal, Shipra and Goyal, Navin · 2012
Cited alongside, same era.
Bayesian posterior sampling via stochastic gradient fisher scoring
Ahn, Sungjin, Balan, Anoop Korattikara, and Welling, Max · 2012
Cited alongside, same era.
Stochastic gradient riemannian langevin dynamics on the probability simplex
Patterson, Sam and Teh, Yee Whye · 2013
Cited alongside, same era.
Expectation propagation as a way of life
Gelman, Andrew, Vehtari, Aki, Jylänki, Pasi, Robert, Christian, Chopin, Nicolas, and Cunningham, John P · 2014
Cited alongside, same era.
Learning to optimize via information-directed sampling
Russo, Dan and Van Roy, Benjamin · 2014
Cited alongside, same era.
Later among the works it cites.
Preconditioned stochastic gradient langevin dynamics for deep neural networks
Li, Chunyuan, Chen, Changyou, Carlson, David, and Carin, Lawrence · 2016
Later among the works it cites.
A variational analysis of stochastic gradient algorithms
Mandt, Stephan, Hoffman, Matthew D., and Blei, David M · 2016
Later among the works it cites.
Deep exploration via bootstrapped dqn
Osband, Ian, Blundell, Charles, Pritzel, Alexander, and Van Roy, Benjamin · 2016
Later among the works it cites.
Prioritized experience replay
Schaul, Tom, Quan, John, Antonoglou, Ioannis, and Silver, David · 2016
Later among the works it cites.
A practical method for solving contextual bandit problems using decision trees
Elmachtoub, Adam N, McNellis, Ryan, Oh, Sechan, and Petrik, Marek · 2017
Later among the works it cites.
Noisy networks for exploration
Fortunato, Meire, Azar, Mohammad Gheshlaghi, Piot, Bilal, Menick, Jacob, Osband, Ian, Graves, Alex, Mnih, Vlad, Munos, Remi, Hassabis, Demis, Pietquin, Olivier, Blundell, Charles, and Legg, Shane · 2017
Later among the works it cites.
Variational gaussian dropout is not bayesian
Hron, Jiri, Matthews, Alexander G de G, and Ghahramani, Zoubin · 2017
Later among the works it cites.
GPflow: A Gaussian process library using TensorFlow
Matthews, Alexander G. de G., van der Wilk, Mark, Nickson, Tom, Fujii, Keisuke., Boukouvalas, Alexis, León-Villagrá, Pablo, Ghahramani, Zoubin, and Hensman, James · 2017
Later among the works it cites.
Parameter space noise for exploration
Plappert, Matthias, Houthooft, Rein, Dhariwal, Prafulla, Sidor, Szymon, Chen, Richard Y., Chen, Xi, Asfour, Tamim, Abbeel, Pieter, and Andrychowicz, Marcin · 2017
Later among the works it cites.
Active learning for accurate estimation of linear models
Riquelme, Carlos, Ghavamzadeh, Mohammad, and Lazaric, Alessandro · 2017
Later among the works it cites.
Time-sensitive bandit learning and satisficing thompson sampling
Russo, Daniel, Tse, David, and Van Roy, Benjamin · 2017
Later among the works it cites.