Fetching the paper…
Reading the bibliography…
Achieving efficient and scalable exploration in complex domains poses a major challenge in reinforcement learning.
M. Kearns and D. Koller, Efficient reinforcement learning in factored MDPs. Proc. IJCAI, 1999
1999
Earlier work this paper cites.
W. D. Smart and L. P. Kaelbling, Practical reinforcement learning in continuous spaces. Proc. ICML, 2000
2000
Earlier work this paper cites.
T. Hastie, R. Tibshirani, and J. Friedman, The Elements of Statistical Learning . Springer series in statistics. Springer, New York, 2001
2001
Earlier work this paper cites.
R. I. Brafman and M. Tennenholtz, R-max, a general polynomial time algorithm for near-optimal reinforcement learning . Journal of Machine Learning Research, 2002
2002
Earlier work this paper cites.
M. Kearns and S. Singh, Near-optimal reinforcement learning in polynomial time. Machine Learning Journal, 2002
2002
Earlier work this paper cites.
Kakade, S., Kearns, M., and Langford, J. (2003). Exploration in metric state spaces. Proc. ICML
2003
Earlier work this paper cites.
M. Belkin, P. Niyogi and V. Sindhwani, Manifold Regularization: A Geometric Framework for Learning from Labeled and Unlabeled Examples . JMLR, vol. 7, pp. 2399-2434, Nov. 2006
2006
Earlier work this paper cites.
G. E. Hinton and R. Salakhutdinov, Reducing the dimensionality of data with neural networks . Science, 313, 504–507, 2006
2006
Cited alongside, same era.
Juergen Schmidhuber Developmental Robotics, Optimal Artificial Curiosity, Creativity, Music, and the Fine Arts. . Connection Science, vol. 18 (2), p 173-187. 2006
2006
Cited alongside, same era.
J. Z. Kolter and A. Y. Ng, Near-Bayesian Exploration in Polynomial Time Proceedings of the 26th Annual International Conference on Machine Learning, pp. 1Ð8, 2009
2009
Cited alongside, same era.
J. Sorg, S. Singh, R. L. Lewis, Variance-Based Rewards for Approximate Bayesian Reinforcement Learning . Proc. UAI, 2010
2010
Cited alongside, same era.
M. Geist and O. Pietquin, Managing Uncertainty within Value Function Approximation in Reinforcement Learning . W. on Active Learning and Experimental Design, 2010
2010
J. Pazis and R. Parr, PAC Optimal Exploration in Continuous Space Markov Decision Processes . Proc. AAAI, 2013
2013
Later among the works it cites.
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling, The Arcade Learning Environment: An Evaluation Platform for General Agents . JAIR. Volume 47, p.235-279. June 2013
2013
Later among the works it cites.
F. Doshi-Velez, D. Wingate, N. Roy, and J. Tenenbaum, Nonparametric Bayesian Policy Priors for Reinforcement Learning . NIPS, 2014
2014
Later among the works it cites.
T. Lang, M. Toussaint, K. Keristing, Exploration in relational domains for model-based reinforcement learning Proc. AAMAS, 2014
2014
Later among the works it cites.
A. Guez, D. Silver, P. Dayan, Efficient Bayes-Adaptive Reinforcement Learning using Sample-Based Search . NIPS, 2014
2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
M. Lopes, T. Lang, M. Toussaint and P.-Y. Oudeyer, Exploration in Model-based Reinforcement Learning by Empirically Estimating Learning Progress . NIPS, 2012
2012
Cited alongside, same era.
M. Araya, O. Buffet, and V. Thomas. Near-optimal BRL Using Optimistic Local Transitions . (ICML-12), ser. ICML â12, J. Langford and J. Pineau, Eds. New York, NY, USA: Omnipress, Jul. 2012, pp. 97-104
2012
Cited alongside, same era.
A. L. Strehl and M. L. Littman, An Analysis of Model-Based Interval Estimation for Markov Decision Processes. Journal of Computer and System Sciences, 74, 1209Ð1331
Cited in the paper.
Y Gal, Z Ghahramani. Dropout as a Bayesian approximation: Insights and applications Deep Learning Workshop, ICML
Cited in the paper.
David Carmel and Shaul Markovitch. Exploration Strategies for Model-based Learning in Multi-agent Systems Autonomous Agents and Multi-Agent Systems Volume 2, Issue 2 , pp 141-172
Cited in the paper.
John Schulman, Sergey Levine, Philipp Moritz, Michael I. Jordan, Pieter Abbeel. Trust Region Policy Optimization Arxiv preprint 1502.05477
Cited in the paper.
X. Guo, S. Singh, H. Lee, R. Lewis, X. Wang, Deep Learning for Real-Time Atari Game Play Using Offline Monte-Carlo Tree Search Planning . NIPS, 2014
2014
Later among the works it cites.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, et al. Human-level Control Through Deep Reinforcement Learning. Nature, 518(7540):529Ð533, 2015
2015
Closest in time.