Fetching the paper…
Reading the bibliography…
Robots that are trained to perform a task in a fixed environment often fail when facing unexpected changes to the environment due to a lack of exploration.
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu, “Asynchronous methods for deep reinforcement learning,” in International Conference on Machine Learning , 2016, pp. 1928–1937
1937
Earlier work this paper cites.
N. Littlestone, “Learning quickly when irrelevant attributes abound: A new linear-threshold algorithm,” Machine learning , vol. 2, no. 4, pp. 285–318, 1988
1988
Earlier work this paper cites.
V. Gullapalli, “A stochastic reinforcement learning algorithm for learning real-valued functions,” Neural networks , vol. 3, no. 6, pp. 671–692, 1990
1990
Earlier work this paper cites.
R. J. Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” Machine Learning , vol. 8, no. 3-4, pp. 229–256, 1992
1992
Earlier work this paper cites.
R. Agrawal, “The continuum-armed bandit problem,” SIAM Journal on Control and Optimization , vol. 33, no. 6, pp. 1926–1951, 1995
1995
Earlier work this paper cites.
R. J. Hyndman, “Computing and graphing highest density regions,” The American Statistician , vol. 50, no. 2, pp. 120–126, 1996
1996
Earlier work this paper cites.
S. P. Choi, D.-Y. Yeung, and N. L. Zhang, “Hidden-mode markov decision processes for nonstationary sequential decision making,” in Sequence Learning . Springer, 2000, pp. 264–287
2000
Earlier work this paper cites.
Y. Sakaguchi and M. Takano, “Reliability of internal prediction/estimation and its application. i. adaptive action selection reflecting reliability of value function,” Neural Networks , vol. 17, no. 7, pp. 935–952, 2004
2004
Earlier work this paper cites.
2004
Earlier work this paper cites.
B. C. Da Silva, E. W. Basso, A. L. Bazzan, and P. M. Engel, “Dealing with non-stationary environments using context detection,” in Proceedings of the 23rd international conference on Machine learning . ACM, 2006, pp. 217–224
2006
Earlier work this paper cites.
L. Li, M. L. Littman, and T. J. Walsh, “Knows what it knows: A framework for self-aware learning,” in Proceedings of the 25th International Conference on Machine Learning , ser. ICML ’08. New York, NY, USA: ACM, 2008, pp. 568–575
2008
Earlier work this paper cites.
Quionero-Candela, Joaquin and Sugiyama, Masashi and Schwaighofer, Anton and Lawrence, Neil D, Dataset Shift in Machine Learning . The MIT Press, 2009
2009
Earlier work this paper cites.
M. Tokic, “Adaptive ϵ \epsilon -Greedy exploration in reinforcement learning based on value differences,” in KI 2010: Advances in Artificial Intelligence , ser. Lecture Notes in Computer Science. Springer, Berlin, Heidelberg, 2010, pp. 203–210
2010
Earlier work this paper cites.
M. Sugiyama and M. Kawanabe, Machine learning in non-stationary environments: Introduction to covariate shift adaptation . MIT press, 2012
2012
Cited alongside, same era.
T. Degris, P. M. Pilarski, and R. S. Sutton, “Model-free reinforcement learning with continuous action in practice,” in American Control Conference (ACC), 2012 . IEEE, 2012, pp. 2177–2182
2012
Cited alongside, same era.
E. Todorov, T. Erez, and Y. Tassa, “Mujoco: A physics engine for model-based control,” in Intelligent Robots and Systems (IROS), 2012 IEEE/RSJ International Conference on . IEEE, 2012, pp. 5026–5033
2012
Cited alongside, same era.
2014
Cited alongside, same era.
2017
Later among the works it cites.
2017
Later among the works it cites.
T. Haarnoja, H. Tang, P. Abbeel, and S. Levine, “Reinforcement learning with deep Energy-Based policies,” in International Conference on Machine Learning , 2017
2017
Later among the works it cites.
O. Nachum, M. Norouzi, K. Xu, and D. Schuurmans, “Bridging the gap between value and policy based reinforcement learning,” in Advances in Neural Information Processing Systems , 2017
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2015
Cited alongside, same era.
2016
Cited alongside, same era.
I. Osband, C. Blundell, A. Pritzel, and B. Van Roy, “Deep exploration via bootstrapped DQN,” in Advances in Neural Information Processing Systems , D. D. Lee, M. Sugiyama, U. V. Luxburg, I. Guyon, and R. Garnett, Eds., 2016, pp. 4026–4034
2016
Cited alongside, same era.
R. Fox, A. Pakman, and N. Tishby, “Taming the noise in reinforcement learning via soft updates,” in Conference on Uncertainty in Artificial Intelligence , 2016
2016
Cited alongside, same era.
M. G. Bellemare, S. Srinivasan, G. Ostrovski, T. Schaul, D. Saxton, and R. Munos, “Unifying Count-Based exploration and intrinsic motivation,” in Advances in Neural Information Processing Systems , 2016
2016
Cited alongside, same era.
R. Houthooft, X. Chen, Y. Duan, J. Schulman, F. De Turck, and P. Abbeel, “Curiosity-driven exploration in deep reinforcement learning via bayesian neural networks,” in Advances in Neural Information Processing Systems , 2016
2016
Cited alongside, same era.
A. Tamar, D. D. Castro, and S. Mannor, “Learning the variance of the reward-to-go,” Journal of Machine Learning Research , vol. 17, no. 13, pp. 1–36, 2016. [Online]. Available: http://jmlr.org/papers/v17/14-335.html
2016
Cited alongside, same era.
A. Rajeswaran, S. Ghotra, B. Ravindran, and S. Levine, “EPOpt: Learning robust neural network policies using model ensembles,” in International Conference on Learning Representations , 2016
2016
Cited alongside, same era.
H. Tang, R. Houthooft, D. Foote, A. Stooke, X. Chen, Y. Duan, J. Schulman, F. De Turck, and P. Abbeel, “#exploration: A study of Count-Based exploration for deep reinforcement learning,” in Advances in Neural Information Processing Systems , 2017
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
C. Finn, P. Abbeel, and S. Levine, “Model-Agnostic Meta-Learning for fast adaptation of deep networks,” in International Conference on Machine Learning , 2017
2017
Later among the works it cites.
2018
Later among the works it cites.
X. B. Peng, M. Andrychowicz, W. Zaremba, and P. Abbeel, “Sim-to-Real transfer of robotic control with dynamics randomization,” in IEEE International Conference on Robotics and Automation , 2018
2018
Later among the works it cites.
M. Al-Shedivat, T. Bansal, Y. Burda, I. Sutskever, I. Mordatch, and P. Abbeel, “Continuous adaptation via Meta-Learning in nonstationary and competitive environments,” in International Conference on Learning Representations , 2018
2018
Later among the works it cites.
L. Yang, “Active learning with a drifting distribution,” in Advances in Neural Information Processing Systems 24 , J. Shawe-Taylor, R. S. Zemel, P. L. Bartlett, F. Pereira, and K. Q. Weinberger, Eds. Curran Associates, Inc., 2011, pp. 2079–2087
2087
Closest in time.