Fetching the paper…
Reading the bibliography…
Deep reinforcement learning algorithms have been shown to learn complex tasks using highly general policy classes.
A possibility for implementing curiosity and boredom in model-building neural controllers
Schmidhuber, Jürgen · 1990
Earlier work this paper cites.
R-max – a general polynomial time algorithm for near-optimal reinforcement learning
Brafman, Ronen I. and Tennenholtz, Moshe · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Kearns, Michael and Singh, Satinder · 2002
Earlier work this paper cites.
Learning Options in Reinforcement Learning
Stolle, Martin and Precup, Doina · 2002
Earlier work this paper cites.
Recent advances in hierarchical reinforcement learning
Barto, Andrew G. and Mahadevan, Sridhar · 2003
Earlier work this paper cites.
Exploration in metric state spaces
Kakade, Sham, Kearns, Michael, and Langford, John · 2003
Earlier work this paper cites.
Intrinsically Motivated Reinforcement Learning
Chentanez, Nuttapong, Barto, Andrew G, and Singh, Satinder P · 2005
Earlier work this paper cites.
Near-bayesian exploration in polynomial time
Kolter, J. Zico and Ng, Andrew Y · 2009
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
Strehl, Alexander L. and Littman, Michael L · 2009
Earlier work this paper cites.
An empirical evaluation of thompson sampling
Chapelle, O. and Li, Lihong · 2011
Earlier work this paper cites.
Ensemble of exemplar-svms for object detection and beyond
Malisiewicz, Tomasz, Gupta, Abhinav, and Efros, Alexei A · 2011
Earlier work this paper cites.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Bubeck, Sébastien and Cesa-Bianchi, Nicolò · 2012
Cited alongside, same era.
Pac optimal exploration in continuous space markov decision processes
Pazis, Jason and Parr, Ronald · 2013
Cited alongside, same era.
Generative adversarial nets
Goodfellow, Ian, Pouget-Abadie, Jean, Mirza, Mehdi, Xu, Bing, Warde-Farley, David, Ozair, Sherjil, Courville, Aaron, and Bengio, Yoshua · 2014
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, Timothy P., Hunt, Jonathan J., Pritzel, Alexander, Heess, Nicolas, Erez, Tom, Tassa, Yuval, Silver, David, and Wierstra, Daan · 2015
Cited alongside, same era.
Deep multi-scale video prediction beyond mean square error
Mathieu, Michaël, Couprie, Camille, and LeCun, Yann · 2015
Cited alongside, same era.
Benchmarking deep reinforcement learning for continuous control
Duan, Yan, Chen, Xi, Houthooft, Rein, Schulman, John, and Abbeel, Pieter · 2016
Later among the works it cites.
Learning and transfer of modulated locomotor controllers
Heess, Nicolas, Wayne, Gregory, Tassa, Yuval, Lillicrap, Timothy P., Riedmiller, Martin A., and Silver, David · 2016
Later among the works it cites.
Vime: Variational information maximizing exploration
Houthooft, Rein, Chen, Xi, Duan, Yan, Schulman, John, Turck, Filip De, and Abbeel, Pieter · 2016
Later among the works it cites.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
Kulkarni, Tejas D, Narasimhan, Karthik, Saeedi, Ardavan, and Tenenbaum, Josh · 2016
Later among the works it cites.
Deep exploration via bootstrapped DQN
Osband, Ian, Blundell, Charles, and Alexander Pritzel, Benjamin Van Roy · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Rusu, Andrei A., Veness, Joel, Bellemare, Marc G., Graves, Alex, Riedmiller, Martin, Fidjeland, Andreas K., Ostrovski, Georg, Petersen, Stig, Beattie, Charles, Sadik, Amir, Antonoglou, Ioannis, King, Helen, Kumaran, Dharshan, Wierstra, Daan, Legg, Shane, and Hassabis, Demis · 2015
Cited alongside, same era.
Action-conditional video prediction using deep networks in atari games
Oh, Junhyuk, Guo, Xiaoxiao, Lee, Honglak, Lewis, Richard, and Singh, Satinder · 2015
Cited alongside, same era.
Trust region policy optimization
Schulman, John, Levine, Sergey, Moritz, Philipp, Jordan, Michael I., and Abbeel, Pieter · 2015
Cited alongside, same era.
Incentivizing exploration in reinforcement learning with deep predictive models
Stadie, Bradly C., Levine, Sergey, and Abbeel, Pieter · 2015
Cited alongside, same era.
Exploratory gradient boosting for reinforcement learning in complex domains
Abel, David, Agarwal, Alekh, Diaz, Fernando, Krishnamurthy, Akshay, and Schapire, Robert E · 2016
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
Bellemare, Marc G., Srinivasan, Sriram, Ostrovski, Georg, Schaul, Tom, Saxton, David, and Munos, Remi · 2016
Cited alongside, same era.
Improved techniques for training gans
Salimans, Tim, Goodfellow, Ian J., Zaremba, Wojciech, Cheung, Vicki, Radford, Alec, and Chen, Xi · 2016
Later among the works it cites.
#exploration: A study of count-based exploration for deep reinforcement learning
Tang, Haoran, Houthooft, Rein, Foote, Davis, Stooke, Adam, Chen, Xi, Duan, Yan, Schulman, John, Turck, Filip De, and Abbeel, Pieter · 2016
Later among the works it cites.
Surprise-based intrinsic motivation for deep reinforcement learning
Achiam, Joshua and Sastry, Shankar · 2017
Closest in time.
Stochastic neural networks for hierarchical reinforcement learning
Florensa, Carlos Campo, Duan, Yan, and Abbeel, Pieter · 2017
Closest in time.
Curiosity-driven exploration by self-supervised prediction
Pathak, Deepak, Agrawal, Pulkit, Efros, Alexei A., and Darrell, Trevor · 2017
Closest in time.