Fetching the paper…
Reading the bibliography…
Deep reinforcement learning algorithms that estimate state and state-action value functions have been shown to be effective in a variety of challenging domains, including learning control strategies from raw image pixels.
Reinforcement Learning Algorithm for Partially Observable Markov Decision Problems
Jaakkola, T., Singh, S. P., and Jordan, M. I · 1994
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Littman, M. L · 1994
Earlier work this paper cites.
Efficient Exploration in Reinforcement Learning with Hidden State
McCallum, A. K · 1997
Earlier work this paper cites.
The Asymptotic Convergence-Rate of Q-learning
Szepesvári, C · 1998
Earlier work this paper cites.
A Simple Adaptive Procedure Leading to Correlated Equilibrium
Hart, S. and Mas-Colell, A · 2000
Earlier work this paper cites.
Potential-Based Algorithms in On-Line Prediction and Game Theory
Cesa-Bianchi, N. and Lugosi, G · 2003
Earlier work this paper cites.
No-regret Algorithms for Online Convex Programs
Gordon, G. J · 2007
Earlier work this paper cites.
Reinforcement Learning in Continuous Action Spaces
van Hasselt, H. and Wiering, M. A · 2007
Earlier work this paper cites.
Regret Minimization in Games with Incomplete Information
Zinkevich, M., Johanson, M., Bowling, M. H., and Piccione, C · 2007
Earlier work this paper cites.
A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning
Ross, S., Gordon, G. J., and Bagnell, J. A · 2011
Earlier work this paper cites.
Lossy Stochastic Game Abstraction with Bounds
Sandholm, T. and Singh, S · 2012
Earlier work this paper cites.
The Arcade Learning Environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Playing Atari with Deep Reinforcement Learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Earlier work this paper cites.
Extensive-Form Game Abstraction With Bounds
Kroer, C. and Sandholm, T · 2014
Earlier work this paper cites.
Reinforcement and Imitation Learning via Interactive No-Regret Learning
Ross, S. and Bagnell, J. A · 2014
Earlier work this paper cites.
Solving Large Imperfect Information Games Using CFR+
Tammelin, O · 2014
Earlier work this paper cites.
Heads-up limit hold’em poker is solved
Bowling, M., Burch, N., Johanson, M., and Tammelin, O · 2015
Cited alongside, same era.
Policy Gradient Reinforcement Learning Without Regret
Dick, T · 2015
Cited alongside, same era.
Memory-based control with recurrent neural networks
Heess, N., Hunt, J. J., Lillicrap, T. P., and Silver, D · 2015
Cited alongside, same era.
Adam: A Method for Stochastic Optimization
Kingma, D. P. and Ba, J. L · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Cited alongside, same era.
Action-Conditional Video Prediction using Deep Networks in Atari Games
Oh, J., Guo, X., Lee, H., Lewis, R. L., and Singh, S · 2015
Terrain-Adaptive Locomotion Skills Using Deep Reinforcement Learning
Peng, X. B., Berseth, G., and van de Penne, M · 2016
Later among the works it cites.
High-Dimensional Continuous Control Using Generalized Advantage Estimation
Schulman, J., Moritz, P., Levine, S., Jordan, M. I., and Abbeel, P · 2016
Later among the works it cites.
Deep Reinforcement Learning and Double Q-Learning
van Hasselt, H., Guez, A., and Silver, D · 2016
Later among the works it cites.
Dueling Network Architectures for Deep Reinforcement Learning
Wang, Z., Schaul, T., Hessel, M., van Hasselt, H., Lanctot, M., and de Freitas, N · 2016
Later among the works it cites.
Averaged-DQN: Variance Reduction and Stabilization for Deep Reinforcement Learning
Anschel, O., Baram, N., and Shimkin, N · 2017
Closest in time.
A Distributional Perspective on Reinforcement Learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Cited alongside, same era.
Solving Heads-up Limit Texas Hold’em
Tammelin, O., Burch, N., Johanson, M., and Bowling, M · 2015
Cited alongside, same era.
Solving Games with Functional Regret Estimation
Waugh, K., Morrill, D., Bagnell, J. A., and Bowling, M · 2015
Cited alongside, same era.
Increasing the Action Gap: New Operators for Reinforcement Learning
Bellemare, M. G., Ostrovski, G., Guez, A., Thomas, P. S., and Munos, R · 2016
Cited alongside, same era.
The Malmo Platform for Artificial Intelligence Experimentation
Johnson, M., Hofmann, K., Hutton, T., and Bignell, D · 2016
Cited alongside, same era.
ViZDoom: A Doom-based AI Research Platform for Visual Reinforcement Learning
Kempka, M., Wydmuch, M., Runc, G., Toczek, J., and Jaśkowski, W · 2016
Cited alongside, same era.
Bellemare, M. G., Dabney, W., and Munos, R · 2017
Closest in time.
Learning to Act by Predicting the Future
Dosovitskiy, A. and Koltun, V · 2017
Closest in time.
Gu, S., Lillicrap, T., Ghahramani, Z., Turner, R. E., Schölkopf, B., and Levine, S · 2017
Closest in time.
Reinforcement Learning with Deep Energy-Based Policies
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S · 2017
Closest in time.
Deep Recurrent Q-Learning for Partially Observable MDPs
Hausknecht, M. and Stone, P · 2017
Closest in time.
Teacher-Student Curriculum Learning
Matiisen, T., Oliver, A., Cohen, T., and Schulman, J · 2017
Closest in time.
Bridging the Gap Between Value and Policy Based Reinforcement Learning
Nachum, O., Norouzi, M., Xu, K., and Schuurmans, D · 2017
Closest in time.
Combining policy gradient and Q-learning
O’Donoghue, B., Munos, R., Kavukcuoglu, K., and Mnih, V · 2017
Closest in time.
Equivalence Between Policy Gradients and Soft Q-Learning
Schulman, J., Chen, X., and Abbeel, P · 2017
Closest in time.
Sample Efficient Actor-Critic with Experience Replay
Wang, Z., Bapst, V., Heess, N., Mnih, V., Munos, R., Kavukcuoglu, K., and de Freitas, N · 2017
Closest in time.
Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Closest in time.