Fetching the paper…
Reading the bibliography…
Reinforcement Learning (RL), a subfield of Artificial Intelligence (AI), focuses on training agents to make decisions by interacting with their environment to maximize cumulative rewards.
P. Erd, “On a new law of large numbers,” J. Anal. Muth
1970
Earlier work this paper cites.
C. J. C. H. Watkins, “Learning from delayed rewards,” 1989
1989
Earlier work this paper cites.
R. S. Sutton, “Integrated architectures for learning, planning, and reacting based on approximating dynamic programming,” in Machine learning proceedings 1990
1990
Earlier work this paper cites.
Carnegie Mellon University, 1992
S. B. Thrun, Efficient exploration in reinforcement learning · 1992
Earlier work this paper cites.
P. Dayan and C. Watkins, “Q-learning,” Machine learning
1992
Earlier work this paper cites.
R. J. Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” Machine learning
1992
Earlier work this paper cites.
University of Cambridge, Department of Engineering Cambridge, UK, 1994
G. A. Rummery and M. Niranjan, On-line Q-learning using connectionist systems · 1994
Earlier work this paper cites.
L. P. Kaelbling, M. L. Littman, and A. W. Moore, “Reinforcement learning: A survey,” Journal of artificial intelligence research
1996
Earlier work this paper cites.
V. Konda and J. Tsitsiklis, “Actor-critic algorithms,” Advances in neural information processing systems
1999
Earlier work this paper cites.
R. S. Sutton, D. McAllester, S. Singh, and Y. Mansour, “Policy gradient methods for reinforcement learning with function approximation,” Advances in neural information processing systems
1999
Earlier work this paper cites.
L. Matignon, G. J. Laurent, and N. Le Fort-Piat, “Hysteretic q-learning: an algorithm for decentralized reinforcement learning in cooperative multi-agent teams,” in 2007 IEEE/RSJ International Conference on Intelligent Robots and Systems
2007
Earlier work this paper cites.
H. Van Seijen, H. Van Hasselt, S. Whiteson, and M. Wiering, “A theoretical and empirical analysis of expected sarsa,” in 2009 ieee symposium on adaptive dynamic programming and reinforcement learning
2009
Earlier work this paper cites.
H. Hasselt, “Double q-learning,” Advances in neural information processing systems
2010
Earlier work this paper cites.
A. Garivier and E. Moulines, “On upper-confidence bound policies for switching bandit problems,” in International conference on algorithmic learning theory
2011
Earlier work this paper cites.
D. P. Bertsekas, “Approximate policy iteration: A survey and some new methods,” Journal of Control Theory and Applications
2011
Earlier work this paper cites.
M. A. Wiering and M. Van Otterlo, “Reinforcement learning,” Adaptation, learning, and optimization
2012
Earlier work this paper cites.
M. Van Otterlo and M. Wiering, “Reinforcement learning and markov decision processes,” in Reinforcement learning: State-of-the-art
2012
Earlier work this paper cites.
Athena scientific, 2012
D. Bertsekas, Dynamic programming and optimal control: Volume I · 2012
Earlier work this paper cites.
I. Grondman, L. Busoniu, G. A. Lopes, and R. Babuska, “A survey of actor-critic reinforcement learning: Standard and natural policy gradients,” IEEE Transactions on Systems, Man, and Cybernetics, part C (applications and reviews)
2012
Earlier work this paper cites.
2013
Earlier work this paper cites.
J. Kober, J. A. Bagnell, and J. Peters, “Reinforcement learning in robotics: A survey,” The International Journal of Robotics Research
2013
Cited alongside, same era.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al
2015
Cited alongside, same era.
2015
Cited alongside, same era.
T. Schaul, “Prioritized experience replay,” arXiv preprint arXiv:1511.05952
2015
Cited alongside, same era.
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz, “Trust region policy optimization,” in International conference on machine learning
G. C. Lopes, M. Ferreira, A. da Silva Simões, and E. L. Colombini, “Intelligent control of a quadrotor with proximal policy optimization reinforcement learning,” in 2018 Latin American Robotic Symposium, 2018 Brazilian Symposium on Robotics (SBR) and 2018 Workshop on Robotics in Education (WRE)
2018
Later among the works it cites.
H. Wei, X. Liu, L. Mashayekhy, and K. Decker, “Mixed-autonomy traffic control with proximal policy optimization,” in 2019 IEEE Vehicular Networking Conference (VNC)
2019
Later among the works it cites.
E. Bøhn, E. M. Coates, S. Moe, and T. A. Johansen, “Deep reinforcement learning attitude control of fixed-wing uavs using proximal policy optimization,” in 2019 international conference on unmanned aircraft systems (ICUAS)
2019
Later among the works it cites.
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2015
Cited alongside, same era.
H. Van Hasselt, A. Guez, and D. Silver, “Deep reinforcement learning with double q-learning,” in Proceedings of the AAAI conference on artificial intelligence
2016
Cited alongside, same era.
Z. Wang, T. Schaul, M. Hessel, H. Hasselt, M. Lanctot, and N. Freitas, “Dueling network architectures for deep reinforcement learning,” in International conference on machine learning
2016
Cited alongside, same era.
V. Mnih, “Asynchronous methods for deep reinforcement learning,” arXiv preprint arXiv:1602.01783
2016
Cited alongside, same era.
2017
Cited alongside, same era.
Y. Li, “Deep reinforcement learning: An overview,” arXiv preprint arXiv:1701.07274
2017
Cited alongside, same era.
2017
Cited alongside, same era.
M. G. Bellemare, W. Dabney, and R. Munos, “A distributional perspective on reinforcement learning,” in International conference on machine learning
2017
Cited alongside, same era.
2020
Later among the works it cites.
B. Zhang, X. Lu, R. Diao, H. Li, T. Lan, D. Shi, and Z. Wang, “Real-time autonomous line flow control using proximal policy optimization,” in 2020 IEEE Power & Energy Society General Meeting (PESGM)
2020
Later among the works it cites.
Y. Guan, Y. Ren, S. E. Li, Q. Sun, L. Luo, and K. Li, “Centralized cooperation for connected and automated vehicles at intersections by proximal policy optimization,” IEEE Transactions on Vehicular Technology
2020
Later among the works it cites.
F. Ye, X. Cheng, P. Wang, C.-Y. Chan, and J. Zhang, “Automated lane change strategy using proximal policy optimization-based deep reinforcement learning,” in 2020 IEEE Intelligent Vehicles Symposium (IV)
2020
Later among the works it cites.
H.-n. Wang, N. Liu, Y.-y. Zhang, D.-w. Feng, F. Huang, D.-s. Li, and Y.-m. Zhang, “Deep reinforcement learning: a survey,” Frontiers of Information Technology & Electronic Engineering
2020
Later among the works it cites.
2021
Later among the works it cites.
J. Jin and Y. Xu, “Optimal policy characterization enhanced proximal policy optimization for multitask scheduling in cloud computing,” IEEE Internet of Things Journal
2021
Later among the works it cites.
L. Zhang, Y. Zhang, X. Zhao, and Z. Zou, “Image captioning via proximal policy optimization,” Image and Vision Computing
2021
Later among the works it cites.
2022
Later among the works it cites.
P. Ladosz, L. Weng, M. Kim, and H. Oh, “Exploration in deep reinforcement learning: A survey,” Information Fusion
2022
Later among the works it cites.
X. Wang, S. Wang, X. Liang, D. Zhao, J. Huang, X. Xu, B. Dai, and Q. Miao, “Deep reinforcement learning: A survey,” IEEE Transactions on Neural Networks and Learning Systems
2022
Later among the works it cites.
J. Ramírez, W. Yu, and A. Perrusquía, “Model-free reinforcement learning from expert demonstrations: a survey,” Artificial Intelligence Review
2022
Later among the works it cites.
S. E. Li, “Deep reinforcement learning,” in Reinforcement learning for sequential decision and optimal control
2023
Later among the works it cites.
T. M. Moerland, J. Broekens, A. Plaat, C. M. Jonker, et al
2023
Later among the works it cites.
F.-M. Luo, T. Xu, H. Lai, X.-H. Chen, W. Zhang, and Y. Yu, “A survey on model-based reinforcement learning,” Science China Information Sciences
2024
Closest in time.