Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) has shown remarkable success in solving complex decision-making and control tasks.
H. van Hasselt, “Double Q-learning,” in 23rd Advances in Neural Information Processing Systems (NeurIPS 2010)
2010
Earlier work this paper cites.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al
2015
Earlier work this paper cites.
J. Schulman, S. Levine, P. Abbeel, M. I. Jordan, and P. Moritz, “Trust region policy optimization,” in Proceedings of the 32nd International Conference on Machine Learning, (ICML 2015)
2015
Earlier work this paper cites.
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al
2016
Earlier work this paper cites.
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra, “Continuous control with deep reinforcement learning,” in 4th International Conference on Learning Representations (ICLR 2016)
2016
Earlier work this paper cites.
H. van Hasselt, A. Guez, and D. Silver, “Deep reinforcement learning with double Q-learning,” in Proceedings of the 30th Conference on Artificial Intelligence (AAAI 2016)
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, et al
2017
Earlier work this paper cites.
T. Haarnoja, H. Tang, P. Abbeel, and S. Levine, “Reinforcement learning with deep energy-based policies,” in Proceedings of the 34th International Conference on Machine Learning, (ICML 2017)
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
M. G. Bellemare, W. Dabney, and R. Munos, “A distributional perspective on reinforcement learning,” in Proceedings of the 34th International Conference on Machine Learning, (ICML 2017)
2017
Earlier work this paper cites.
S. Fujimoto, H. van Hoof, and D. Meger, “Addressing function approximation error in actor-critic methods,” in Proceedings of the 35th International Conference on Machine Learning (ICML 2018)
2018
Earlier work this paper cites.
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, “Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,” in Proceedings of the 35th International Conference on Machine Learning (ICML 2018)
2018
Earlier work this paper cites.
2018
Cited alongside, same era.
W. Dabney, M. Rowland, M. G. Bellemare, and R. Munos, “Distributional reinforcement learning with quantile regression,” in Proceedings of the 32nd Conference on Artificial Intelligence, (AAAI 2018)
2018
Cited alongside, same era.
W. Dabney, G. Ostrovski, D. Silver, and R. Munos, “Implicit quantile networks for distributional reinforcement learning,” in Proceedings of the 35th International Conference on Machine Learning (ICML 2018)
2018
Cited alongside, same era.
G. Barth-Maron, M. W. Hoffman, D. Budden, W. Dabney, D. Horgan, D. TB, A. Muldal, N. Heess, and T. P. Lillicrap, “Distributed distributional deterministic policy gradients,” in 6th International Conference on Learning Representations, (ICLR 2018)
2018
Cited alongside, same era.
Y. Chandak, S. Niekum, B. C. da Silva, E. G. Learned-Miller, E. Brunskill, and P. S. Thomas, “Universal off-policy evaluation,” in Neural Information Processing Systems
2021
Later among the works it cites.
G. Li, Y. Yang, S. Li, X. Qu, N. Lyu, and S. E. Li, “Decision making of autonomous vehicles in lane change scenarios: Deep reinforcement learning approaches with risk awareness,” Transportation research part C: emerging technologies
2022
Later among the works it cites.
J. Duan, W. Cao, Y. Zheng, and L. Zhao, “On the optimization landscape of dynamic output feedback linear quadratic control,” IEEE Transactions on Automatic Control
2023
Closest in time.
S. Li, Q. Tang, Y. Pang, X. Ma, and G. Wang, “Realistic actor-critic: A framework for balance between value overestimation and underestimation,” Frontiers in Neurorobotics
2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Yang, L. Zhao, Z. Lin, T. Qin, J. Bian, and T.-Y. Liu, “Fully parameterized quantile function for distributional reinforcement learning,” Advances in Neural Information Processing Systems
2019
Cited alongside, same era.
M. Rowland, R. Dadashi, S. Kumar, R. Munos, M. G. Bellemare, and W. Dabney, “Statistics and samples in distributional reinforcement learning,” in Proceedings of the 36th International Conference on Machine Learning, (ICML 2019)
2019
Cited alongside, same era.
B. Mavrin, H. Yao, L. Kong, K. Wu, and Y. Yu, “Distributional reinforcement learning for efficient exploration,” in Proceedings of the 36th International Conference on Machine Learning, (ICML 2019)
2019
Cited alongside, same era.
C. Tessler, G. Tennenholtz, and S. Mannor, “Distributional policy optimization: An alternative approach for continuous control,” Advances in Neural Information Processing Systems
2019
Cited alongside, same era.
2019
Cited alongside, same era.
W. Dabney, Z. Kurth-Nelson, N. Uchida, C. K. Starkweather, D. Hassabis, R. Munos, and M. Botvinick, “A distributional code for value in dopamine-based reinforcement learning,” Nature
2020
Cited alongside, same era.
Y. Ren, J. Duan, S. E. Li, Y. Guan, and Q. Sun, “Improving generalization of reinforcement learning with minimax distributional soft actor-critic,” in 23rd IEEE International Conference on Intelligent Transportation Systems (IEEE ITSC 2020)
2020
Cited alongside, same era.
D. Brown, S. Niekum, and M. Petrik, “Bayesian robust optimization for imitation learning,” in Advances in Neural Information Processing Systems
2020
Cited alongside, same era.
J. Li, J. Wang, S. E. Li, and K. Li, “Learning optimal robust control of connected vehicles in mixed traffic flow,” in 2023 62nd IEEE Conference on Decision and Control (CDC)
2023
Closest in time.
W. Wang, Y. Zhang, J. Gao, Y. Jiang, Y. Yang, Z. Zheng, W. Zou, J. Li, C. Zhang, W. Cao, et al
2023
Closest in time.
Springer Verlag, Singapore, 2023
S. E. Li, Reinforcement Learning for Sequential Decision and Optimal Control · 2023
Closest in time.
Adaptive computation and machine learning, Cambridge, Massachusetts London: The MIT Press, 2023
M. G. Bellemare, W. Dabney, and M. Rowland, Distributional reinforcement learning · 2023
Closest in time.
J. Duan, Y. Ren, F. Zhang, J. Li, S. E. Li, Y. Guan, and K. Li, “Encoding distributional soft actor-critic for autonomous driving in multi-lane scenarios [research frontier],” IEEE Computational Intelligence Magazine
2024
Closest in time.
X. He, J. Wu, Z. Huang, Z. Hu, J. Wang, A. Sangiovanni-Vincentelli, and C. Lv, “Fear-neuro-inspired reinforcement learning for safe autonomous driving,” IEEE Transactions on Pattern Analysis and Machine Intelligence
2024
Closest in time.
J. Zhang, S. Han, X. Xiong, S. Zhu, and S. Lü, “Explorer-actor-critic: Better actors for deep reinforcement learning,” Information Sciences
2024
Closest in time.
X. He, W. Huang, and C. Lv, “Toward trustworthy decision-making for autonomous vehicles: A robust reinforcement learning approach with safety guarantees,” Engineering
2024
Closest in time.
L. Xiao, Y. Lyu, F. Zhang, L. Chen, G. Yu, S. E. Li, F. Ma, and J. Duan, “Multi-style distributional soft actor-critic: Learning a unified policy for diverse control behaviors,” IEEE Transactions on Intelligent Vehicles
2024
Closest in time.