Fetching the paper…
Reading the bibliography…
We develop a new class of model-free deep reinforcement learning algorithms for data-driven, learning-based control.
A. Kong, “A note on importance sampling using standardized weights,” Tech. Rep. 348, Dept. Statist., Univ. Chicago, 1992
1992
Earlier work this paper cites.
S. Kakade and J. Langford, “Approximately optimal approximate reinforcement learning,” in Proc. 19th Int. Conf. Mach. Learn. , 2002, pp. 267–274
2002
Earlier work this paper cites.
A. B. Tsybakov, Introduction to Nonparametric Estimation . Springer New York, NY, 2009
2009
Earlier work this paper cites.
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz, “Trust region policy optimization,” in Proc. 32nd Int. Conf. Mach. Learn. , vol. 37, 2015, pp. 1889–1897
2015
Earlier work this paper cites.
Y. Duan, X. Chen, R. Houthooft, J. Schulman, and P. Abbeel, “Benchmarking deep reinforcement learning for continuous control,” in Proc. 33rd Int. Conf. Mach. Learn. , vol. 48, 2016, pp. 1329–1338
2016
Earlier work this paper cites.
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra, “Continuous control with deep reinforcement learning,” in Proc. 4th Int. Conf. Learn. Representations , 2016
2016
Earlier work this paper cites.
T. Schaul, J. Quan, I. Antonoglou, and D. Silver, “Prioritized experience replay,” in Proc. 4th Int. Conf. Learn. Representations , 2016
2016
Earlier work this paper cites.
J. Schulman, P. Moritz, S. Levine, M. I. Jordan, and P. Abbeel, “High-dimensional continuous control using generalized advantage estimation,” in Proc. 4th Int. Conf. Learn. Representations , 2016
2016
Earlier work this paper cites.
J. Achiam, D. Held, A. Tamar, and P. Abbeel, “Constrained policy optimization,” in Proc. 34th Int. Conf. Mach. Learn. , vol. 70, 2017, pp. 22–31
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
B. O’Donoghue, R. Munos, K. Kavukcuoglu, and V. Mnih, “Combining policy gradient and Q-learning,” in Proc. 5th Int. Conf. Learn. Representations , 2017
2017
Earlier work this paper cites.
S. Gu, T. Lillicrap, Z. Ghahramani, R. E. Turner, and S. Levine, “Q-Prop: Sample-efficient policy gradient with an off-policy critic,” in Proc. 5th Int. Conf. Learn. Representations , 2017
2017
Earlier work this paper cites.
S. Gu, T. Lillicrap, R. E. Turner, Z. Ghahramani, B. Schölkopf, and S. Levine, “Interpolated policy gradient: Merging on-policy and off-policy gradient estimation for deep reinforcement learning,” in Proc. Adv. Neural Inf. Process. Syst. , vol. 30, 2017
2017
Earlier work this paper cites.
Z. Wang, V. Bapst, N. Heess, V. Mnih, R. Munos, K. Kavukcuoglu, and N. de Freitas, “Sample efficient actor-critic with experience replay,” in Proc. 5th Int. Conf. Learn. Representations , 2017
2017
Earlier work this paper cites.
L. Buşoniu, T. de Bruin, D. Tolić, J. Kober, and I. Palunko, “Reinforcement learning for control: Performance, stability, and deep approximators,” Annu. Rev. Control , vol. 46, pp. 8–28, 2018
2018
Earlier work this paper cites.
T. Kurutach, I. Clavera, Y. Duan, A. Tamar, and P. Abbeel, “Model-ensemble trust-region policy optimization,” in Proc. 6th Int. Conf. Learn. Representations , 2018
2018
Earlier work this paper cites.
P. Henderson, R. Islam, P. Bachman, J. Pineau, D. Precup, and D. Meger, “Deep reinforcement learning that matters,” in Proc. AAAI Conf. Artif. Intell. , vol. 32, no. 1, 2018, pp. 3207–3214
2018
Cited alongside, same era.
S. Fujimoto, H. van Hoof, and D. Meger, “Addressing function approximation error in actor-critic methods,” in Proc. 35th Int. Conf. Mach. Learn. , vol. 80, 2018, pp. 1587–1596
2018
Cited alongside, same era.
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, “Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,” in Proc. 35th Int. Conf. Mach. Learn. , vol. 80, 2018, pp. 1861–1870
2018
Cited alongside, same era.
A. Abdolmaleki, J. T. Springenberg, Y. Tassa, R. Munos, N. Heess, and M. Riedmiller, “Maximum a posteriori policy optimisation,” in Proc. 6th Int. Conf. Learn. Representations , 2018
2018
Cited alongside, same era.
R. Kidambi, A. Rajeswaran, P. Netrapalli, and T. Joachims, “MOReL: Model-based offline reinforcement learning,” in Proc. Adv. Neural Inf. Process. Syst. , vol. 33, 2020
2020
Later among the works it cites.
H. F. Song, A. Abdolmaleki, J. T. Springenberg, A. Clark, H. Soyer, J. W. Rae, S. Noury, A. Ahuja, S. Liu, D. Tirumala, N. Heess, D. Belov, M. Riedmiller, and M. M. Botvinick, “V-MPO: On-policy maximum a posteriori policy optimization for discrete and continuous control,” in Proc. 8th Int. Conf. Learn. Representations , 2020
2020
Later among the works it cites.
S. Tunyasuvunakool, A. Muldal, Y. Doron, S. Liu, S. Bohez, J. Merel, T. Erez, T. Lillicrap, N. Heess, and Y. Tassa, “dm_control: Software and tasks for continuous control,” Softw. Impacts , vol. 6, p. 100022, 2020
2020
Later among the works it cites.
L. Engstrom, A. Ilyas, S. Santurkar, D. Tsipras, F. Janoos, L. Rudolph, and A. Madry, “Implementation matters in deep RL: A case study on PPO and TRPO,” in Proc. 8th Int. Conf. Learn. Representations , 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. de Bruin, J. Kober, K. Tuyls, and R. Babuška, “Experience selection in deep reinforcement learning for control,” J. Mach. Learn. Res. , vol. 19, no. 9, pp. 1–56, 2018
2018
Cited alongside, same era.
Z.-W. Hong, T.-Y. Shann, S.-Y. Su, Y.-H. Chang, T.-J. Fu, and C.-Y. Lee, “Diversity-driven exploration strategy for deep reinforcement learning,” in Proc. Adv. Neural Inf. Process. Syst. , vol. 31, 2018
2018
Cited alongside, same era.
L. Espeholt, H. Soyer, R. Munos, K. Simonyan, V. Mnih, T. Ward, Y. Doron, V. Firoiu, T. Harley, I. Dunning, S. Legg, and K. Kavukcuoglu, “IMPALA: Scalable distributed deep-RL with importance weighted actor-learner architectures,” in Proc. 35th Int. Conf. Mach. Learn. , vol. 80, 2018, pp. 1407–1416
2018
Cited alongside, same era.
B. Recht, “A tour of reinforcement learning: The view from continuous control,” Annu. Rev. Control, Robot., Auton. Syst. , vol. 2, no. 1, pp. 253–279, 2019
2019
Cited alongside, same era.
M. Janner, J. Fu, M. Zhang, and S. Levine, “When to trust your model: Model-based policy optimization,” in Proc. Adv. Neural Inf. Process. Syst. , vol. 32, 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
J. F. Fisac, A. K. Akametalu, M. N. Zeilinger, S. Kaynama, J. Gillula, and C. J. Tomlin, “A general safety framework for learning-based control in uncertain robotic systems,” IEEE Trans. Autom. Control , vol. 64, no. 7, pp. 2737–2752, 2019
2019
Cited alongside, same era.
Y. Wang, H. He, X. Tan, and Y. Gan, “Trust region-guided proximal policy optimization,” in Proc. Adv. Neural Inf. Process. Syst. , vol. 32, 2019
2019
Cited alongside, same era.
2020
Later among the works it cites.
Y. Wang, H. He, and X. Tan, “Truly proximal policy optimization,” in Proc. 35th Uncertainty Artif. Intell. Conf. , vol. 115, 2020, pp. 113–122
2020
Later among the works it cites.
R. Fakoor, P. Chaudhari, and A. J. Smola, “P3O: Policy-on policy-off policy optimization,” in Proc. 35th Uncertainty Artif. Intell. Conf. , vol. 115, 2020, pp. 1017–1027
2020
Later among the works it cites.
C. Wang, Y. Wu, Q. Vuong, and K. Ross, “Striving for simplicity and performance in off-policy DRL: Output normalization and non-uniform sampling,” in Proc. 37th Int. Conf. Mach. Learn. , vol. 119, 2020, pp. 10 070–10 080
2020
Later among the works it cites.
J. Queeney, I. C. Paschalidis, and C. G. Cassandras, “Generalized proximal policy optimization with sample reuse,” in Proc. Adv. Neural Inf. Process. Syst. , vol. 34, 2021
2021
Later among the works it cites.
M. Andrychowicz, A. Raichuk, P. Stańczyk, M. Orsini, S. Girgin, R. Marinier, L. Hussenot, M. Geist, O. Pietquin, M. Michalski, S. Gelly, and O. Bachem, “What matters for on-policy deep actor-critic methods? A large-scale study,” in Proc. 9th Int. Conf. Learn. Representations , 2021
2021
Later among the works it cites.
J. Queeney, I. C. Paschalidis, and C. G. Cassandras, “Uncertainty-aware policy optimization: A robust, adaptive trust region approach,” in Proc. AAAI Conf. Artif. Intell. , vol. 35, no. 11, 2021, pp. 9377–9385
2021
Later among the works it cites.
K. P. Wabersich, L. Hewing, A. Carron, and M. N. Zeilinger, “Probabilistic model predictive safety certification for learning-based control,” IEEE Trans. Autom. Control , vol. 67, no. 1, pp. 176–188, 2022
2022
Closest in time.
L. Brunke, M. Greeff, A. W. Hall, Z. Yuan, S. Zhou, J. Panerati, and A. P. Schoellig, “Safe learning in robotics: From learning-based control to safe reinforcement learning,” Annu. Rev. Control, Robot., Auton. Syst. , vol. 5, no. 1, pp. 411–444, 2022
2022
Closest in time.
Y. Cheng, L. Huang, and X. Wang, “Authentic boundary proximal policy optimization,” IEEE Trans. Cybern. , vol. 52, no. 9, pp. 9428–9438, 2022
2022
Closest in time.
W. Meng, Q. Zheng, Y. Shi, and G. Pan, “An off-policy trust region policy optimization method with monotonic improvement guarantee for deep reinforcement learning,” IEEE Trans. Neural Netw. Learn. Syst. , vol. 33, no. 5, pp. 2223–2235, 2022
2022
Closest in time.
S. Paternain, M. Calvo-Fullana, L. F. O. Chamon, and A. Ribeiro, “Safe policies for reinforcement learning via primal-dual methods,” IEEE Trans. Autom. Control , vol. 68, no. 3, pp. 1321–1336, 2023
2023
Closest in time.
L. Grossman and B. Plancher, “Just round: Quantized observation spaces enable memory efficient learning of dynamic locomotion,” in IEEE Int. Conf. Robot. Automat. (ICRA) , 2023, pp. 3002–3007
2023
Closest in time.