Fetching the paper…
Reading the bibliography…
In this paper, we present a new class of Markov decision processes (MDPs), called Tsallis MDPs, with Tsallis entropy maximization, which generalizes existing maximum entropy reinforcement learning (RL).
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. P. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu, “Asynchronous methods for deep reinforcement learning,” in Proceedings of the 33nd International Conference on Machine Learning, ICML 2016 , New York City, NY, USA, 2016, pp. 1928–1937
1937
Earlier work this paper cites.
C. Tsallis, “Possible generalization of boltzmann-gibbs statistics,” Journal of statistical physics , vol. 52, no. 1-2, pp. 479–487, 1988
1988
Earlier work this paper cites.
D. G. Luenberger, Optimization by vector space methods . John Wiley & Sons, 1997
1997
Earlier work this paper cites.
U. Syed, M. H. Bowling, and R. E. Schapire, “Apprenticeship learning using linear programming,” in Machine Learning, Proceedings of the Twenty-Fifth International Conference (ICML 2008) , Helsinki, Finland, 2008, pp. 1032–1039
2008
Earlier work this paper cites.
B. D. Ziebart, “Modeling purposeful adaptive behavior with the principle of maximum causal entropy,” Ph.D. dissertation, Carnegie Mellon University, 2010
2010
Earlier work this paper cites.
S. Amari and A. Ohara, “Geometry of q-exponential family of probability distributions,” Entropy , vol. 13, no. 6, pp. 1170–1185, 2011
2011
Earlier work this paper cites.
M. L. Puterman, Markov decision processes: discrete stochastic dynamic programming . John Wiley & Sons, 2014
2014
Earlier work this paper cites.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. A. Riedmiller, A. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis, “Human-level control through deep reinforcement learning,” Nature , vol. 518, no. 7540, pp. 529–533, 2015
2015
Earlier work this paper cites.
J. Schulman, S. Levine, P. Abbeel, M. I. Jordan, and P. Moritz, “Trust region policy optimization,” in Proceedings of the 32nd International Conference on Machine Learning, ICML 2015 , Lille, France, 2015, pp. 1889–1897
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Cited alongside, same era.
Y. Duan, X. Chen, R. Houthooft, J. Schulman, and P. Abbeel, “Benchmarking deep reinforcement learning for continuous control,” in Proceedings of the 33nd International Conference on Machine Learning, ICML 2016 . New York City, NY, USA: JMLR.org, 2016, pp. 1329–1338
2016
Cited alongside, same era.
S. Gu, T. P. Lillicrap, I. Sutskever, and S. Levine, “Continuous deep q-learning with model-based acceleration,” in Proceedings of the 33nd International Conference on Machine Learning, ICML 2016 , New York City, NY, USA, 2016, pp. 2829–2838
2016
Cited alongside, same era.
X. B. Peng, G. Berseth, K. Yin, and M. van de Panne, “Deeploco: dynamic locomotion skills using hierarchical deep reinforcement learning,” ACM Trans. Graph. , vol. 36, no. 4, pp. 41:1–41:13, 2017
2017
Later among the works it cites.
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, “Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,” in Proceedings of the 35th International Conference on Machine Learning, ICML 2018 , Stockholmsmässan, Stockholm, Sweden, 2018, pp. 1856–1865
2018
Later among the works it cites.
B. Dai, A. Shaw, L. Li, L. Xiao, N. He, Z. Liu, J. Chen, and L. Song, “SBEED: convergent reinforcement learning with nonlinear function approximation,” in Proceedings of the 35th International Conference on Machine Learning, (ICML 2018) , Stockholmsmässan, Stockholm, Sweden, 2018, pp. 1133–1142
2018
Later among the works it cites.
K. Lee, S. Choi, and S. Oh, “Sparse markov decision processes with causal sparse tsallis entropy regularization for reinforcement learning,” IEEE Robotics and Automation Letters , vol. 3, no. 3, pp. 1466–1473, 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
T. Haarnoja, H. Tang, P. Abbeel, and S. Levine, “Reinforcement learning with deep energy-based policies,” in Proceedings of the 34th International Conference on Machine Learning, ICML 2017 , Sydney, NSW, Australia, 2017, pp. 1352–1361
2017
Cited alongside, same era.
2017
Cited alongside, same era.
O. Nachum, M. Norouzi, K. Xu, and D. Schuurmans, “Bridging the gap between value and policy based reinforcement learning,” in Advances in Neural Information Processing Systems 30 NeurIPS 2017 , Long Beach, CA, USA, 2017, pp. 2772–2782
2017
Cited alongside, same era.
B. O’Donoghue, R. Munos, K. Kavukcuoglu, and V. Mnih, “PGQ: combining policy gradient and q-learning,” in Proceedings of the International Conference on Learning Representations (ICLR) , 2017. [Online]. Available: https://openreview.net/forum?id=B1kJ6H9ex
2017
Cited alongside, same era.
N. Cesa-Bianchi, C. Gentile, G. Neu, and G. Lugosi, “Boltzmann exploration done right,” in Advances in Neural Information Processing Systems 30 NeurIPS 2017 , Long Beach, CA, USA, 2017, pp. 6287–6296
2017
Cited alongside, same era.
2018
Later among the works it cites.
Y. Chow, O. Nachum, and M. Ghavamzadeh, “Path consistency learning in tsallis entropy regularized mdps,” in Proceedings of the 35th International Conference on Machine Learning, ICML 2018 , Stockholmsmässan, Stockholm, Sweden, 2018, pp. 978–987
2018
Later among the works it cites.
2018
Later among the works it cites.
S. Fujimoto, H. van Hoof, and D. Meger, “Addressing function approximation error in actor-critic methods,” in Proceedings of the 35th International Conference on Machine Learning, ICML 2018 , Stockholmsmässan, Stockholm, Sweden, 2018, pp. 1582–1591
2018
Later among the works it cites.
B. Eysenbach, A. Gupta, J. Ibarz, and S. Levine, “Diversity is all you need: Learning skills without a reward function,” in International Conference on Learning Representations ICLR , 2019. [Online]. Available: https://openreview.net/forum?id=SJx63jRqFm
2019
Closest in time.
J. Grau-Moya, F. Leibfried, and P. Vrancx, “Soft q-learning with mutual-information regularization,” in International Conference on Learning Representations , 2019. [Online]. Available: https://openreview.net/forum?id=HyEtjoCqFX
2019
Closest in time.