Fetching the paper…
Reading the bibliography…
Discrete-continuous hybrid action space is a natural setting in many practical problems, such as robot control and game AI.
Visualizing data using t-SNE
L. V. D. Maaten and G. E. Hinton · 2008
Earlier work this paper cites.
Auto-encoding variational Bayes
D. P. Kingma and M. Welling · 2014
Earlier work this paper cites.
Deterministic policy gradient algorithms
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. A. Riedmiller · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. A. Riedmiller, A. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis · 2015
Earlier work this paper cites.
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M. I. Jordan, and P. Moritz · 2015
Earlier work this paper cites.
Deep reinforcement learning in parameterized action space
M. Hausknecht and P. Stone · 2016
Earlier work this paper cites.
Reinforcement learning with parameterized actions
W. Masson, P. Ranchod, and G. D. Konidaris · 2016
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
J. Schulman, P. Moritz, S. Levine, M. I. Jordan, and P. Abbeel · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. P. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis · 2016
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
S. Fujimoto, H. v. Hoof, and D. Meger · 2018
Earlier work this paper cites.
Hierarchical approaches for reinforcement learning in parameterized action space
E. Wei, D. Wicke, and S. Luke · 2018
Earlier work this paper cites.
J. Xiong, Q. Wang, Z. Yang, P. Sun, L. Han, Y. Zheng, H. Fu, T. Zhang, J. Liu, and H. Liu · 2018
Earlier work this paper cites.
A deep bayesian policy reuse approach against non-stationary agents
Y. Zheng, Z. Meng, J. Hao, Z. Zhang, T. Yang, and C. Fan · 2018
Cited alongside, same era.
Multi-pass q-networks for deep reinforcement learning with parameterised action spaces
C. J. Bester, S. D. James, and G. D. Konidaris · 2019
Cited alongside, same era.
Learning action representations for reinforcement learning
Y. Chandak, G. Theocharous, J. Kostas, S. M. Jordan, and P. S. Thomas · 2019
Cited alongside, same era.
Discrete and continuous action representation for practical RL in video games
O. Delalleau, M. Peter, E. Alonso, and A. Logut · 2019
Cited alongside, same era.
Hybrid actor-critic reinforcement learning in parameterized action space
Z. Fan, R. Su, W. Zhang, and Y. Yu · 2019
Cited alongside, same era.
Continuous multiagent control using collective behavior entropy for large-scale home energy management
J. Sun, Y. Zheng, J. Hao, Z. Meng, and Y. Liu · 2020
Later among the works it cites.
I 2 hrl: Interactive influence-based hierarchical reinforcement learning
R. Wang, R. Yu, B. An, and Z. Rabinovich · 2020
Later among the works it cites.
Dynamics-aware embeddings
W. F. Whitney, R. Agarwal, K. Cho, and A. Gupta · 2020
Later among the works it cites.
PLAS: latent action space for offline reinforcement learning
W. Zhou, S. Bajracharya, and D. Held · 2020
Later among the works it cites.
Mine your own view: Self-supervised learning through across-sample prediction
M. Azabou, M. G. Azar, R. Liu, C. H. Lin, E. C. Johnson, K. B. Nair, M. D., K. B. Hengen, W. G. Roncal, M. V., and E. Dyer · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep multi-agent reinforcement learning with discrete-continuous hybrid action spaces
H. Fu, H. Tang, J. Hao, Z. Lei, Y. Chen, and C. Fan · 2019
Cited alongside, same era.
Continuous-discrete reinforcement learning for hybrid control in robotics
M. Neunert, A. Abdolmaleki, M. Wulfmeier, T. Lampe, J. T. Springenberg, R. Hafner, F. Romano, J. Buchli, N. Heess, and M. A. Riedmiller · 2019
Cited alongside, same era.
Improving action branching for deep reinforcement learning with a multi-dimensional hybrid action space
L. Peng and Y. Tsuruoka · 2019
Cited alongside, same era.
Wuji: Automatic online combat game testing using evolutionary deep reinforcement learning
Y. Zheng, X. Xie, T. Su, L. Ma, J. Hao, Z. Meng, Y. Liu, R. Shen, Y. Chen, and C. Fan · 2019
Cited alongside, same era.
The impact of non-stationarity on generalisation in deep reinforcement learning
M. Igl, G. Farquhar, J. Luketina, W. Boehmer, and S. Whiteson · 2020
Cited alongside, same era.
Neural ordinary differential equation value networks for parametrized action spaces
S. Massaroli, M. Poli, S. Bakhtiyarov, A. Yamashita, H. Asama, and J. Park · 2020
Cited alongside, same era.
Data-efficient reinforcement learning with momentum predictive representations
M. Schwarzer, A. Anand, R. Goel, R. D. Hjelm, A. C. Courville, and P. Bachman · 2020
Cited alongside, same era.
Exploration-driven representation learning in reinforcement learning
A. Erraqabi, M. Zhao, M. C. Machado, Y. Bengio, S. Sukhbaatar, L. Denoyer, and A. Lazaric · 2021
Closest in time.
High-dimensional bayesian optimisation with variational autoencoders and deep metric learning
A. Grosnit, R. Tutunov, A. Maraval, R. Griffiths, A. Cowen-Rivers, L. Yang, L. Zhu, W. Lyu, Z. Chen, J. Wang, J. Peters, and H. Bou-Ammar · 2021
Closest in time.
Distributed reinforcement learning with self-play in parameterized action space
J. Ma, S. Yao, G. Chen, J. Song, and J. Ji · 2021
Closest in time.
Improving black-box optimization in VAE latent space using decoder uncertainty
P. Notin, J. M. Hernández-Lobato, and Y. Gal · 2021
Closest in time.
Decoupling representation learning from reinforcement learning
A. Stooke, K. Lee, P. Abbeel, and M. Laskin · 2021
Closest in time.
Foresee then evaluate: Decomposing value estimation with latent future prediction
H. Tang, Z. Meng, G. Chen, P. Chen, C. Chen, Y. Yang, L. Zhang, W. Liu, and J. Hao · 2021
Closest in time.
Reinforcement learning with prototypical representations
D. Yarats, R. Fergus, A. Lazaric, and L. Pinto · 2021
Closest in time.
Api: Boosting multi-agent reinforcement learning via agent-permutation-invariant networks
X. Hao, W. Wang, H. Mao, Y. Yang, D. Li, Y. Zheng, Z. Wang, and J. Hao · 2022
Closest in time.