Fetching the paper…
Reading the bibliography…
In this paper, we propose a novel Reinforcement Learning (RL) framework for problems with continuous action spaces: Action Quantization from Demonstrations (AQuaDem).
Differential equations with a discontinuous forcing term
Bushaw, D. W · 1953
Earlier work this paper cites.
On the “bang-bang” control problem
Bellman, R., Glicksberg, I., and Gross, O · 1956
Earlier work this paper cites.
A markovian decision process
Bellman, R · 1957
Earlier work this paper cites.
Efficient training of artificial neural networks for autonomous navigation
Pomerleau, D. A · 1991
Earlier work this paper cites.
Q-learning
Watkins, C. J. and Dayan, P · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Mixture density networks
Bishop, C. M · 1994
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Ziebart, B. D., Maas, A. L., Bagnell, J. A., and Dey, A. K · 1994
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Ng, A. Y., Harada, D., and Russell, S · 1999
Earlier work this paper cites.
Dynamic programming and optimal control
Bertsekas, D. P · 2000
Earlier work this paper cites.
Actor-critic algorithms
Konda, V. R. and Tsitsiklis, J. N · 2000
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Ng, A. Y., Russell, S. J., et al · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y · 2000
Earlier work this paper cites.
Incremental learning of gestures by imitation in a humanoid robot
Calinon, S. and Billard, A · 2007
Earlier work this paper cites.
Confidence-based policy learning from demonstration using gaussian mixture models
Chernova, S. and Veloso, M · 2007
Earlier work this paper cites.
Learning and generalization of motor skills by learning from demonstration
Pastor, P., Hoffmann, H., Asfour, T., and Schaal, S · 2009
Earlier work this paper cites.
A survey of monte carlo tree search methods
Browne, C. B., Powley, E., Whitehouse, D., Lucas, S. M., Cowling, P. I., Rohlfshagen, P., Tavener, S., Perez, D., Samothrakis, S., and Colton, S · 2012
Earlier work this paper cites.
Novelty or surprise?
Barto, A., Mirolli, M., and Baldassarre, G · 2013
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L · 2014
Earlier work this paper cites.
Deep reinforcement learning in large discrete action spaces
Dulac-Arnold, G., Evans, R., van Hasselt, H., Sunehag, P., Lillicrap, T., Hunt, J., Mann, T., Weber, T., Degris, T., and Coppin, B · 2015
Earlier work this paper cites.
Towards learning hierarchical skills for multi-phase manipulation tasks
Kroemer, O., Daniel, C., Neumann, G., Van Hoof, H., and Peters, J · 2015
Earlier work this paper cites.
Learning movement primitive attractor goals and sequential skills from kinesthetic demonstrations
Manschitz, S., Kober, J., Gienger, M., and Peters, J · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Earlier work this paper cites.
Continuous deep q-learning with model-based acceleration
Gu, S., Lillicrap, T., Sutskever, I., and Levine, S · 2016
Earlier work this paper cites.
Generative adversarial imitation learning
Ho, J. and Ermon, S · 2016
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2016
Earlier work this paper cites.
Prioritized experience replay
Schaul, T., Quan, J., Antonoglou, I., and Silver, D · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Cited alongside, same era.
Deep reinforcement learning with double q-learning
Van Hasselt, H., Guez, A., and Silver, D · 2016
Cited alongside, same era.
Input convex neural networks
Amos, B., Xu, L., and Kolter, J. Z · 2017
Cited alongside, same era.
Ddco: Discovery of deep continuous options for robot learning from demonstrations
Krishnan, S., Fox, R., Stoica, I., and Goldberg, K · 2017
Cited alongside, same era.
Discrete sequential prediction of continuous actions for deep rl
Metz, L., Ibarz, J., Jaitly, N., and Davidson, J · 2017
Cited alongside, same era.
Learning complex dexterous manipulation with deep reinforcement learning and demonstrations
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., et al · 2019
Later among the works it cites.
Random expert distillation: Imitation learning via expert policy support estimation
Wang, R., Ciliberto, C., Amadori, P. V., and Demiris, Y · 2019
Later among the works it cites.
Behavior regularized offline reinforcement learning
Wu, Y., Tucker, G., and Nachum, O · 2019
Later among the works it cites.
Learning dexterous in-hand manipulation
Andrychowicz, O. M., Baker, B., Chociej, M., Jozefowicz, R., McGrew, B., Pachocki, J., Petron, A., Plappert, M., Powell, G., Ray, A., et al · 2020
Later among the works it cites.
Acme: A research framework for distributed reinforcement learning
Hoffman, M., Shahriari, B., Aslanides, J., Barth-Maron, G., Behbahani, F., Norman, T., Abdolmaleki, A., Cassirer, A., Yang, F., Baumli, K., et al · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rajeswaran, A., Kumar, V., Gupta, A., Vezzani, G., Schulman, J., Todorov, E., and Levine, S · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
# exploration: A study of count-based exploration for deep reinforcement learning
Tang, H., Houthooft, R., Foote, D., Stooke, A., Chen, X., Duan, Y., Schulman, J., De Turck, F., and Abbeel, P · 2017
Cited alongside, same era.
Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards
Vecerik, M., Hester, T., Scholz, J., Wang, F., Pietquin, O., Piot, B., Heess, N., Rothörl, T., Lampe, T., and Riedmiller, M · 2017
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
Fujimoto, S., Hoof, H., and Meger, D · 2018
Cited alongside, same era.
Deep q-learning from demonstrations
Hester, T., Vecerik, M., Pietquin, O., Lanctot, M., Schaul, T., Piot, B., Horgan, D., Quan, J., Sendonaris, A., Osband, I., et al · 2018
Cited alongside, same era.
Qt-opt: Scalable deep reinforcement learning for vision-based robotic manipulation
Kalashnikov, D., Irpan, A., Pastor, P., Ibarz, J., Herzog, A., Jang, E., Quillen, D., Holly, E., Kalakrishnan, M., Vanhoucke, V., et al · 2018
Cited alongside, same era.
Later among the works it cites.
Strictly batch imitation learning by energy-based distribution matching
Jarrett, D., Bica, I., and van der Schaar, M · 2020
Later among the works it cites.
Conservative q-learning for offline reinforcement learning
Kumar, A., Zhou, A., Tucker, G., and Levine, S · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., Kumar, A., Tucker, G., and Fu, J · 2020
Later among the works it cites.
Learning latent plans from play
Lynch, C., Khansari, M., Xiao, T., Kumar, V., Tompson, J., Levine, S., and Sermanet, P · 2020
Later among the works it cites.
Continuous-discrete reinforcement learning for hybrid control in robotics
Neunert, M., Abdolmaleki, A., Wulfmeier, M., Lampe, T., Springenberg, T., Hafner, R., Romano, F., Buchli, J., Heess, N., and Riedmiller, M · 2020
Later among the works it cites.
Caql: Continuous action q-learning
Ryu, M., Chow, Y., Anderson, R., Tjandraatmadja, C., and Boutilier, C · 2020
Later among the works it cites.
Inferring DQN structure for high-dimensional continuous control
Sakryukin, A., Raissi, C., and Kankanhalli, M · 2020
Later among the works it cites.
Parrot: Data-driven behavioral priors for reinforcement learning
Singh, A., Liu, H., Zhou, G., Yu, A., Rhinehart, N., and Levine, S · 2020
Later among the works it cites.
Discretizing continuous action space for on-policy optimization
Tang, Y. and Agrawal, S · 2020
Later among the works it cites.
dm_control: Software and tasks for continuous control, 2020
Tassa, Y., Tunyasuvunakool, S., Muldal, A., Doron, Y., Liu, S., Bohez, S., Merel, J., Erez, T., Lillicrap, T., and Heess, N · 2020
Later among the works it cites.
Munchausen reinforcement learning
Vieillard, N., Pietquin, O., and Geist, M · 2020
Later among the works it cites.
Deep radial-basis value functions for continuous control
Asadi, K., Parikh, N., Parr, R. E., Konidaris, G. D., and Littman, M. L · 2021
Closest in time.
Primal wasserstein imitation learning
Dadashi, R., Hussenot, L., Geist, M., and Pietquin, O · 2021
Closest in time.
Pot: Python optimal transport
Flamary, R., Courty, N., Gramfort, A., Alaya, M. Z., Boisbunon, A., Chambon, S., Chapel, L., Corenflos, A., Fatras, K., Fournier, N., Gautheron, L., Gayraud, N. T., Janati, H., Rakotomamonjy, A., Redko, I., Rolet, A., Schutz, A., Seguy, V., Sutherland, D. J., Tavenard, R., Tong, A., and Vayer, T · 2021
Closest in time.
Florence, P., Lynch, C., Zeng, A., Ramirez, O., Wahid, A., Downs, L., Wong, A., Lee, J., Mordatch, I., and Tompson, J · 2021
Closest in time.
A minimalist approach to offline reinforcement learning
Fujimoto, S. and Gu, S. S · 2021
Closest in time.
Hyperparameter selection for imitation learning
Hussenot, L., Andrychowicz, M., Vincent, D., Dadashi, R., Raichuk, A., Ramos, S., Momchev, N., Girgin, S., Marinier, R., Stafiniak, L., et al · 2021
Closest in time.
Robodesk: A multi-task reinforcement learning benchmark
Kannan, H., Hafner, D., Finn, C., and Erhan, D · 2021
Closest in time.
Offline reinforcement learning with implicit q-learning
Kostrikov, I., Nair, A., and Levine, S · 2021
Closest in time.
What matters in learning from offline human demonstrations for robot manipulation
Mandlekar, A., Xu, D., Wong, J., Nasiriany, S., Wang, C., Kulkarni, R., Fei-Fei, L., Savarese, S., Zhu, Y., and Martín-Martín, R · 2021
Closest in time.
What matters for adversarial imitation learning?
Orsini, M., Raichuk, A., Hussenot, L., Vincent, D., Dadashi, R., Girgin, S., Geist, M., Bachem, O., Pietquin, O., and Andrychowicz, M · 2021
Closest in time.
Is bang-bang control all you need? solving continuous control with bernoulli policies
Seyde, T., Gilitschenski, I., Schwarting, W., Stellato, B., Riedmiller, M., Wulfmeier, M., and Rus, D · 2021
Closest in time.
Learning to represent action values as a hypergraph on the action vertices
Tavakoli, A., Fatemi, M., and Kormushev, P · 2021
Closest in time.
Implicitly regularized rl with implicit q-values
Vieillard, N., Andrychowicz, M., Raichuk, A., Pietquin, O., and Geist, M · 2021
Closest in time.