Fetching the paper…
Reading the bibliography…
We introduce Wasserstein Policy Optimization (WPO), an actor-critic algorithm for reinforcement learning in continuous action spaces.
Wasserstein robust reinforcement learning, 2019
Abdullah, M. A., Ren, H., Ammar, H. B., Milenkovic, V., Luo, R., Zhang, M., and Wang, J · 1907
Earlier work this paper cites.
Dynamic Programming and Markov Processes
Howard, R. A · 1960
Earlier work this paper cites.
Beyond regression: New tools for prediction and analysis in the behavioral sciences
Werbos, P · 1974
Earlier work this paper cites.
Neuronlike adaptive elements that can solve difficult learning control problems
Barto, A. G., Sutton, R. S., and Anderson, C. W · 1983
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Adaptive critic designs
Prokhorov, D. V. and Wunsch, D. C · 1997
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D., Singh, S., and Mansour, Y · 1999
Earlier work this paper cites.
A computational fluid mechanics solution to the monge-kantorovich mass transfer problem
Benamou, J.-D. and Brenier, Y · 2000
Earlier work this paper cites.
A natural policy gradient
Kakade, S. M · 2001
Earlier work this paper cites.
Metrics for finite markov decision processes
Ferns, N., Panangaden, P., and Precup, D · 2004
Earlier work this paper cites.
Reinforcement learning in continuous action spaces
van Hasselt, H. and Wiering, M. A · 2007
Earlier work this paper cites.
Gradient flows: in metric spaces and in the space of probability measures
Ambrosio, L., Gigli, N., and Savaré, G · 2008
Earlier work this paper cites.
Double q-learning
van Hasselt, H · 2010
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
Revisiting natural gradient for deep networks
Pascanu, R. and Bengio, Y · 2013
Earlier work this paper cites.
Deterministic policy gradient algorithms
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M · 2014
Earlier work this paper cites.
Learning continuous control policies by stochastic value gradients
Heess, N., Wayne, G., Silver, D., Lillicrap, T., Erez, T., and Tassa, Y · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T · 2015
Cited alongside, same era.
Optimizing neural networks with Kronecker-factored approximate curvature
Martens, J. and Grosse, R · 2015
Cited alongside, same era.
Fast and accurate deep network learning by exponential linear units (elus). arxiv 2015
Clevert, D.-A., Unterthiner, T., and Hochreiter, S · 2016
Cited alongside, same era.
A distributional perspective on reinforcement learning
Bellemare, M. G., Dabney, W., and Munos, R · 2017
Cited alongside, same era.
Categorical reparameterization with Gumbel-softmax
Jang, E., Gu, S., and Poole, B · 2017
Cited alongside, same era.
The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables
V-MPO: On-policy maximum a posteriori policy optimization for discrete and continuous control
Song, H. F., Abdolmaleki, A., Springenberg, J. T., Clark, A., Soyer, H., Rae, J. W., Noury, S., Ahuja, A., Liu, S., Tirumala, D., et al · 2019
Later among the works it cites.
Acme: A research framework for distributed reinforcement learning
Hoffman, M. W., Shahriari, B., Aslanides, J., Barth-Maron, G., Momchev, N., Sinopalnikov, D., Stańczyk, P., Ramos, S., Raichuk, A., Vincent, D., et al · 2020
Later among the works it cites.
Efficient Wasserstein natural gradients for reinforcement learning
Moskovitz, T., Arbel, M., Huszar, F., and Gretton, A · 2020
Later among the works it cites.
Learning to score behaviors for guided policy optimization
Pacchiano, A., Parker-Holder, J., Tang, Y., Choromanski, K., Choromanska, A., and Jordan, M · 2020
Later among the works it cites.
dm_control: Software and tasks for continuous control
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Maddison, C. J., Mnih, A., and Teh, Y. W · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
Maximum a posteriori policy optimisation
Abdolmaleki, A., Springenberg, J. T., Tassa, Y., Munos, R., Heess, N., and Riedmiller, M · 2018
Cited alongside, same era.
Distributed distributional deterministic policy gradients
Barth-Maron, G., Hoffman, M. W., Budden, D., Dabney, W., Horgan, D., Tb, D., Muldal, A., Heess, N., and Lillicrap, T · 2018
Cited alongside, same era.
Neural ordinary differential equations
Chen, R. T., Rubanova, Y., Bettencourt, J., and Duvenaud, D. K · 2018
Cited alongside, same era.
Sigmoid-weighted linear units for neural network function approximation in reinforcement learning
Elfwing, S., Uchibe, E., and Doya, K · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Cited alongside, same era.
Tunyasuvunakool, S., Muldal, A., Doron, Y., Liu, S., Bohez, S., Merel, J., Erez, T., Lillicrap, T., Heess, N., and Tassa, Y · 2020
Later among the works it cites.
On the theory of policy gradient methods: Optimality, approximation, and distribution shift
Agarwal, A., Kakade, S. M., Lee, J. D., and Mahajan, G · 2021
Later among the works it cites.
Development of free-boundary equilibrium and transport solvers for simulation and real-time interpretation of tokamak experiments
Carpanese, F · 2021
Later among the works it cites.
Wasserstein unsupervised reinforcement learning, 2021
He, S., Jiang, Y., Zhang, H., Shao, J., and Ji, X · 2021
Later among the works it cites.
Castro, P. S., Kastner, T., Panangaden, P., and Rowland, M · 2022
Later among the works it cites.
Magnetic control of tokamak plasmas through deep reinforcement learning
Degrave, J., Felici, F., Buchli, J., Neunert, M., Tracey, B., Carpanese, F., Ewalds, T., Hafner, R., Abdolmaleki, A., de Las Casas, D., et al · 2022
Later among the works it cites.
The 37 implementation details of proximal policy optimization
Huang, S., Dossa, R. F. J., Raffin, A., Kanervisto, A., and Wang, W · 2022
Later among the works it cites.
Wasserstein actor-critic: Directed exploration via optimism for continuous-actions control, 2023
Likmeta, A., Sacco, M., Metelli, A. M., and Restelli, M · 2023
Later among the works it cites.
Wasserstein quantum Monte Carlo: a novel approach for solving the quantum many-body schrödinger equation
Neklyudov, K., Nys, J., Thiede, L., Carrasquilla, J., Liu, Q., Welling, M., and Makhzani, A · 2023
Later among the works it cites.
Experimental research on the TCV tokamak
Duval, B., Abdolmaleki, A., Agostini, M., Ajay, C., Alberti, S., Alessi, E., Anastasiou, G., Andrèbe, Y., Apruzzese, G., Auriemma, F., et al · 2024
Later among the works it cites.
Cale: Continuous arcade learning environment
Farebrother, J. and Castro, P. S · 2024
Later among the works it cites.
Learning agile soccer skills for a bipedal robot with deep reinforcement learning
Haarnoja, T., Moran, B., Lever, G., Huang, S. H., Tirumala, D., Humplik, J., Wulfmeier, M., Tunyasuvunakool, S., Siegel, N. Y., Hafner, R., et al · 2024
Later among the works it cites.
Towards practical reinforcement learning for tokamak magnetic control
Tracey, B. D., Michi, A., Chervonyi, Y., Davies, I., Paduraru, C., Lazic, N., Felici, F., Ewalds, T., Donner, C., Galperti, C., et al · 2024
Later among the works it cites.