Fetching the paper…
Reading the bibliography…
Policy gradients methods often achieve better performance when the change in policy is limited to a small Kullback-Leibler divergence.
Uber die umkehrung der naturgesetze
E. Schrodinger · 1931
Earlier work this paper cites.
On the transfer of masses (in russian)
L. Kantorovich · 1942
Earlier work this paper cites.
The relaxation method of finding the common point of convex sets and its application to the solution of problems in convex programming
L.M. Bregman · 1967
Earlier work this paper cites.
Concerning nonnegative matrices and doubly stochastic matrices
R. Sinkhorn and P . Knopp · 1967
Earlier work this paper cites.
Function optimization using connectionist reinforcement learning algorithms
R. J. Williams and J. Peng · 1991
Earlier work this paper cites.
Asymptotic analysis of the exponential penalty trajectory in linear programming
R. Cominetti and J. San Martin · 1994
Earlier work this paper cites.
Interior-point polynomial algorithms in convex programming
Y. Nesterov, A. Nemirovskii, and Y. Ye · 1994
Earlier work this paper cites.
Reinforcement learning by probability matching
P. Sabes and M. Jordan · 1996
Earlier work this paper cites.
The variational formulation of the Fokker-Planck equation
R. Jordan, D. Kinderlehrer, and F. Otto · 1998
Earlier work this paper cites.
Continuous martingales and Brownian motion
D. Revuz and M. Yor · 1999
Cited alongside, same era.
Gradient flows: in metric spaces and in the space of probability measures
L. Ambrosio, N. Gigli, and G. Savaré · 2006
Cited alongside, same era.
Optimal Transport : Old and New
C. Villani · 2008
Cited alongside, same era.
Transport inequalities. a survey
N. Gozlan and C. Léonard · 2010
Cited alongside, same era.
Sinkhorn distances: Lightspeed computation of optimal transport
M. Cuturi · 2013
Cited alongside, same era.
A survey of the Schrodinger problem and some of its connections with optimal transport
C. Léonard · 2014
Cited alongside, same era.
Entropic wasserstein gradient flows
G. Peyré · 2015
Later among the works it cites.
Optimal Transport for Applied Mathematicians : Calculus of Variations, PDEs and Modeling
F. Santambrogio · 2015
Later among the works it cites.
Information Geometry and Its Applications
S. Amari · 2016
Later among the works it cites.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. Puigdomenech Badia, M. Mirza, A. Graves, T. P Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Later among the works it cites.
M. Arjovsky, S. Chintala, and L. Bottou · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G. Carlier, V.Duval, G. Peyre, and B. Schmitzer · 2015
Cited alongside, same era.
Learning with a wasserstein loss
C. Frogner, C. Zhang, H. Mobahi, M. Araya-Polo, and T. Poggio · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, and et al · 2015
Cited alongside, same era.
A distributional perspective on reinforcement learning
M. G. Bellemare, W. Dabney, and R. Munos
Cited in the paper.
The cramer distance as a solution to biased wasserstein gradients
M. G. Bellemare, I. Danihelka, W. Dabney, S. Mohamed, B. Lakshminarayanan, S. Hoyer, and R. Munos
Cited in the paper.
High-dimensional continuous control using generalized advantage estimation
J. Schulman, S. Levine, P. Moritz, M. I. Jordan, and P. Abbeel
Cited in the paper.
O. Bousquet, S. Gelly, I. Tolstikhin, C. J. Simon-Gabriel, and B. Scholkopf · 2017
Closest in time.
Deep relaxation: partial differential equations for optimizing deep neural networks
P. Chaudhari, A. Oberman, S. Osher, S. Soatto, and G. Carlier · 2017
Closest in time.
Gan and vae from an optimal transport point of view
A. Genevay, G. Peyré, and M. Cuturi · 2017
Closest in time.