Fetching the paper…
Reading the bibliography…
The empirical success of distributional reinforcement learning (RL) highly relies on the choice of distribution divergence equipped with an appropriate distribution representation.
Information theory and statistical mechanics
Edwin T Jaynes · 1957
Earlier work this paper cites.
Diagonal equivalence to matrices with prescribed row and column sums
Richard Sinkhorn · 1967
Earlier work this paper cites.
Generalized iterative scaling for log-linear models
John N Darroch and Douglas Ratcliff · 1972
Earlier work this paper cites.
Asymptotic evaluation of certain markov process expectations for large time—iii
Monroe D Donsker and SR Srinivasa Varadhan · 1976
Earlier work this paper cites.
On the scaling of multidimensional matrices
Joel Franklin and Jens Lorenz · 1989
Earlier work this paper cites.
Closedness of sum spaces andthe generalized schrödinger problem
Ludger Rüschendorf and Wolfgang Thomsen · 1998
Earlier work this paper cites.
E-statistics: The energy of statistical samples
Gábor J Székely · 2003
Earlier work this paper cites.
A kernel method for the two-sample-problem
Arthur Gretton, Karsten Borgwardt, Malte Rasch, Bernhard Schölkopf, and Alex Smola · 2006
Earlier work this paper cites.
Kernel choice and classifiability for rkhs embeddings of probability distributions
Kenji Fukumizu, Arthur Gretton, Gert Lanckriet, Bernhard Schölkopf, and Bharath K Sriperumbudur · 2009
Earlier work this paper cites.
Efficient reinforcement learning with multiple reward functions for randomized controlled trial analysis
Daniel J Lizotte, Michael H Bowling, and Susan A Murphy · 2010
Earlier work this paper cites.
Super-samples from kernel herding
Yutian Chen, Max Welling, and Alex Smola · 2012
Earlier work this paper cites.
Sinkhorn distances: Lightspeed computation of optimal transport
Marco Cuturi · 2013
Earlier work this paper cites.
Christian Léonard · 2013
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Wasserstein generative adversarial networks
Martin Arjovsky, Soumith Chintala, and Léon Bottou · 2017
Earlier work this paper cites.
A distributional perspective on reinforcement learning
Marc G Bellemare, Will Dabney, and Rémi Munos · 2017
Earlier work this paper cites.
The cramer distance as a solution to biased wasserstein gradients
Marc G Bellemare, Ivo Danihelka, Will Dabney, Shakir Mohamed, Balaji Lakshminarayanan, Stephan Hoyer, and Rémi Munos · 2017
Earlier work this paper cites.
Near-linear time approximation algorithms for optimal transport via sinkhorn iteration, 2017
Philippe Rigollet Jason Altschuler, Jonathan Weed · 2017
Earlier work this paper cites.
On wasserstein two-sample testing and related families of nonparametric tests
Aaditya Ramdas, Nicolás García Trillos, and Marco Cuturi · 2017
Earlier work this paper cites.
Hybrid reward architecture for reinforcement learning
Harm Van Seijen, Mehdi Fatemi, Joshua Romoff, Romain Laroche, Tavian Barnes, and Jeffrey Tsang · 2017
Earlier work this paper cites.
Implicit quantile networks for distributional reinforcement learning
Will Dabney, Georg Ostrovski, David Silver, and Rémi Munos · 2018
Cited alongside, same era.
Distributional reinforcement learning with quantile regression
Will Dabney, Mark Rowland, Marc G Bellemare, and Rémi Munos · 2018
Cited alongside, same era.
Learning generative models with sinkhorn divergences
Aude Genevay, Gabriel Peyré, and Marco Cuturi · 2018
Cited alongside, same era.
Screening sinkhorn algorithm for regularized optimal transport
Mokhtar Z Alaya, Maxime Berar, Gilles Gasso, and Alain Rakotomamonjy · 2019
Cited alongside, same era.
Interpolating between optimal transport and mmd using sinkhorn divergences
Jean Feydy, Thibault Séjourné, François-Xavier Vialard, Shun-ichi Amari, Alain Trouvé, and Gabriel Peyré · 2019
Cited alongside, same era.
Sample complexity of sinkhorn divergences
Aude Genevay, Lénaic Chizat, Francis Bach, Marco Cuturi, and Gabriel Peyré · 2019
Towards deeper deep reinforcement learning with spectral normalization
Nils Bjorck, Carla P. Gomes, and Kilian Q. Weinberger · 2021
Later among the works it cites.
Don’t generate me: Training differentially private generative models with sinkhorn divergence
Tianshi Cao, Alex Bie, Arash Vahdat, Sanja Fidler, and Karsten Kreis · 2021
Later among the works it cites.
Unbalanced minibatch optimal transport; applications to domain adaptation
Kilian Fatras, Thibault Séjourné, Rémi Flamary, and Nicolas Courty · 2021
Later among the works it cites.
Distributional reinforcement learning with monotonic splines
Yudong Luo, Guiliang Liu, Haonan Duan, Oliver Schulte, and Pascal Poupart · 2021
Later among the works it cites.
Conservative offline distributional reinforcement learning
Yecheng Ma, Dinesh Jayaraman, and Osbert Bastani · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Distributional reward decomposition for reinforcement learning
Zichuan Lin, Li Zhao, Derek Yang, Tao Qin, Tie-Yan Liu, and Guangwen Yang · 2019
Cited alongside, same era.
Distributional reinforcement learning for efficient exploration
Borislav Mavrin, Shangtong Zhang, Hengshuai Yao, Linglong Kong, Kaiwen Wu, and Yaoliang Yu · 2019
Cited alongside, same era.
Computational optimal transport: With applications to data science
Gabriel Peyré, Marco Cuturi, et al · 2019
Cited alongside, same era.
Statistics and samples in distributional reinforcement learning
Mark Rowland, Robert Dadashi, Saurabh Kumar, Rémi Munos, Marc G Bellemare, and Will Dabney · 2019
Cited alongside, same era.
Wasserstein adversarial examples via projected sinkhorn iterations
Eric Wong, Frank Schmidt, and Zico Kolter · 2019
Cited alongside, same era.
Fully parameterized quantile function for distributional reinforcement learning
Derek Yang, Li Zhao, Zichuan Lin, Tao Qin, Jiang Bian, and Tie-Yan Liu · 2019
Cited alongside, same era.
Ke Sun, Yingnan Zhao, Yi Liu, Enze Shi, Yafei Wang, Xiaodong Yan, Bei Jiang, and Linglong Kong · 2021
Later among the works it cites.
Distributional reinforcement learning for multi-dimensional reward functions
Pushi Zhang, Xiaoyu Chen, Li Zhao, Wei Xiong, Tao Qin, and Tie-Yan Liu · 2021
Later among the works it cites.
Quantitative stability of regularized optimal transport and convergence of sinkhorn’s algorithm
Stephan Eckstein and Marcel Nutz · 2022
Closest in time.
Reinforcement learning can be more efficient with multiple rewards
Christoph Dann, Yishay Mansour, and Mehryar Mohri · 2023
Closest in time.
The statistical benefits of quantile temporal-difference learning for value estimation
Mark Rowland, Yunhao Tang, Clare Lyle, Rémi Munos, Marc G Bellemare, and Will Dabney · 2023
Closest in time.
Exploring the training robustness of distributional reinforcement learning against noisy state observations
Ke Sun, Yingnan Zhao, Shangling Jui, and Linglong Kong · 2023
Closest in time.
The benefits of being distributional: Small-loss bounds for reinforcement learning
Kaiwen Wang, Kevin Zhou, Runzhe Wu, Nathan Kallus, and Wen Sun · 2023
Closest in time.
Distributional offline policy evaluation with predictive error guarantees
Runzhe Wu, Masatoshi Uehara, and Wen Sun · 2023
Closest in time.
Provable risk-sensitive distributional reinforcement learning with general function approximation
Yu Chen, Xiangcheng Zhang, Siwei Wang, and Longbo Huang · 2024
Closest in time.
Pitfall of optimism: Distributional reinforcement learning by randomizing risk criterion
Taehyun Cho, Seungyub Han, Heesoo Lee, Kyungjae Lee, and Jungwoo Lee · 2024
Closest in time.
An analysis of quantile temporal-difference learning
Mark Rowland, Rémi Munos, Mohammad Gheshlaghi Azar, Yunhao Tang, Georg Ostrovski, Anna Harutyunyan, Karl Tuyls, Marc G Bellemare, and Will Dabney · 2024
Closest in time.
How does return distribution in distributional reinforcement learning help optimization?
Ke Sun, Bei Jiang, and Linglong Kong · 2024
Closest in time.
More benefits of being distributional: Second-order bounds for reinforcement learning
Kaiwen Wang, Owen Oertell, Alekh Agarwal, Nathan Kallus, and Wen Sun · 2024
Closest in time.
Distributional bellman operators over mean embeddings
Li Kevin Wenliang, Grégoire Déletang, Matthew Aitchison, Marcus Hutter, Anian Ruoss, Arthur Gretton, and Mark Rowland · 2024
Closest in time.