Fetching the paper…
Reading the bibliography…
We present a new behavioural distance over the state space of a Markov decision process, and demonstrate the use of this distance as an effective means of shaping the learnt representations of deep reinforcement learning agents.
Robust Estimation of a Location Parameter
Peter J. Huber · 1964
Earlier work this paper cites.
Communication and Concurrency
R. Milner · 1989
Earlier work this paper cites.
Bisimulation through probablistic testing
Kim G Larsen and Arne Skou · 1991
Earlier work this paper cites.
Improving generalization for temporal difference learning: The successor representation
Peter Dayan · 1993
Earlier work this paper cites.
Partial metric topology
Steve Matthews · 1994
Earlier work this paper cites.
On the generation of Markov decision processes
T. W. Archibald, K. I. M. McKinnon, and L. C. Thomas · 1995
Earlier work this paper cites.
Neuro-Dynamic Programming
Dimitri P. Bertsekas and John N. Tsitsiklis · 1996
Earlier work this paper cites.
Reinforcement Learning with Selective Perception and Hidden State
Andrew Kachites McCallum · 1996
Earlier work this paper cites.
Metrics for labeled Markov systems
Josée Desharnais, Vineet Gupta, Radhakrishnan Jagadeesan, and Prakash Panangaden · 1999
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
R.S. Sutton, D. Precup, and S. Singh · 1999
Earlier work this paper cites.
Equivalence notions and model minimization in Markov decision processes
Robert Givan, Thomas Dean, and Matthew Greig · 2003
Earlier work this paper cites.
SMDP homomorphisms: An algebraic approach to abstraction in semi-Markov decision processes
Balaraman Ravindran and Andrew G. Barto · 2003
Earlier work this paper cites.
A metric for labelled Markov processes
Josée Desharnais, Vineet Gupta, Radhakrishnan Jagadeesan, and Prakash Panangaden · 2004
Earlier work this paper cites.
Metrics for finite Markov decision processes
Norm Ferns, Prakash Panangaden, and Doina Precup · 2004
Earlier work this paper cites.
A new concept of probability metric and its applications in approximation of scattered data sets
Szymon Łukaszyk · 2004
Earlier work this paper cites.
Metrics for Markov decision processes with infinite state spaces
Norman Ferns, Prakash Panangaden, and Doina Precup · 2005
Earlier work this paper cites.
State abstraction discovery from irrelevant state variables
Nicholas K Jong and Peter Stone · 2005
Earlier work this paper cites.
Methods for computing state similarity in Markov decision processes
Norm Ferns, Pablo Samuel Castro, Doina Precup, and Prakash Panangaden · 2006
Earlier work this paper cites.
Towards a unified theory of state abstraction for MDPs
Lihong Li, Thomas J Walsh, and Michael L Littman · 2006
Earlier work this paper cites.
Coinductive proof principles for stochastic processes
Dexter Kozen · 2007
Earlier work this paper cites.
Proto-value functions: A Laplacian framework for learning representation and control in Markov decision processes
Sridhar Mahadevan and Mauro Maggioni · 2007
Earlier work this paper cites.
Lax probabilistic bisimulation
Jonathan Taylor · 2008
Earlier work this paper cites.
Optimal Transport
Cédric Villani · 2008
Cited alongside, same era.
Fast and robust earth mover’s distances
Ofir Pele and Michael Werman · 2009
Cited alongside, same era.
Bounding performance loss in approximate MDP homomorphisms
Jonathan Taylor, Doina Precup, and Prakash Panagaden · 2009
Cited alongside, same era.
Using bisimulation for policy transfer in MDPs
Pablo Samuel Castro and Doina Precup · 2010
Cited alongside, same era.
Bisimulation metrics for continuous Markov decision processes
Norm Ferns, Prakash Panangaden, and Doina Precup · 2011
Cited alongside, same era.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Richard S Sutton, Joseph Modayil, Michael Delp, Thomas Degris, Patrick M Pilarski, Adam White, and Doina Precup · 2011
Cited alongside, same era.
Rainbow: Combining Improvements in Deep Reinforcement learning
Matteo Hessel, Joseph Modayil, Hado van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, Mohammad Azar, and David Silver · 2018
Later among the works it cites.
Revisiting the Arcade Learning Environment: Evaluation protocols and open problems for general agents
Marlos C Machado, Marc G Bellemare, Erik Talvitie, Joel Veness, Matthew Hausknecht, and Michael Bowling · 2018
Later among the works it cites.
UMAP: Uniform manifold approximation and projection
Leland McInnes, John Healy, Nathaniel Saul, and Lukas Grossberger · 2018
Later among the works it cites.
Deepmind control suite
Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, Timothy P. Lillicrap, and Martin A. Riedmiller · 2018
Later among the works it cites.
Temporally extended metrics for Markov decision processes
Philip Amortila, Marc G Bellemare, Prakash Panangaden, and Doina Precup · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On-the-fly algorithms for bisimulation metrics
Gheorghe Comanici, Prakash Panangaden, and Doina Precup · 2012
Cited alongside, same era.
Approximation metrics based on probabilistic bisimulations for general state-space Markov processes: A survey
Alessandro Abate · 2013
Cited alongside, same era.
The Arcade Learning Environment: An evaluation platform for general agents
Marc G. Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Cited alongside, same era.
Sinkhorn distances: Lightspeed computation of optimal transport
Marco Cuturi · 2013
Cited alongside, same era.
Bisimulation metrics are optimal value functions
Norm Ferns and Doina Precup · 2014
Cited alongside, same era.
Path finding methods for linear programming: Solving linear programs in
Yin Tat Lee and Aaron Sidford · 2014
Cited alongside, same era.
Hyperbolic discounting and learning over multiple horizons
William Fedus, Carles Gelada, Yoshua Bengio, Marc G Bellemare, and Hugo Larochelle · 2019
Later among the works it cites.
Combined reinforcement learning via abstract representations
Vincent François-Lavet, Yoshua Bengio, Doina Precup, and Joelle Pineau · 2019
Later among the works it cites.
DeepMDP: Learning continuous latent space models for representation learning
Carles Gelada, Saurabh Kumar, Jacob Buckman, Ofir Nachum, and Marc G Bellemare · 2019
Later among the works it cites.
Computational optimal transport
Gabriel Peyré and Marco Cuturi · 2019
Later among the works it cites.
ExTra: Transfer-guided exploration
Anirban Santara, Rishabh Madan, Balaraman Ravindran, and Pabitra Mitra · 2019
Later among the works it cites.
Improving sample efficiency in model-free reinforcement learning from images
Denis Yarats, Amy Zhang, Ilya Kostrikov, Brandon Amos, Joelle Pineau, and Rob Fergus · 2019
Later among the works it cites.
Scalable methods for computing state similarity in deterministic Markov Decision Processes
Pablo Samuel Castro · 2020
Later among the works it cites.
Faster wasserstein distance estimation with the sinkhorn divergence
Lenaic Chizat, Pierre Roussillon, Flavien Léger, François-Xavier Vialard, and Gabriel Peyré · 2020
Later among the works it cites.
Learning with minibatch Wasserstein: asymptotic and gradient properties
Kilian Fatras, Younes Zine, Rémi Flamary, Rémi Gribonval, and Nicolas Courty · 2020
Later among the works it cites.
Learning to score behaviors for guided policy optimization
Aldo Pacchiano, Jack Parker-Holder, Yunhao Tang, Krzysztof Choromanski, Anna Choromanska, and Michael Jordan · 2020
Later among the works it cites.
Munchausen reinforcement learning
Nino Vieillard, Olivier Pietquin, and Matthieu Geist · 2020
Later among the works it cites.
Minibatch optimal transport distances; analysis and applications
Kilian Fatras, Younes Zine, Szymon Majewski, Rémi Flamary, Rémi Gribonval, and Nicolas Courty · 2021
Closest in time.
Metrics and continuity in reinforcement learning
Charline Le Lan, Marc G. Bellemare, and Pablo Samuel Castro · 2021
Closest in time.
Efficient Wasserstein natural gradients for reinforcement learning
Ted Moskovitz, Michael Arbel, Ferenc Huszar, and Arthur Gretton · 2021
Closest in time.
Revisiting rainbow: Promoting more insightful and inclusive deep reinforcement learning research
Johan S Obando-Ceron and Pablo Samuel Castro · 2021
Closest in time.
Invariant representations for reinforcement learning without reconstruction
Amy Zhang, Rowan McAllister, Roberto Calandra, Yarin Gal, and Sergey Levine · 2021
Closest in time.