Fetching the paper…
Reading the bibliography…
Learned representations in deep reinforcement learning (DRL) have to extract task-relevant information from complex observations, balancing between robustness to distraction and informativeness to the policy.
The distance between two random vectors with given dispersion matrices
Ingram Olkin and Friedrich Pukelsheim · 1982
Earlier work this paper cites.
Efficient memory-based learning for robot control
Andrew William Moore · 1990
Earlier work this paper cites.
Equivalence notions and model minimization in markov decision processes
Robert Givan, Thomas Dean, and Matthew Greig · 2003
Earlier work this paper cites.
Metrics for finite Markov decision processes
Norm Ferns, Prakash Panangaden, and Doina Precup · 2004
Earlier work this paper cites.
An algebraic approach to abstraction in reinforcement learning
Balaraman Ravindran · 2004
Earlier work this paper cites.
Approximate homomorphisms: A framework for non-exact minimization in Markov decision processes
Balaraman Ravindran and Andrew G Barto · 2004
Earlier work this paper cites.
Towards a unified theory of state abstraction for MDPs
Lihong Li, Thomas J. Walsh, and Michael L. Littman · 2006
Earlier work this paper cites.
An elementary proof of the triangle inequality for the Wasserstein metric
Philippe Clement and Wolfgang Desch · 2008
Earlier work this paper cites.
Bounding performance loss in approximate MDP homomorphisms
Jonathan Taylor, Doina Precup, and Prakash Panagaden · 2008
Earlier work this paper cites.
Optimal transport: old and new
Cédric Villani · 2008
Earlier work this paper cites.
Biasing approximate dynamic programming with a lower discount factor
Marek Petrik and Bruno Scherrer · 2009
Earlier work this paper cites.
Using bisimulation for policy transfer in MDPs
Pablo Castro and Doina Precup · 2010
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
Brian D Ziebart · 2010
Earlier work this paper cites.
Bisimulation metrics for continuous Markov decision processes
Norm Ferns, Prakash Panangaden, and Doina Precup · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Bisimulation metrics are optimal value functions
Norm Ferns and Doina Precup · 2014
Earlier work this paper cites.
Empowerment–an introduction
Christoph Salge, Cornelius Glackin, and Daniel Polani · 2014
Earlier work this paper cites.
Optimal transport for applied mathematicians
Filippo Santambrogio · 2015
Earlier work this paper cites.
FaceNet: A unified embedding for face recognition and clustering
Florian Schroff, Dmitry Kalenichenko, and James Philbin · 2015
Earlier work this paper cites.
Incentivizing exploration in reinforcement learning with deep predictive models
Bradly C Stadie, Sergey Levine, and Pieter Abbeel · 2015
Cited alongside, same era.
Embed to control: a locally linear latent dynamics model for control from raw images
Manuel Watter, Jost Tobias Springenberg, Joschka Boedecker, and Martin Riedmiller · 2015
Cited alongside, same era.
Learning to poke by poking: experiential learning of intuitive physics
Pulkit Agrawal, Ashvin Nair, Pieter Abbeel, Jitendra Malik, and Sergey Levine · 2016
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
Marc G Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Cited alongside, same era.
Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, et al · 2018
Later among the works it cites.
A survey on intrinsic motivation in reinforcement learning
Arthur Aubret, Laetitia Matignon, and Salima Hassas · 2019
Later among the works it cites.
Online abstraction with MDP homomorphisms for deep learning
Ondrej Biza and Robert Platt · 2019
Later among the works it cites.
Combined reinforcement learning via abstract representations
Vincent François-Lavet, Yoshua Bengio, Doina Precup, and Joelle Pineau · 2019
Later among the works it cites.
DeepMDP: Learning continuous latent space models for representation learning
Carles Gelada, Saurabh Kumar, Jacob Buckman, Ofir Nachum, and Marc G Bellemare · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Karol Gregor, Danilo Jimenez Rezende, and Daan Wierstra · 2016
Cited alongside, same era.
Surprise-based intrinsic motivation for deep reinforcement learning
Joshua Achiam and Shankar Sastry · 2017
Cited alongside, same era.
In defense of the triplet loss for person re-identification
Alexander Hermans, Lucas Beyer, and Bastian Leibe · 2017
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
Deepak Pathak, Pulkit Agrawal, Alexei A Efros, and Trevor Darrell · 2017
Cited alongside, same era.
Sampling matters in deep embedding learning
Chao-Yuan Wu, R. Manmatha, Alexander J. Smola, and Philipp Krähenbühl · 2017
Cited alongside, same era.
Large-scale study of curiosity-driven learning
Yuri Burda, Harri Edwards, Deepak Pathak, Amos Storkey, Trevor Darrell, and Alexei A Efros · 2018
Cited alongside, same era.
Efficient model-based deep reinforcement learning with variational state tabulation
Dane Corneil, Wulfram Gerstner, and Johanni Brea · 2018
Cited alongside, same era.
Learning latent dynamics for planning from pixels
Danijar Hafner, Timothy Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, and James Davidson · 2019
Later among the works it cites.
Stochastic latent actor-critic: Deep reinforcement learning with a latent variable model
Alex X Lee, Anusha Nagabandi, Pieter Abbeel, and Sergey Levine · 2019
Later among the works it cites.
Solar: Deep structured representations for model-based reinforcement learning
Marvin Zhang, Sharad Vikram, Laura Smith, Pieter Abbeel, Matthew Johnson, and Sergey Levine · 2019
Later among the works it cites.
Scalable methods for computing state similarity in deterministic Markov decision processes
Pablo Samuel Castro · 2020
Later among the works it cites.
The value equivalence principle for model-based reinforcement learning
Christopher Grimm, Andre Barreto, Satinder Singh, and David Silver · 2020
Later among the works it cites.
Curl: Contrastive unsupervised representations for reinforcement learning
Michael Laskin, Aravind Srinivas, and Pieter Abbeel · 2020
Later among the works it cites.
Dreaming: Model-based reinforcement learning by latent imagination without reconstruction
Masashi Okada and Tadahiro Taniguchi · 2020
Later among the works it cites.
Decoupling representation learning from reinforcement learning
Adam Stooke, Kimin Lee, Pieter Abbeel, and Michael Laskin · 2020
Later among the works it cites.
Plannable approximations to MDP homomorphisms: Equivariance under actions
Elise van der Pol, Thomas Kipf, Frans A Oliehoek, and Max Welling · 2020
Later among the works it cites.
MDP homomorphic networks: Group symmetries in reinforcement learning
Elise van der Pol, Daniel Worrall, Herke van Hoof, Frans Oliehoek, and Max Welling · 2020
Later among the works it cites.
Contrastive behavioral similarity embeddings for generalization in reinforcement learning
Rishabh Agarwal, Marlos C. Machado, Pablo Samuel Castro, and Marc G. Bellemare · 2021
Closest in time.
Metrics and continuity in reinforcement learning
Charline Le Lan, Marc G Bellemare, and Pablo Samuel Castro · 2021
Closest in time.
Invariant representations for reinforcement learning without reconstruction
Amy Zhang, Rowan Thomas McAllister, Roberto Calandra, Yarin Gal, and Sergey Levine · 2021
Closest in time.