Fetching the paper…
Reading the bibliography…
We study how representation learning can accelerate reinforcement learning from rich observations, such as images, without relying either on domain knowledge or pixel-reconstruction.
Provably efficient RL with rich observations via latent state decoding
Simon S. Du, Akshay Krishnamurthy, Nan Jiang, Alekh Agarwal, Miroslav Dudík, and John Langford · 1901
Earlier work this paper cites.
Bisimulation through probabilistic testing (preliminary report)
K. G. Larsen and A. Skou · 1989
Earlier work this paper cites.
Towards quantitative verification of probabilistic transition systems
Franck van Breugel and James Worrell · 2001
Earlier work this paper cites.
Equivalence notions and model minimization in Markov decision processes
Robert Givan, Thomas L. Dean, and Matthew Greig · 2003
Earlier work this paper cites.
Topics in optimal transportation
Cédric Villani · 2003
Earlier work this paper cites.
Metrics for finite Markov decision processes
Norm Ferns, Prakash Panangaden, and Doina Precup · 2004
Earlier work this paper cites.
Causal graph based decomposition of factored MDPs
Anders Jonsson and Andrew Barto · 2006
Earlier work this paper cites.
Towards a unified theory of state abstraction for MDPs
Lihong Li, Thomas J Walsh, and Michael L Littman · 2006
Earlier work this paper cites.
Bounding performance loss in approximate MDP homomorphisms
Jonathan Taylor, Doina Precup, and Prakash Panagaden · 2009
Earlier work this paper cites.
Deep auto-encoder neural networks in reinforcement learning
Sascha Lange and Martin Riedmiller · 2010
Earlier work this paper cites.
Bisimulation metrics for continuous Markov decision processes
Norm Ferns, Prakash Panangaden, and Doina Precup · 2011
Earlier work this paper cites.
Autonomous reinforcement learning on raw visual input data in a real world application
Sascha Lange, Martin Riedmiller, and Arne Voigtländer · 2012
Earlier work this paper cites.
Bisimulation metrics are optimal value functions
Norman Ferns and Doina Precup · 2014
Earlier work this paper cites.
Learning state representations with robotic priors
Rico Jonschkowski and Oliver Brock · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Cited alongside, same era.
From pixels to torques: Policy learning with deep dynamical models
Niklas Wahlström, Thomas Schön, and Marc Deisenroth · 2015
Cited alongside, same era.
Embed to control: A locally linear latent dynamics model for control from raw images
Manuel Watter, Jost Springenberg, Joschka Boedecker, and Martin Riedmiller · 2015
Cited alongside, same era.
CARLA: An open urban driving simulator
Alexey Dosovitskiy, German Ros, Felipe Codevilla, Antonio Lopez, and Vladlen Koltun · 2017
Cited alongside, same era.
DeepMDP: Learning continuous latent space models for representation learning
Carles Gelada, Saurabh Kumar, Jacob Buckman, Ofir Nachum, and Marc G. Bellemare · 2019
Later among the works it cites.
Learning latent dynamics for planning from pixels
Danijar Hafner, Timothy Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, and James Davidson · 2019
Later among the works it cites.
Efficient off-policy meta-reinforcement learning via probabilistic context variables
Kate Rakelly, Aurick Zhou, Deirdre Quillen, Chelsea Finn, and Sergey Levine · 2019
Later among the works it cites.
Causality for machine learning, 2019
Bernhard Schölkopf · 2019
Later among the works it cites.
Scalable methods for computing state similarity in deterministic Markov decision processes
Pablo Samuel Castro · 2020
Closest in time.
A simple framework for contrastive learning of visual representations
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The kinetics human action video dataset
Will Kay, João Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, Mustafa Suleyman, and Andrew Zisserman · 2017
Cited alongside, same era.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Kurtland Chua, Roberto Calandra, Rowan McAllister, and Sergey Levine · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Cited alongside, same era.
Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, Timothy Lillicrap, and Martin Riedmiller · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
Aäron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Cited alongside, same era.
Natural environment benchmarks for reinforcement learning
Amy Zhang, Yuxin Wu, and Joelle Pineau · 2018
Cited alongside, same era.
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton · 2020
Closest in time.
Data-efficient image recognition with contrastive predictive coding
Olivier J Hénaff, Aravind Srinivas, Jeffrey De Fauw, Ali Razavi, Carl Doersch, SM Eslami, and Aaron van den Oord · 2020
Closest in time.
CURL: Contrastive unsupervised representations for reinforcement learning
Michael Laskin, Aravind Srinivas, and Pieter Abbeel · 2020
Closest in time.
Stochastic latent actor-critic: Deep reinforcement learning with a latent variable model
Alex Lee, Anusha Nagabandi, Pieter Abbeel, and Sergey Levine · 2020
Closest in time.
Soft actor-critic (SAC) implementation in PyTorch
Denis Yarats and Ilya Kostrikov · 2020
Closest in time.
Invariant causal prediction for block MDPs
Amy Zhang, Clare Lyle, Shagun Sodhani, Angelos Filos, Marta Kwiatkowska, Joelle Pineau, Yarin Gal, and Doina Precup · 2020
Closest in time.
Improving sample efficiency in model-free reinforcement learning from images
Denis Yarats, Amy Zhang, Ilya Kostrikov, Brandon Amos, Joelle Pineau, and Rob Fergus · 2021
Closest in time.