Fetching the paper…
Reading the bibliography…
Abstraction has been widely studied as a way to improve the efficiency and generalization of reinforcement learning algorithms.
Dynamic Programming
Richard Bellman · 1957
Earlier work this paper cites.
Communication and concurrency
Robin Milner · 1989
Earlier work this paper cites.
Bisimulation through probabilistic testing
Kim G Larsen and Arne Skou · 1991
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Bisimulation for labelled markov processes
Richard Blute, Josée Desharnais, Abbas Edalat, and Prakash Panangaden · 1997
Earlier work this paper cites.
Metrics for labeled Markov systems
J. Desharnais, V. Gupta, R. Jagadeesan, and P. Panangaden · 1999
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
Richard S Sutton, Doina Precup, and Satinder Singh · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Symmetries and model minimization in markov decision processes, 2001
Balaraman Ravindran and Andrew G Barto · 2001
Earlier work this paper cites.
Bisimulation for labeled Markov processes
J. Desharnais, A. Edalat, and P. Panangaden · 2002
Earlier work this paper cites.
Equivalence notions and model minimization in Markov decision processes
Robert Givan, Thomas Dean, and Matthew Greig · 2003
Earlier work this paper cites.
Relativized options: Choosing the right transformation
Balaraman Ravindran and Andrew G Barto · 2003
Earlier work this paper cites.
An algebraic approach to abstraction in reinforcement learning
Balaraman Ravindran · 2004
Earlier work this paper cites.
Approximate homomorphisms: A framework for non-exact minimization in Markov Decision Processes, 2004
Balaraman Ravindran and Andrew G Barto · 2004
Earlier work this paper cites.
Reinforcement learning with factored states and actions
Brian Sallans and Geoffrey E Hinton · 2004
Earlier work this paper cites.
Metrics for Markov decision processes with infinite state spaces
Norm Ferns, Prakash Panangaden, and Doina Precup · 2005
Earlier work this paper cites.
Methods for computing state similarity in Markov decision processes
Norm Ferns, Pablo Samuel Castro, Doina Precup, and Prakash Panangaden · 2006
Earlier work this paper cites.
Towards a unified theory of state abstraction for MDPs
Lihong Li, Thomas J Walsh, and Michael L Littman · 2006
Earlier work this paper cites.
Using homomorphisms to transfer options across continuous reinforcement learning domains
Vishal Soni and Satinder Singh · 2006
Earlier work this paper cites.
Decision tree methods for finding reusable mdp homomorphisms
Alicia P Wolfe and Andrew G Barto · 2006
Earlier work this paper cites.
Defining object types and options using mdp homomorphisms
Alicia Peregrin Wolfe and Andrew G Barto · 2006
Earlier work this paper cites.
Measure theory
Vladimir Igorevich Bogachev and Maria Aparecida Soares Ruas · 2007
Earlier work this paper cites.
On the hardness of finding symmetries in markov decision processes
Shravan Matthur Narayanamurthy and Balaraman Ravindran · 2008
Earlier work this paper cites.
Bounding performance loss in approximate mdp homomorphisms
Jonathan Taylor, Doina Precup, and Prakash Panagaden · 2008
Earlier work this paper cites.
Learning to generalize and reuse skills using approximate partial policy homomorphisms
Srividhya Rajendran and Manfred Huber · 2009
Earlier work this paper cites.
Transfer via soft homomorphisms
Jonathan Sorg and Satinder Singh · 2009
Earlier work this paper cites.
Using bisimulation for policy transfer in MDPs
Pablo Samuel Castro and Doina Precup · 2010
Earlier work this paper cites.
Automatic construction of temporally extended actions for MDPs using bisimulation metrics
Pablo Samuel Castro and Doina Precup · 2011
Earlier work this paper cites.
Bisimulation metrics for continuous Markov decision processes
Norm Ferns, Prakash Panangaden, and Doina Precup · 2011
Earlier work this paper cites.
Dynamic programming and optimal control: Volume I
Dimitri Bertsekas · 2012
Earlier work this paper cites.
Differential and Riemannian manifolds
Serge Lang · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Playing Atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Cited alongside, same era.
Deep reinforcement learning in large discrete action spaces
Gabriel Dulac-Arnold, Richard Evans, Hado van Hasselt, Peter Sunehag, Timothy Lillicrap, Jonathan Hunt, Timothy Mann, Theophane Weber, Thomas Degris, and Ben Coppin · 2015
Cited alongside, same era.
Learning state representations with robotic priors
Rico Jonschkowski and Oliver Brock · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Near optimal behavior via approximate state abstraction
David Abel, David Hershkowitz, and Michael Littman · 2016
Cited alongside, same era.
Dynamics-aware embeddings
William Whitney, Rajat Agarwal, Kyunghyun Cho, and Abhinav Gupta · 2019
Later among the works it cites.
Value preserving state-action abstractions
David Abel, Nate Umbanhowar, Khimya Khetarpal, Dilip Arumugam, Doina Precup, and Michael Littman · 2020
Later among the works it cites.
Contrastive behavioral similarity embeddings for generalization in reinforcement learning
Rishabh Agarwal, Marlos C Machado, Pablo Samuel Castro, and Marc G Bellemare · 2020
Later among the works it cites.
Scalable methods for computing state similarity in deterministic markov decision processes
Pablo Samuel Castro · 2020
Later among the works it cites.
Miles Cranmer, Sam Greydanus, Stephan Hoyer, Peter Battaglia, David Spergel, and Shirley Ho · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Cited alongside, same era.
Reinforcement learning with unsupervised auxiliary tasks
Max Jaderberg, Volodymyr Mnih, Wojciech Marian Czarnecki, Tom Schaul, Joel Z Leibo, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Symmetry learning for function approximation in reinforcement learning
Anuj Mahajan and Theja Tulabandhula · 2017
Cited alongside, same era.
Sahil Sharma, Aravind Suresh, Rahul Ramesh, and Balaraman Ravindran · 2017
Cited alongside, same era.
Independently controllable factors
Valentin Thomas, Jules Pondard, Emmanuel Bengio, Marc Sarfati, Philippe Beaudoin, Marie-Jean Meurs, Joelle Pineau, Doina Precup, and Yoshua Bengio · 2017
Cited alongside, same era.
Distributed distributional deterministic policy gradients
Gabriel Barth-Maron, Matthew W Hoffman, David Budden, Will Dabney, Dan Horgan, Dhruva Tb, Alistair Muldal, Nicolas Heess, and Timothy Lillicrap · 2018
Cited alongside, same era.
Neural scene representation and rendering
SM Ali Eslami, Danilo Jimenez Rezende, Frederic Besse, Fabio Viola, Ari S Morcos, Marta Garnelo, Avraham Ruderman, Andrei A Rusu, Ivo Danihelka, Karol Gregor, et al · 2018
Cited alongside, same era.
Christopher Grimm, André Barreto, Satinder Singh, and David Silver · 2020
Later among the works it cites.
Self-supervised policy adaptation during deployment
Nicklas Hansen, Rishabh Jangir, Yu Sun, Guillem Alenyà, Pieter Abbeel, Alexei A Efros, Lerrel Pinto, and Xiaolong Wang · 2020
Later among the works it cites.
Tonic: A deep reinforcement learning library for fast prototyping and benchmarking
Fabio Pardo · 2020
Later among the works it cites.
Learning group structure and disentangled representations of dynamical environments
Robin Quessard, Thomas D Barrett, and William R Clements · 2020
Later among the works it cites.
Plannable approximations to mdp homomorphisms: Equivariance under actions
Elise van der Pol, Thomas Kipf, Frans A Oliehoek, and Max Welling · 2020
Later among the works it cites.
Mdp homomorphic networks: Group symmetries in reinforcement learning
Elise van der Pol, Daniel Worrall, Herke van Hoof, Frans Oliehoek, and Max Welling · 2020
Later among the works it cites.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
Denis Yarats, Ilya Kostrikov, and Rob Fergus · 2020
Later among the works it cites.
Learning invariant representations for reinforcement learning without reconstruction
Amy Zhang, Rowan Thomas McAllister, Roberto Calandra, Yarin Gal, and Sergey Levine · 2020
Later among the works it cites.
Deep reinforcement learning at the edge of the statistical precipice
Rishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C Courville, and Marc Bellemare · 2021
Later among the works it cites.
Learning discrete state abstractions with deep variational inference
O Biza, R Platt, JW van de Meent, and L Wong · 2021
Later among the works it cites.
Mico: Improved representations via sampling-based state similarity for Markov decision processes
Pablo Samuel Castro, Tyler Kastner, Prakash Panangaden, and Mark Rowland · 2021
Later among the works it cites.
Secant: Self-expert cloning for zero-shot generalization of visual policies
Linxi Fan, Guanzhi Wang, De-An Huang, Zhiding Yu, Li Fei-Fei, Yuke Zhu, and Animashree Anandkumar · 2021
Later among the works it cites.
Christopher Grimm, André Barreto, Gregory Farquhar, David Silver, and Satinder Singh · 2021
Later among the works it cites.
Generalization in reinforcement learning by soft data augmentation
Nicklas Hansen and Xiaolong Wang · 2021
Later among the works it cites.
Symetric: Measuring the quality of learnt hamiltonian dynamics inferred from vision
Irina Higgins, Peter Wirnsberger, Andrew Jaegle, and Aleksandar Botev · 2021
Later among the works it cites.
Towards robust bisimulation metric learning
Mete Kemertas and Tristan Aumentado-Armstrong · 2021
Later among the works it cites.
Continuous control benchmark of DeepMind control suite and MuJoCo
Qing Li · 2021
Later among the works it cites.
Aps: Active pretraining with successor features
Hao Liu and Pieter Abbeel · 2021
Later among the works it cites.
On the effect of auxiliary tasks on representation dynamics
Clare Lyle, Mark Rowland, Georg Ostrovski, and Will Dabney · 2021
Later among the works it cites.
S4rl: Surprisingly simple self-supervision for offline reinforcement learning in robotics
Samarth Sinha, Ajay Mandlekar, and Animesh Garg · 2021
Later among the works it cites.
Decoupling representation learning from reinforcement learning
Adam Stooke, Kimin Lee, Pieter Abbeel, and Michael Laskin · 2021
Later among the works it cites.
Multi-agent MDP homomorphic networks
Elise van der Pol, Herke van Hoof, Frans A Oliehoek, and Max Welling · 2021
Later among the works it cites.
So(2)-equivariant reinforcement learning
Dian Wang, Robin Walters, and Robert Platt · 2021
Later among the works it cites.
Mastering visual continuous control: Improved data-augmented reinforcement learning
Denis Yarats, Rob Fergus, Alessandro Lazaric, and Lerrel Pinto · 2021
Later among the works it cites.
Improving sample efficiency in model-free reinforcement learning from images
Denis Yarats, Amy Zhang, Ilya Kostrikov, Brandon Amos, Joelle Pineau, and Rob Fergus · 2021
Later among the works it cites.
Eqr: Equivariant representations for data-efficient reinforcement learning
Arnab Kumar Mondal, Vineet Jain, Kaleem Siddiqi, and Siamak Ravanbakhsh · 2022
Closest in time.
Learning symmetric embeddings for equivariant world models
Jung Yeon Park, Ondrej Biza, Linfeng Zhao, Jan Willem van de Meent, and Robin Walters · 2022
Closest in time.