Fetching the paper…
Reading the bibliography…
Learning representations for reinforcement learning (RL) has shown much promise for continuous control.
Dream to control: Learning behaviors by latent imagination
Danijar Hafner, Timothy Lillicrap, Jimmy Ba, and Mohammad Norouzi · 1912
Earlier work this paper cites.
A Markovian Decision Process
Richard Bellman · 1957
Earlier work this paper cites.
Reinforcement Learning with Soft State Aggregation
Satinder Singh, Tommi Jaakkola, and Michael Jordan · 1994
Earlier work this paper cites.
Abstraction and approximate decision-theoretic planning
Richard Dearden and Craig Boutilier · 1997
Earlier work this paper cites.
Optimization of computer simulation models with rare events
Reuven Y Rubinstein · 1997
Earlier work this paper cites.
State abstraction for programmable reinforcement learning agents
David Andre and Stuart J. Russell · 2002
Earlier work this paper cites.
Dynamic abstraction in reinforcement learning via clustering
Shie Mannor, Ishai Menache, Amit Hoze, and Uri Klein · 2004
Earlier work this paper cites.
Towards a Unified Theory of State Abstraction for MDPs
Lihong Li, Thomas Walsh, and Michael Littman · 2006
Earlier work this paper cites.
Data-efficient reinforcement learning with self-predictive representations
Max Schwarzer, Ankesh Anand, Rishab Goel, R Devon Hjelm, Aaron Courville, and Philip Bachman · 2007
Earlier work this paper cites.
Autonomous reinforcement learning on raw visual input data in a real world application
Sascha Lange, Martin Riedmiller, and Arne Voigtländer · 2012
Earlier work this paper cites.
Auto-Encoding Variational Bayes
Diederik P. Kingma and M. Welling · 2014
Earlier work this paper cites.
Embed to Control: A Locally Linear Latent Dynamics Model for Control from Raw Images
Manuel Watter, Jost Springenberg, Joschka Boedecker, and Martin Riedmiller · 2015
Earlier work this paper cites.
Model predictive path integral control using covariance variable importance sampling
Grady Williams, Andrew Aldrich, and Evangelos Theodorou · 2015
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Near Optimal Behavior via Approximate State Abstraction
David Abel, David Hershkowitz, and Michael Littman · 2016
Earlier work this paper cites.
Stable reinforcement learning with autoencoders for tactile and visual data
Herke van Hoof, Nutan Chen, Maximilian Karl, Patrick van der Smagt, and Jan Peters · 2016
Earlier work this paper cites.
Deep spatial autoencoders for visuomotor learning
Chelsea Finn, Xin Yu Tan, Yan Duan, Trevor Darrell, Sergey Levine, and Pieter Abbeel · 2016
Earlier work this paper cites.
Layer Normalization, July 2016
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton · 2016
Cited alongside, same era.
DARLA: Improving Zero-Shot Transfer in Reinforcement Learning
Irina Higgins, Arka Pal, Andrei Rusu, Loic Matthey, Christopher Burgess, Alexander Pritzel, Matthew Botvinick, Charles Blundell, and Alexander Lerchner · 2017
Cited alongside, same era.
Reinforcement Learning, Second Edition: An Introduction
R.S. Sutton and A.G. Barto · 2018
Cited alongside, same era.
Addressing Function Approximation Error in Actor-Critic Methods
Scott Fujimoto, Herke Hoof, and David Meger · 2018
Cited alongside, same era.
Towards a Definition of Disentangled Representations, December 2018
Irina Higgins, David Amos, David Pfau, Sebastien Racaniere, Loic Matthey, Danilo Rezende, and Alexander Lerchner · 2018
Cited alongside, same era.
Varibad: Variational bayes-adaptive deep rl via meta-learning
Luisa Zintgraf, Sebastian Schulze, Cong Lu, Leo Feng, Maximilian Igl, Kyriacos Shiarlis, Yarin Gal, Katja Hofmann, and Shimon Whiteson · 2021
Later among the works it cites.
Learning representations for pixel-based control: What matters and why?
Manan Tomar, Utkarsh A Mishra, Amy Zhang, and Matthew E Taylor · 2021
Later among the works it cites.
Mastering atari games with limited data
Weirui Ye, Shaohuai Liu, Thanard Kurutach, Pieter Abbeel, and Yang Gao · 2021
Later among the works it cites.
Learning dynamics models for model predictive agents
Michael Lutter, Leonard Hasenclever, Arunkumar Byravan, Gabriel Dulac-Arnold, Piotr Trochim, Nicolas Heess, Josh Merel, and Yuval Tassa · 2021
Later among the works it cites.
Temporal Difference Learning for Model Predictive Control
Nicklas A. Hansen, Hao Su, and Xiaolong Wang · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Amy Zhang, Yuxin Wu, and Joelle Pineau · 2018
Cited alongside, same era.
Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, et al · 2018
Cited alongside, same era.
Recurrent World Models Facilitate Policy Evolution
David Ha and Jürgen Schmidhuber · 2018
Cited alongside, same era.
Unsupervised State Representation Learning in Atari
Ankesh Anand, Evan Racah, Sherjil Ozair, Yoshua Bengio, Marc-Alexandre Côté, and R Devon Hjelm · 2019
Cited alongside, same era.
DeepMDP: Learning Continuous Latent Space Models for Representation Learning
Carles Gelada, Saurabh Kumar, Jacob Buckman, Ofir Nachum, and Marc G. Bellemare · 2019
Cited alongside, same era.
PyTorch: An Imperative Style, High-Performance Deep Learning Library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Cited alongside, same era.
CURL: Contrastive Unsupervised Representations for Reinforcement Learning
Michael Laskin, Aravind Srinivas, and Pieter Abbeel · 2020
Cited alongside, same era.
Continuous MDP Homomorphisms and Homomorphic Policy Gradient
Sahand Rezaei-Shoshtari, Rosie Zhao, Prakash Panangaden, David Meger, and Doina Precup · 2022
Later among the works it cites.
Approximate information state for approximate planning and reinforcement learning in partially observed systems
Jayakumar Subramanian, Amit Sinha, Raihan Seraj, and Aditya Mahajan · 2022
Later among the works it cites.
The Mechanism of Prediction Head in Non-contrastive Self-supervised Learning
Zixin Wen and Yuanzhi Li · 2022
Later among the works it cites.
Mastering Atari with Discrete World Models
Danijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, and Jimmy Ba · 2022
Later among the works it cites.
For SALE: State-Action Representation Learning for Deep Reinforcement Learning
Scott Fujimoto, Wei-Di Chang, Edward Smith, Shixiang (Shane) Gu, Doina Precup, and David Meger · 2023
Later among the works it cites.
Simplified Temporal Consistency Reinforcement Learning
Yi Zhao, Wenshuai Zhao, Rinu Boney, Juho Kannala, and Joni Pajarinen · 2023
Later among the works it cites.
Finite Scalar Quantization: VQ-VAE Made Simple, September 2023
Fabian Mentzer, David Minnen, Eirikur Agustsson, and Michael Tschannen · 2023
Later among the works it cites.
Bridging State and History Representations: Understanding Self-Predictive RL
Tianwei Ni, Benjamin Eysenbach, Erfan SeyedSalehi, Michel Ma, Clement Gehring, Aditya Mahajan, and Pierre-Luc Bacon · 2023
Later among the works it cites.
TD-MPC2: Scalable, Robust World Models for Continuous Control
Nicklas Hansen, Hao Su, and Xiaolong Wang · 2023
Later among the works it cites.
Mastering diverse domains through world models
Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap · 2023
Later among the works it cites.
Taco: Temporal latent action-driven contrastive loss for visual reinforcement learning
Ruijie Zheng, Xiyao Wang, Yanchao Sun, Shuang Ma, Jieyu Zhao, Huazhe Xu, Hal Daumé III, and Furong Huang · 2024
Closest in time.