Fetching the paper…
Reading the bibliography…
Representations are at the core of all deep reinforcement learning (RL) methods for both Markov decision processes (MDPs) and partially observable Markov decision processes (POMDPs).
Theory of Self-Adaptive Control Systems
H. Kwakernaak · 1965
Earlier work this paper cites.
Sufficient statistics in the optimum control of stochastic systems
Charlotte Striebel · 1965
Earlier work this paper cites.
Information pattern for linear discrete-time models with stochastic coefficients
Torsten Bohlin · 1970
Earlier work this paper cites.
Stochastic Systems: Estimation, Identification and Adaptive Control
P. R. Kumar and Pravin Varaiya · 1986
Earlier work this paper cites.
Probabilistic reasoning in intelligent systems: networks of plausible inference
Judea Pearl · 1988
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Richard S Sutton · 1988
Earlier work this paper cites.
Finding structure in time
Jeffrey L Elman · 1990
Earlier work this paper cites.
Efficient memory-based learning for robot control
Andrew William Moore · 1990
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Richard S Sutton · 1990
Earlier work this paper cites.
Improving generalization for temporal difference learning: The successor representation
Peter Dayan · 1993
Earlier work this paper cites.
Acting optimally in partially observable stochastic domains
Anthony R. Cassandra, Leslie Pack Kaelbling, and Michael L. Littman · 1994
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 1994
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Leemon Baird · 1995
Earlier work this paper cites.
Generalization in reinforcement learning: Successful examples using sparse coarse coding
Richard S Sutton · 1995
Earlier work this paper cites.
Model minimization in markov decision processes
Thomas Dean and Robert Givan · 1997
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Leslie Pack Kaelbling, Michael L. Littman, and Anthony R. Cassandra · 1998
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David McAllester, Satinder Singh, and Yishay Mansour · 1999
Earlier work this paper cites.
Predictive representations of state
Michael L. Littman, Richard S. Sutton, and Satinder P. Singh · 2001
Earlier work this paper cites.
Computational mechanics: Pattern and prediction, structure and simplicity
Cosma Rohilla Shalizi and James P Crutchfield · 2001
Earlier work this paper cites.
Equivalence notions and model minimization in markov decision processes
Robert Givan, Thomas Dean, and Matthew Greig · 2003
Earlier work this paper cites.
Learning predictive state representations
Satinder P Singh, Michael L Littman, Nicholas K Jong, David Pardoe, and Peter Stone · 2003
Earlier work this paper cites.
Metrics for finite markov decision processes
Norm Ferns, Prakash Panangaden, and Doina Precup · 2004
Earlier work this paper cites.
Planning with predictive state representations
Michael R James, Satinder Singh, and Michael L Littman · 2004
Earlier work this paper cites.
Using rewards for belief state updates in partially observable Markov decision processes
Masoumeh T. Izadi and Doina Precup · 2005
Earlier work this paper cites.
Towards a unified theory of state abstraction for mdps
Lihong Li, Thomas J Walsh, and Michael L Littman · 2006
Earlier work this paper cites.
Equivalence relations in fully and partially observable markov decision processes
Pablo Samuel Castro, Prakash Panangaden, and Doina Precup · 2009
Earlier work this paper cites.
Mastering atari with discrete world models
Danijar Hafner, Timothy Lillicrap, Mohammad Norouzi, and Jimmy Ba · 2010
Earlier work this paper cites.
Deep auto-encoder neural networks in reinforcement learning
Sascha Lange and Martin Riedmiller · 2010
Earlier work this paper cites.
Value function approximation in reinforcement learning using the fourier basis
George Konidaris, Sarah Osentoski, and Philip Thomas · 2011
Earlier work this paper cites.
Dyna-style planning with linear function approximation and prioritized sweeping
Richard S Sutton, Csaba Szepesvári, Alborz Geramifard, and Michael P Bowling · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Solving pomdps by searching the space of finite policies
Nicolas Meuleau, Kee-Eung Kim, Leslie Pack Kaelbling, and Anthony R Cassandra · 2013
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M Saxe, James L McClelland, and Surya Ganguli · 2013
Earlier work this paper cites.
Equivalence of distance-based and rkhs-based statistics in hypothesis testing
Dino Sejdinovic, Bharath Sriperumbudur, Arthur Gretton, and Kenji Fukumizu · 2013
Earlier work this paper cites.
Fast and accurate deep network learning by exponential linear units (elus)
Djork-Arné Clevert, Thomas Unterthiner, and Sepp Hochreiter · 2015
Earlier work this paper cites.
Deep recurrent Q-learning for partially observable mdps
Matthew J. Hausknecht and Peter Stone · 2015
Earlier work this paper cites.
Learning continuous control policies by stochastic value gradients
Nicolas Heess, Gregory Wayne, David Silver, Timothy Lillicrap, Tom Erez, and Yuval Tassa · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Embed to control: A locally linear latent dynamics model for control from raw images
Manuel Watter, Jost Tobias Springenberg, Joschka Boedecker, and Martin A. Riedmiller · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Reinforcement learning with unsupervised auxiliary tasks
Max Jaderberg, Volodymyr Mnih, Wojciech Marian Czarnecki, Tom Schaul, Joel Z Leibo, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Deep successor reinforcement learning
Tejas D Kulkarni, Ardavan Saeedi, Simanta Gautam, and Samuel J Gershman · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2016
Cited alongside, same era.
Learning state representation for deep actor-critic control
Jelle Munk, Jens Kober, and Robert Babuška · 2016
Data-efficient reinforcement learning with self-predictive representations
Max Schwarzer, Ankesh Anand, Rishab Goel, R Devon Hjelm, Aaron Courville, and Philip Bachman · 2020
Later among the works it cites.
Learning invariant representations for reinforcement learning without reconstruction
Amy Zhang, Rowan McAllister, Roberto Calandra, Yarin Gal, and Sergey Levine · 2020
Later among the works it cites.
Deep reinforcement learning at the edge of the statistical precipice
Rishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville, and Marc G. Bellemare · 2021
Later among the works it cites.
Learning markov state abstractions for deep reinforcement learning
Cameron Allen, Neev Parikh, Omer Gottesman, and George Konidaris · 2021
Later among the works it cites.
Reconciling rewards with predictive state representations
Andrea Baisero and Christopher Amato · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Loss is its own reward: Self-supervision for reinforcement learning
Evan Shelhamer, Parsa Mahmoudieh, Max Argus, and Trevor Darrell · 2016
Cited alongside, same era.
Value iteration networks
Aviv Tamar, Yi Wu, Garrett Thomas, Sergey Levine, and Pieter Abbeel · 2016
Cited alongside, same era.
Deep reinforcement learning with double q-learning
Hado Van Hasselt, Arthur Guez, and David Silver · 2016
Cited alongside, same era.
Successor features for transfer in reinforcement learning
André Barreto, Will Dabney, Rémi Munos, Jonathan J Hunt, Tom Schaul, Hado P van Hasselt, and David Silver · 2017
Cited alongside, same era.
Value prediction network
Junhyuk Oh, Satinder Singh, and Honglak Lee · 2017
Cited alongside, same era.
The predictron: End-to-end learning and planning
David Silver, Hado van Hasselt, Matteo Hessel, Tom Schaul, Arthur Guez, Tim Harley, Gabriel Dulac-Arnold, David P. Reichert, Neil C. Rabinowitz, André Barreto, and Thomas Degris · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Mico: Improved representations via sampling-based state similarity for markov decision processes
Pablo Samuel Castro, Tyler Kastner, Prakash Panangaden, and Mark Rowland · 2021
Later among the works it cites.
Robust predictable control
Benjamin Eysenbach, Russ R Salakhutdinov, and Sergey Levine · 2021
Later among the works it cites.
Learning task informed abstractions
Xiang Fu, Ge Yang, Pulkit Agrawal, and Tommi Jaakkola · 2021
Later among the works it cites.
Understanding dimensional collapse in contrastive self-supervised learning
Li Jing, Pascal Vincent, Yann LeCun, and Yuandong Tian · 2021
Later among the works it cites.
On the effect of auxiliary tasks on representation dynamics
Clare Lyle, Mark Rowland, Georg Ostrovski, and Will Dabney · 2021
Later among the works it cites.
Dreaming: Model-based reinforcement learning by latent imagination without reconstruction
Masashi Okada and Tadahiro Taniguchi · 2021
Later among the works it cites.
Which mutual-information representation learning objectives are sufficient for control?
Kate Rakelly, Abhishek Gupta, Carlos Florensa, and Sergey Levine · 2021
Later among the works it cites.
The distracting control suite – a challenging benchmark for reinforcement learning from pixels
Austin Stone, Oscar Ramirez, Kurt Konolige, and Rico Jonschkowski · 2021
Later among the works it cites.
Learning representations for pixel-based control: What matters and why?
Manan Tomar, Utkarsh A Mishra, Amy Zhang, and Matthew E Taylor · 2021
Later among the works it cites.
Improving sample efficiency in model-free reinforcement learning from images
Denis Yarats, Amy Zhang, Ilya Kostrikov, Brandon Amos, Joelle Pineau, and Rob Fergus · 2021
Later among the works it cites.
Mastering atari games with limited data
Weirui Ye, Shaohuai Liu, Thanard Kurutach, Pieter Abbeel, and Yang Gao · 2021
Later among the works it cites.
Metacure: Meta reinforcement learning with empowerment-driven exploration
Jin Zhang, Jianhao Wang, Hao Hu, Tong Chen, Yingfeng Chen, Changjie Fan, and Chongjie Zhang · 2021
Later among the works it cites.
Exploration in approximate hyper-state space for meta reinforcement learning
Luisa M. Zintgraf, Leo Feng, Cong Lu, Maximilian Igl, Kristian Hartikainen, Katja Hofmann, and Shimon Whiteson · 2021
Later among the works it cites.
Dreamerpro: Reconstruction-free model-based reinforcement learning with prototypical representations
Fei Deng, Ingook Jang, and Sungjin Ahn · 2022
Later among the works it cites.
Raj Ghugare, Homanga Bharadhwaj, Benjamin Eysenbach, Sergey Levine, and Ruslan Salakhutdinov · 2022
Later among the works it cites.
Byol-explore: Exploration by bootstrapped prediction
Zhaohan Daniel Guo, Shantanu Thakoor, Miruna Pîslar, Bernardo Avila Pires, Florent Altché, Corentin Tallec, Alaa Saade, Daniele Calandriello, Jean-Bastien Grill, Yunhao Tang, et al · 2022
Later among the works it cites.
Temporal difference learning for model predictive control
Nicklas Hansen, Xiaolong Wang, and Hao Su · 2022
Later among the works it cites.
Recurrent networks, hidden states and beliefs in partially observable environments
Gaspard Lambrechts, Adrien Bolland, and Damien Ernst · 2022
Later among the works it cites.
Does self-supervised learning really improve reinforcement learning from pixels?
Xiang Li, Jinghuan Shang, Srijan Das, and Michael Ryoo · 2022
Later among the works it cites.
Recurrent model-free rl can be a strong baseline for many pomdps
Tianwei Ni, Benjamin Eysenbach, and Ruslan Salakhutdinov · 2022
Later among the works it cites.
Control-oriented model-based reinforcement learning with implicit differentiation
Evgenii Nikishin, Romina Abachi, Rishabh Agarwal, and Pierre-Luc Bacon · 2022
Later among the works it cites.
On learning history based policies for controlling markov decision processes
Gandharv Patil, Aditya Mahajan, and Doina Precup · 2022
Later among the works it cites.
Approximate information state for approximate planning and reinforcement learning in partially observed systems
Jayakumar Subramanian, Amit Sinha, Raihan Seraj, and Aditya Mahajan · 2022
Later among the works it cites.
Understanding self-predictive learning for reinforcement learning
Yunhao Tang, Zhaohan Daniel Guo, Pierre Harvey Richemond, Bernardo Ávila Pires, Yash Chandak, Rémi Munos, Mark Rowland, Mohammad Gheshlaghi Azar, Charline Le Lan, Clare Lyle, et al · 2022
Later among the works it cites.
Denoised mdps: Learning world models better than the world itself
Tongzhou Wang, Simon S Du, Antonio Torralba, Phillip Isola, Amy Zhang, and Yuandong Tian · 2022
Later among the works it cites.
Discrete approximate information states in partially observable environments
Lujie Yang, Kaiqing Zhang, Alexandre Amice, Yunzhu Li, and Russ Tedrake · 2022
Later among the works it cites.
Minigrid & miniworld: Modular & customizable reinforcement learning environments for goal-oriented tasks
Maxime Chevalier-Boisvert, Bolun Dai, Mark Towers, Rodrigo de Lazcano, Lucas Willems, Salem Lahlou, Suman Pal, Pablo Samuel Castro, and Jordan Terry · 2023
Later among the works it cites.
Contrabar: Contrastive bayes-adaptive deep rl
Era Choshen and Aviv Tamar · 2023
Later among the works it cites.
For sale: State-action representation learning for deep reinforcement learning
Scott Fujimoto, Wei-Di Chang, Edward J Smith, Shixiang Shane Gu, Doina Precup, and David Meger · 2023
Later among the works it cites.
Comparing auxiliary tasks for learning representations for reinforcement learning, 2023
Moritz Lange, Noah Krystiniak, Raphael Engelhardt, Wolfgang Konen, and Laurenz Wiskott · 2023
Later among the works it cites.
Structure in reinforcement learning: A survey and open problems
Aditya Mohan, Amy Zhang, and Marius Lindauer · 2023
Later among the works it cites.
Popgym: Benchmarking partially observable reinforcement learning
Steven Morad, Ryan Kortvelesy, Matteo Bettini, Stephan Liwicki, and Amanda Prorok · 2023
Later among the works it cites.
When do transformers shine in rl? decoupling memory from credit assignment
Tianwei Ni, Michel Ma, Benjamin Eysenbach, and Pierre-Luc Bacon · 2023
Later among the works it cites.
Approximate information state based convergence analysis of recurrent q-learning
Erfan Seyedsalehi, Nima Akbarzadeh, Amit Sinha, and Aditya Mahajan · 2023
Later among the works it cites.
Simplified temporal consistency reinforcement learning
Yi Zhao, Wenshuai Zhao, Rinu Boney, Juho Kannala, and Joni Pajarinen · 2023
Later among the works it cites.