Fetching the paper…
Reading the bibliography…
Two desiderata of reinforcement learning (RL) algorithms are the ability to learn from relatively little experience and the ability to learn policies that generalize to a range of problem specifications.
Learning causal state representations of partially observable environments
Zhang, A.; Lipton, Z. C.; Pineda, L.; Azizzadenesheli, K.; Anandkumar, A.; Itti, L.; Pineau, J.; and Furlanello, T. 2019 · 1906
Earlier work this paper cites.
Model minimization in Markov decision processes
Dean, T.; and Givan, R. 1997 · 1997
Earlier work this paper cites.
A New Learning Algorithm for Mean Field Boltzmann Machines
Welling, M.; and Hinton, G. E. 2002 · 2002
Earlier work this paper cites.
Energy-Based Models for Sparse Overcomplete Representations
Teh, Y. W.; Welling, M.; Osindero, S.; and Hinton, G. E. 2003 · 2003
Earlier work this paper cites.
A Tutorial on Energy-Based Learning
LeCun, Y.; Chopra, S.; Hadsell, R.; Ranzato, A.; and Huang, F. J. 2006 · 2006
Earlier work this paper cites.
Towards a unified theory of state abstraction for MDPs
Li, L.; Walsh, T. J.; and Littman, M. L. 2006 · 2006
Earlier work this paper cites.
Learning invariant representations for reinforcement learning without reconstruction
Zhang, A.; McAllister, R.; Calandra, R.; Gal, Y.; and Levine, S. 2020b · 2006
Earlier work this paper cites.
Causality
Pearl, J. 2009 · 2009
Earlier work this paper cites.
robosuite: A Modular Simulation Framework and Benchmark for Robot Learning
Zhu, Y.; Wong, J.; Mandlekar, A.; and Martín-Martín, R. 2020 · 2009
Earlier work this paper cites.
Bisimulation metrics for continuous Markov decision processes
Ferns, N.; Panangaden, P.; and Precup, D. 2011 · 2011
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Jang, E.; Gu, S.; and Poole, B. 2016 · 2016
Earlier work this paper cites.
Reinforcement Learning with Deep Energy-Based Policies
Haarnoja, T.; Tang, H.; Abbeel, P.; and Levine, S. 2017 · 2017
Earlier work this paper cites.
Information theoretic MPC for model-based reinforcement learning
Williams, G.; Wagener, N.; Goldfain, B.; Drews, P.; Rehg, J. M.; Boots, B.; and Theodorou, E. A. 2017 · 2017
Earlier work this paper cites.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Chua, K.; Calandra, R.; McAllister, R.; and Levine, S. 2018 · 2018
Cited alongside, same era.
Model-based value estimation for efficient model-free reinforcement learning
Feinberg, V.; Wan, A.; Stoica, I.; Jordan, M. I.; Gonzalez, J. E.; and Levine, S. 2018 · 2018
Cited alongside, same era.
Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
Haarnoja, T.; Zhou, A.; Abbeel, P.; and Levine, S. 2018 · 2018
Cited alongside, same era.
Model-ensemble trust-region policy optimization
Kurutach, T.; Clavera, I.; Duan, Y.; Tamar, A.; and Abbeel, P. 2018 · 2018
Cited alongside, same era.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
On the model-based stochastic value gradient for continuous reinforcement learning
Amos, B.; Stanton, S.; Yarats, D.; and Wilson, A. G. 2021 · 2021
Later among the works it cites.
Residual Energy-Based Models for Text
Bakhtin, A.; Deng, Y.; Gross, S.; Ott, M.; Ranzato, M.; and Szlam, A. 2021 · 2021
Later among the works it cites.
Learning task informed abstractions
Fu, X.; Yang, G.; Agrawal, P.; and Jaakkola, T. 2021 · 2021
Later among the works it cites.
Joint Energy-based Model Training for Better Calibrated Natural Language Understanding Models
He, T.; McCann, B.; Xiong, C.; and Hosseini-Asl, E. 2021 · 2021
Later among the works it cites.
Necessary and sufficient conditions for causal feature selection in time series with latent common causes
Mastakouri, A. A.; Schölkopf, B.; and Janzing, D. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nagabandi, A.; Kahn, G.; Fearing, R. S.; and Levine, S. 2018 · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
Oord, A. v. d.; Li, Y.; and Vinyals, O. 2018 · 2018
Cited alongside, same era.
Deep Energy Estimator Networks
Saremi, S.; Mehrjou, A.; Schölkopf, B.; and Hyvärinen, A. 2018 · 2018
Cited alongside, same era.
Implicit Generation and Modeling with Energy Based Models
Du, Y.; and Mordatch, I. 2019 · 2019
Cited alongside, same era.
When to Trust Your Model: Model-Based Policy Optimization
Janner, M.; Fu, J.; Zhang, M.; and Levine, S. 2019 · 2019
Cited alongside, same era.
Sliced Score Matching: A Scalable Approach to Density and Score Estimation
Song, Y.; Garg, S.; Shi, J.; and Ermon, S. 2019 · 2019
Cited alongside, same era.
ContactNets: Learning of Discontinuous Contact Dynamics with Smooth, Implicit Representations
Pfrommer, S.; Halm, M.; and Posa, M. 2020 · 2020
Cited alongside, same era.
dm control: Software and tasks for continuous control
Tunyasuvunakool, S.; Muldal, A.; Doron, Y.; Liu, S.; Bohez, S.; Merel, J.; Erez, T.; Lillicrap, T.; Heess, N.; and Tassa, Y. 2020 · 2020
Cited alongside, same era.
Song, Y.; and Kingma, D. P. 2021 · 2021
Later among the works it cites.
Decomposed mutual information estimation for contrastive representation learning
Sordoni, A.; Dziri, N.; Schulz, H.; Gordon, G.; Bachman, P.; and Des Combes, R. T. 2021 · 2021
Later among the works it cites.
Task-Independent Causal State Abstraction
Wang, Z.; Xiao, X.; Zhu, Y.; and Stone, P. 2021 · 2021
Later among the works it cites.
Generalizing Goal-Conditioned Reinforcement Learning with Variational Causal Reasoning
Ding, W.; Lin, H.; Li, B.; and Zhao, D. 2022 · 2022
Later among the works it cites.
Implicit behavioral cloning
Florence, P.; Lynch, C.; Zeng, A.; Ramirez, O. A.; Wahid, A.; Downs, L.; Wong, A.; Lee, J.; Mordatch, I.; and Tompson, J. 2022 · 2022
Later among the works it cites.
Action-sufficient state representation learning for control with structural constraints
Huang, B.; Lu, C.; Leqi, L.; Hernández-Lobato, J. M.; Glymour, C.; Schölkopf, B.; and Zhang, K. 2022 · 2022
Later among the works it cites.
The Primacy Bias in Deep Reinforcement Learning
Nikishin, E.; Schwarzer, M.; D’Oro, P.; Bacon, P.-L.; and Courville, A. 2022 · 2022
Later among the works it cites.
Tianshou: A Highly Modularized Deep Reinforcement Learning Library
Weng, J.; Chen, H.; Yan, D.; You, K.; Duburcq, A.; Zhang, M.; Su, Y.; Su, H.; and Zhu, J. 2022 · 2022
Later among the works it cites.