Fetching the paper…
Reading the bibliography…
Model-based reinforcement learning (RL) methods are appealing in the offline setting because they allow an agent to reason about the consequences of actions without interacting with the environment.
Efficient memory-based learning for robot control
A. W. Moore · 1990
Earlier work this paper cites.
Improving generalization for temporal difference learning: The successor representation
P. Dayan · 1993
Earlier work this paper cites.
Error bounds for approximate policy iteration
R. Munos · 2003
Earlier work this paper cites.
Random features for large-scale kernel machines
A. Rahimi and B. Recht · 2007
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
M. Gutmann and A. Hyvärinen · 2010
Earlier work this paper cites.
Efficient estimation of word representations in vector space
T. Mikolov, K. Chen, G. Corrado, and J. Dean · 2013
Earlier work this paper cites.
J. L. Ba, J. R. Kiros, and G. E. Hinton · 2016
Earlier work this paper cites.
Successor features for transfer in reinforcement learning
A. Barreto, W. Dabney, R. Munos, J. J. Hunt, T. Schaul, H. Van Hasselt, and D. Silver · 2016
Earlier work this paper cites.
Densely connected convolutional networks
G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger · 2017
Earlier work this paper cites.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
L. Espeholt, H. Soyer, R. Munos, K. Simonyan, V. Mnih, T. Ward, Y. Doron, V. Firoiu, T. Harley, I. Dunning, et al · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Earlier work this paper cites.
Learning deep representations by mutual information estimation and maximization
R. D. Hjelm, A. Fedorov, S. Lavoie-Marchildon, K. Grewal, P. Bachman, A. Trischler, and Y. Bengio · 2018
Earlier work this paper cites.
Emi: Exploration with mutual information
H. Kim, J. Kim, Y. Jeong, S. Levine, and H. O. Song · 2018
Earlier work this paper cites.
Z. Ma and M. Collins · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
A. v. d. Oord, Y. Li, and O. Vinyals · 2018
Cited alongside, same era.
Implicit generation and modeling with energy based models
Y. Du and I. Mordatch · 2019
Cited alongside, same era.
Off-policy deep reinforcement learning without exploration
S. Fujimoto, D. Meger, and D. Precup · 2019
Cited alongside, same era.
Behavior regularized offline reinforcement learning
Y. Wu, G. Tucker, and O. Nachum · 2019
Cited alongside, same era.
Curl: Contrastive unsupervised representations for reinforcement learning
A. Srinivas, M. Laskin, and P. Abbeel · 2020
Later among the works it cites.
Understanding contrastive representation learning through alignment and uniformity on the hypersphere
T. Wang and P. Isola · 2020
Later among the works it cites.
Learning successor states and goal-dependent values: A mathematical viewpoint
L. Blier, C. Tallec, and Y. Ollivier · 2021
Later among the works it cites.
Phasic policy gradient
K. W. Cobbe, J. Hilton, O. Klimov, and J. Schulman · 2021
Later among the works it cites.
Provable guarantees for self-supervised deep learning with spectral contrastive loss
J. Z. HaoChen, C. Wei, A. Gaidon, and T. Ma · 2021
Later among the works it cites.
Provable representation learning for imitation with contrastive fourier features
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Argenson and G. Dulac-Arnold · 2020
Cited alongside, same era.
A simple framework for contrastive learning of visual representations
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton · 2020
Cited alongside, same era.
C-learning: Learning to achieve goals via recursive classification
B. Eysenbach, R. Salakhutdinov, and S. Levine · 2020
Cited alongside, same era.
Momentum contrast for unsupervised visual representation learning
K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick · 2020
Cited alongside, same era.
Generative temporal difference learning for infinite-horizon prediction
M. Janner, I. Mordatch, and S. Levine · 2020
Cited alongside, same era.
Morel: Model-based offline reinforcement learning
R. Kidambi, A. Rajeswaran, P. Netrapalli, and T. Joachims · 2020
Cited alongside, same era.
Conservative q-learning for offline reinforcement learning
A. Kumar, A. Zhou, G. Tucker, and S. Levine · 2020
Cited alongside, same era.
O. Nachum and M. Yang · 2021
Later among the works it cites.
Pretraining representations for data-efficient reinforcement learning
M. Schwarzer, N. Rajkumar, M. Noukhovitch, A. Anand, L. Charlin, D. Hjelm, P. Bachman, and A. Courville · 2021
Later among the works it cites.
Learning one representation to optimize all rewards
A. Touati and Y. Ollivier · 2021
Later among the works it cites.
Combo: Conservative offline model-based policy optimization
T. Yu, A. Kumar, R. Rafailov, A. Rajeswaran, S. Levine, and C. Finn · 2021
Later among the works it cites.
Adaptive behavior cloning regularization for stable offline-to-online reinforcement learning
Y. Zhao, R. Boney, A. Ilin, J. Kannala, and J. Pajarinen · 2021
Later among the works it cites.
Contrastive learning as goal-conditioned reinforcement learning
B. Eysenbach, T. Zhang, R. Salakhutdinov, and S. Levine · 2022
Closest in time.
How to leverage unlabeled data in offline reinforcement learning
T. Yu, A. Kumar, Y. Chebotar, K. Hausman, C. Finn, and S. Levine · 2022
Closest in time.