Fetching the paper…
Reading the bibliography…
The objective of offline RL is to learn optimal policies when a fixed exploratory demonstrations data-set is available and sampling additional observations is impossible (typically if this operation is either costly or rises ethical questions).
Maximum entropy inverse reinforcement learning
B. D. Ziebart, A. L. Maas, J. A. Bagnell, and A. K. Dey · 2008
Earlier work this paper cites.
Relative entropy policy search
J. Peters, K. Mulling, and Y. Altun · 2010
Earlier work this paper cites.
Modeling interaction via the principle of maximum causal entropy
B. D. Ziebart, J. A. Bagnell, and A. K. Dey · 2010
Earlier work this paper cites.
A kernel two-sample test
A. Gretton, K. M. Borgwardt, M. J. Rasch, B. Schölkopf, and A. Smola · 2012
Earlier work this paper cites.
Generative adversarial networks
I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
Conditional generative adversarial nets
M. Mirza and S. Osindero · 2014
Earlier work this paper cites.
Generative adversarial imitation learning
J. Ho and S. Ermon · 2016
Earlier work this paper cites.
Learning robust rewards with adversarial inverse reinforcement learning
J. Fu, K. Luo, and S. Levine · 2017
Earlier work this paper cites.
Convex formulation of multiple instance learning from positive and unlabeled bags
H. Bao, T. Sakai, I. Sato, and M. Sugiyama · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Earlier work this paper cites.
Adversarial imitation via variational inverse reinforcement learning
A. H. Qureshi, B. Boots, and M. C. Yip · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Cited alongside, same era.
A theory of regularized markov decision processes
M. Geist, B. Scherrer, and O. Pietquin · 2019
Cited alongside, same era.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog
N. Jaques, A. Ghandeharioun, J. H. Shen, C. Ferguson, A. Lapedriza, N. Jones, S. Gu, and R. Picard · 2019
Cited alongside, same era.
Situated gail: Multitask imitation using task-conditioned adversarial inverse reinforcement learning
K. Kobayashi, T. Horii, R. Iwaki, Y. Nagai, and M. Asada · 2019
Cited alongside, same era.
Stabilizing off-policy q-learning via bootstrapping error reduction
Regularized inverse reinforcement learning
W. Jeon, C.-Y. Su, P. Barde, T. Doan, D. Nowrouzezahrai, and J. Pineau · 2020
Later among the works it cites.
Morel: Model-based offline reinforcement learning
R. Kidambi, A. Rajeswaran, P. Netrapalli, and T. Joachims · 2020
Later among the works it cites.
Semi-supervised reward learning for offline reinforcement learning
K. Konyushkova, K. Zolna, Y. Aytar, A. Novikov, S. Reed, S. Cabi, and N. de Freitas · 2020
Later among the works it cites.
Conservative q-learning for offline reinforcement learning
A. Kumar, A. Zhou, G. Tucker, and S. Levine · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Kumar, J. Fu, G. Tucker, and S. Levine · 2019
Cited alongside, same era.
Sqil: Imitation learning via reinforcement learning with sparse rewards
S. Reddy, A. D. Dragan, and S. Levine · 2019
Cited alongside, same era.
Behavior regularized offline reinforcement learning
Y. Wu, G. Tucker, and O. Nachum · 2019
Cited alongside, same era.
Learning from positive and unlabeled data: A survey
J. Bekker and J. Davis · 2020
Cited alongside, same era.
C-learning: Learning to achieve goals via recursive classification
B. Eysenbach, R. Salakhutdinov, and S. Levine · 2020
Cited alongside, same era.
D4rl: Datasets for deep data-driven reinforcement learning
J. Fu, A. Kumar, O. Nachum, G. Tucker, and S. Levine · 2020
Cited alongside, same era.
S. Levine, A. Kumar, G. Tucker, and J. Fu · 2020
Later among the works it cites.
Accelerating online reinforcement learning with offline datasets
A. Nair, M. Dalal, A. Gupta, and S. Levine · 2020
Later among the works it cites.
Mopo: Model-based offline policy optimization
T. Yu, G. Thomas, L. Yu, S. Ermon, J. Zou, S. Levine, C. Finn, and T. Ma · 2020
Later among the works it cites.
Offline learning from demonstrations and unlabeled experience
K. Zolna, A. Novikov, K. Konyushkova, C. Gulcehre, Z. Wang, Y. Aytar, M. Denil, N. de Freitas, and S. Reed · 2020
Later among the works it cites.
A generalised inverse reinforcement learning framework
F. Jarboui and V. Perchet · 2021
Closest in time.
Combo: Conservative offline model-based policy optimization
T. Yu, A. Kumar, R. Rafailov, A. Rajeswaran, S. Levine, and C. Finn · 2021
Closest in time.