Fetching the paper…
Reading the bibliography…
Reward-free data is abundant and contains rich prior knowledge of human behaviors, but it is not well exploited by offline reinforcement learning (RL) algorithms.
Alvinn: An autonomous land vehicle in a neural network
Pomerleau, D. A · 1988
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y · 2010
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Earlier work this paper cites.
Learning to optimize via posterior sampling
Russo, D. and Van Roy, B · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Earlier work this paper cites.
Generative adversarial imitation learning
Ho, J. and Ermon, S · 2016
Earlier work this paper cites.
Deep exploration via bootstrapped dqn
Osband, I., Blundell, C., Pritzel, A., and Van Roy, B · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Earlier work this paper cites.
Hindsight experience replay
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Pieter Abbeel, O., and Zaremba, W · 2017
Earlier work this paper cites.
Imitation learning: A survey of learning methods
Hussein, A., Gaber, M. M., Elyan, E., and Jayne, C · 2017
Earlier work this paper cites.
Why is posterior sampling better than optimism for reinforcement learning?
Osband, I. and Van Roy, B · 2017
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
Pathak, D., Agrawal, P., Efros, A. A., and Darrell, T · 2017
Earlier work this paper cites.
Learning complex dexterous manipulation with deep reinforcement learning and demonstrations
Rajeswaran, A., Kumar, V., Gupta, A., Vezzani, G., Schulman, J., Todorov, E., and Levine, S · 2017
Earlier work this paper cites.
Generalization properties of learning with random features
Rudi, A. and Rosasco, L · 2017
Earlier work this paper cites.
Large-scale study of curiosity-driven learning
Burda, Y., Edwards, H., Pathak, D., Storkey, A., Darrell, T., and Efros, A. A · 2018
Earlier work this paper cites.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Chua, K., Calandra, R., McAllister, R., and Levine, S · 2018
Earlier work this paper cites.
Diversity is all you need: Learning skills without a reward function
Eysenbach, B., Gupta, A., Ibarz, J., and Levine, S · 2018
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
Fujimoto, S., Hoof, H., and Meger, D · 2018
Earlier work this paper cites.
Fast task inference with variational intrinsic successor features
Hansen, S., Dabney, W., Barreto, A., Van de Wiele, T., Warde-Farley, D., and Mnih, V · 2019
Earlier work this paper cites.
Information-theoretic confidence bounds for reinforcement learning
Lu, X. and Van Roy, B · 2019
Earlier work this paper cites.
Self-supervised exploration via disagreement
Pathak, D., Gandhi, D., and Gupta, A · 2019
Earlier work this paper cites.
Dynamics-aware unsupervised discovery of skills
Sharma, A., Gu, S., Levine, S., Kumar, V., and Hausman, K · 2019
Earlier work this paper cites.
Wasserstein adversarial imitation learning
Xiao, H., Herman, M., Wagner, J., Ziesche, S., Etesami, J., and Linh, T. H · 2019
Earlier work this paper cites.
Opal: Offline primitive discovery for accelerating offline reinforcement learning
Ajay, A., Kumar, A., Agrawal, P., Levine, S., and Nachum, O · 2020
Cited alongside, same era.
Provably efficient exploration in policy optimization
Cai, Q., Yang, Z., Jin, C., and Wang, Z · 2020
Cited alongside, same era.
Primal wasserstein imitation learning
Dadashi, R., Hussenot, L., Geist, M., and Pietquin, O · 2020
Cited alongside, same era.
D4rl: Datasets for deep data-driven reinforcement learning
Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S · 2020
Cited alongside, same era.
Provably efficient reinforcement learning with linear function approximation
Jin, C., Yang, Z., Wang, Z., and Jordan, M. I · 2020
Cited alongside, same era.
Believe what you see: Implicit constraint approach for offline multi-agent reinforcement learning
Yang, Y., Ma, X., Li, C., Zheng, Z., Zhang, Q., Huang, G., Yang, J., and Zhao, Q · 2021
Later among the works it cites.
Conservative data sharing for multi-task offline reinforcement learning
Yu, T., Kumar, A., Chebotar, Y., Hausman, K., Levine, S., and Finn, C · 2021
Later among the works it cites.
Magnetic control of tokamak plasmas through deep reinforcement learning
Degrave, J., Felici, F., Buchli, J., Neunert, M., Tracey, B., Carpanese, F., Ewalds, T., Hafner, R., Abdolmaleki, A., de Las Casas, D., et al · 2022
Later among the works it cites.
Offline rl policies should be trained to be adaptive
Ghosh, D., Ajay, A., Agrawal, P., and Levine, S · 2022
Later among the works it cites.
On the role of discount factor in offline reinforcement learning
Hu, H., Yang, Y., Zhao, Q., and Zhang, C · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Conservative q-learning for offline reinforcement learning
Kumar, A., Zhou, A., Tucker, G., and Levine, S · 2020
Cited alongside, same era.
A policy gradient method for task-agnostic exploration
Mutti, M., Pratissoli, L., and Restelli, M · 2020
Cited alongside, same era.
Awac: Accelerating online reinforcement learning with offline datasets
Nair, A., Gupta, A., Dalal, M., and Levine, S · 2020
Cited alongside, same era.
Parrot: Data-driven behavioral priors for reinforcement learning
Singh, A., Liu, H., Zhou, G., Yu, A., Rhinehart, N., and Levine, S · 2020
Cited alongside, same era.
Self-supervised equivariant attention mechanism for weakly supervised semantic segmentation
Wang, Y., Zhang, J., Kan, M., Shan, S., and Chen, X · 2020
Cited alongside, same era.
Reinforcement learning in feature space: Matrix bandit, kernels, and regret bound
Yang, L. and Wang, M · 2020
Cited alongside, same era.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Yu, T., Quillen, D., He, Z., Julian, R., Hausman, K., Finn, C., and Levine, S · 2020
Cited alongside, same era.
Lee, S., Seo, Y., Lee, K., Abbeel, P., and Shin, J · 2022
Later among the works it cites.
Challenges and opportunities in offline reinforcement learning from visual observations
Lu, C., Ball, P. J., Rudner, T. G., Parker-Holder, J., Osborne, M. A., and Teh, Y. W · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Later among the works it cites.
Hybrid rl: Using both offline and online data can make rl efficient
Song, Y., Zhou, Y., Sekhari, A., Bagnell, J. A., Krishnamurthy, A., and Sun, W · 2022
Later among the works it cites.
The role of coverage in online reinforcement learning
Xie, T., Foster, D. J., Bai, Y., Jiang, N., and Kakade, S. M · 2022
Later among the works it cites.
Xiong, W., Zhong, H., Shi, C., Shen, C., Wang, L., and Zhang, T · 2022
Later among the works it cites.
Flow to control: Offline reinforcement learning with lossless primitive discovery
Yang, Y., Hu, H., Li, W., Li, S., Yang, J., Zhao, Q., and Zhang, C · 2022
Later among the works it cites.
Become a proficient player with limited data through watching pure videos
Ye, W., Zhang, Y., Abbeel, P., and Gao, Y · 2022
Later among the works it cites.
How to leverage unlabeled data in offline reinforcement learning
Yu, T., Kumar, A., Chebotar, Y., Hausman, K., Finn, C., and Levine, S · 2022
Later among the works it cites.
Cup: Critic-guided policy reuse
Zhang, J., Li, S., and Zhang, C · 2022
Later among the works it cites.
Efficient online reinforcement learning with offline data
Ball, P. J., Smith, L., Kostrikov, I., and Levine, S · 2023
Closest in time.
Reinforcement learning from passive data via latent intentions
Ghosh, D., Bhateja, C., and Levine, S · 2023
Closest in time.
The provable benefits of unsupervised data sharing for offline reinforcement learning
Hu, H., Yang, Y., Zhao, Q., and Zhang, C · 2023
Closest in time.
Survival instinct in offline reinforcement learning
Li, A., Misra, D., Kolobov, A., and Cheng, C.-A · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
Benchmarks and algorithms for offline preference-based reward learning
Shin, D., Dragan, A. D., and Brown, D. S · 2023
Closest in time.
Policy expansion for bridging offline-to-online reinforcement learning
Zhang, H., Xu, W., and Yu, H · 2023
Closest in time.