Fetching the paper…
Reading the bibliography…
Recent works have shown that tackling offline reinforcement learning (RL) with a conditional policy produces promising results.
Kumar, A., Peng, X. B., and Levine, S · 1912
Earlier work this paper cites.
Reinforcement Learning
Sutton, R. S. and Barto, A. G · 1998
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Precup, D · 2000
Earlier work this paper cites.
Double q-learning
Hasselt, H · 2010
Earlier work this paper cites.
Exploiting multi-step sample trajectories for approximate value iteration
Wright, R., Loscalzo, S., Dexter, P., and Yu, L · 2013
Earlier work this paper cites.
Off-policy learning with eligibility traces: a survey
Geist, M., Scherrer, B., et al · 2014
Earlier work this paper cites.
Safe and efficient off-policy reinforcement learning
Munos, R., Stepleton, T., Harutyunyan, A., and Bellemare, M · 2016
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Łukasz Kaiser, and Polosukhin, I · 2017
Earlier work this paper cites.
End-to-end driving via conditional imitation learning
Codevilla, F., Müller, M., López, A., Koltun, V., and Dosovitskiy, A · 2018
Earlier work this paper cites.
Multi-step reinforcement learning: A unifying algorithm
De Asis, K., Hernandez-Garcia, J., Holland, G., and Sutton, R · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I · 2018
Earlier work this paper cites.
Exponentially weighted imitation learning for batched historical data
Wang, Q., Xiong, J., Han, L., Liu, H., Zhang, T., et al · 2018
Earlier work this paper cites.
Off-policy deep reinforcement learning without exploration
Fujimoto, S., Meger, D., and Precup, D · 2019
Earlier work this paper cites.
Learning to reach goals via iterated supervised learning
Ghosh, D., Gupta, A., Reddy, A., Fu, J., Devin, C., Eysenbach, B., and Levine, S · 2019
Earlier work this paper cites.
Understanding multi-step deep reinforcement learning: A systematic study of the dqn target
Hernandez-Garcia, J. F. and Sutton, R. S · 2019
Earlier work this paper cites.
When to trust your model: Model-based policy optimization
Janner, M., Fu, J., Zhang, M., and Levine, S · 2019
Cited alongside, same era.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
Peng, X. B., Kumar, A., Zhang, G., and Levine, S · 2019
Cited alongside, same era.
Training agents using upside-down reinforcement learning
Srivastava, R. K., Shyam, P., Mutz, F., Jaśkowski, W., and Schmidhuber, J · 2019
Cited alongside, same era.
Behavior regularized offline reinforcement learning
Wu, Y., Tucker, G., and Nachum, O · 2019
Cited alongside, same era.
An optimistic perspective on offline reinforcement learning
Agarwal, R., Schuurmans, D., and Norouzi, M · 2020
Cited alongside, same era.
Decision transformer: Reinforcement learning via sequence modeling
Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., Abbeel, P., Srinivas, A., and Mordatch, I · 2021
Later among the works it cites.
Rvs: What is essential for offline rl via supervised learning?
Emmons, S., Eysenbach, B., Kostrikov, I., and Levine, S · 2021
Later among the works it cites.
Benchmarks for deep off-policy evaluation
Fu, J., Norouzi, M., Nachum, O., Tucker, G., Wang, Z., Novikov, A., Yang, M., Zhang, M. R., Chen, Y., Kumar, A., et al · 2021
Later among the works it cites.
A minimalist approach to offline reinforcement learning
Fujimoto, S. and Gu, S. S · 2021
Later among the works it cites.
Generalized decision transformer for offline hindsight information matching
Furuta, H., Matsuo, Y., and Gu, S. S · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bail: Best-action imitation learning for batch deep reinforcement learning
Chen, X., Zhou, Z., Wang, Z., Wang, C., Wu, Y., and Ross, K · 2020
Cited alongside, same era.
D4rl: Datasets for deep data-driven reinforcement learning, 2020
Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S · 2020
Cited alongside, same era.
Morel: Model-based offline reinforcement learning
Kidambi, R., Rajeswaran, A., Netrapalli, P., and Joachims, T · 2020
Cited alongside, same era.
Conservative q-learning for offline reinforcement learning
Kumar, A., Zhou, A., Tucker, G., and Levine, S · 2020
Cited alongside, same era.
Learning latent plans from play
Lynch, C., Khansari, M., Xiao, T., Kumar, V., Tompson, J., Levine, S., and Sermanet, P · 2020
Cited alongside, same era.
Accelerating online reinforcement learning with offline datasets
Nair, A., Dalal, M., Gupta, A., and Levine, S · 2020
Cited alongside, same era.
Hyperparameter selection for offline reinforcement learning
Paine, T. L., Paduraru, C., Michi, A., Gulcehre, C., Zolna, K., Novikov, A., Wang, Z., and de Freitas, N · 2020
Cited alongside, same era.
Later among the works it cites.
d3rlpy: An offline deep reinforcement library
Imai, M. and Seno, T · 2021
Later among the works it cites.
A workflow for offline model-free robotic reinforcement learning
Kumar, A., Singh, A., Tian, S., Finn, C., and Levine, S · 2021
Later among the works it cites.
Andrew ng launches a campaign for data-centric ai
Press, G · 2021
Later among the works it cites.
Zero-shot text-to-image generation
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I · 2021
Later among the works it cites.
Combo: Conservative offline model-based policy optimization
Yu, T., Kumar, A., Rafailov, R., Rajeswaran, A., Levine, S., and Finn, C · 2021
Later among the works it cites.
Model-based trajectory stitching for improved offline reinforcement learning
Hepburn, C. A. and Montana, G · 2022
Closest in time.
When should we prefer offline reinforcement learning over behavioral cloning?
Kumar, A., Hong, J., Singh, A., and Levine, S · 2022
Closest in time.
A survey on offline reinforcement learning: Taxonomy, review, and open problems
Prudencio, R. F., Maximo, M. R. O. A., and Colombini, E. L · 2022
Closest in time.
Wilcox, A., Balakrishna, A., Dedieu, J., Benslimane, W., Brown, D., and Goldberg, K · 2022
Closest in time.
Trajectory-aware eligibility traces for off-policy reinforcement learning
Daley, B., White, M., Amato, C., and Machado, M. C · 2023
Closest in time.