Fetching the paper…
Reading the bibliography…
As a data-driven paradigm, offline reinforcement learning (RL) has been formulated as sequence modeling that conditions on the hindsight information including returns, goal or future trajectory.
On the estimation of production frontiers: maximum likelihood estimation of the parameters of a discontinuous density function
Aigner, D. J., Amemiya, T., and Poirier, D. J · 1976
Earlier work this paper cites.
Asymmetric least squares estimation and testing
Newey, W. K. and Powell, J. L · 1987
Earlier work this paper cites.
Alvinn: An autonomous land vehicle in a neural network
Pomerleau, D. A · 1988
Earlier work this paper cites.
Q-learning
Watkins, C. J. and Dayan, P · 1992
Earlier work this paper cites.
Introduction to reinforcement learning
Sutton, R. S., Barto, A. G., et al · 1998
Earlier work this paper cites.
Actor-critic algorithms
Konda, V. and Tsitsiklis, J · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D., Singh, S., and Mansour, Y · 1999
Earlier work this paper cites.
Degris, T., White, M., and Sutton, R. S · 2012
Earlier work this paper cites.
Geoadditive expectile regression
Sobotka, F. and Kneib, T · 2012
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Earlier work this paper cites.
Off-policy deep reinforcement learning without exploration
Fujimoto, S., Meger, D., and Precup, D · 2019
Earlier work this paper cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
Kumar, A., Fu, J., Soh, M., Tucker, G., and Levine, S · 2019
Cited alongside, same era.
Reinforcement learning upside down: Don’t predict rewards–just map them to actions
Schmidhuber, J · 2019
Cited alongside, same era.
Training agents using upside-down reinforcement learning
Srivastava, R. K., Shyam, P., Mutz, F., Jaśkowski, W., and Schmidhuber, J · 2019
Cited alongside, same era.
Behavior regularized offline reinforcement learning
Wu, Y., Tucker, G., and Nachum, O · 2019
Cited alongside, same era.
Large batch optimization for deep learning: Training bert in 76 minutes
You, Y., Li, J., Reddi, S., Hseu, J., Kumar, S., Bhojanapalli, S., Song, X., Demmel, J., Keutzer, K., and Hsieh, C.-J · 2019
Pessimistic bootstrapping for uncertainty-driven offline reinforcement learning
Bai, C., Wang, L., Yang, Z., Deng, Z., Garg, A., Liu, P., and Wang, Z · 2022
Later among the works it cites.
When does return-conditioned supervised learning work for offline reinforcement learning?
Brandfonbrener, D., Bietti, A., Buckman, J., Laroche, R., and Bruna, J · 2022
Later among the works it cites.
Decision s4: Efficient sequence-based rl via state spaces layers
David, S. B., Zimerman, I., Nachmani, E., and Wolf, L · 2022
Later among the works it cites.
Offline q-learning on diverse multi-task data both scales and generalizes
Kumar, A., Agarwal, R., Geng, X., Tucker, G., and Levine, S · 2022
Later among the works it cites.
Multi-game decision transformers
Lee, K.-H., Nachum, O., Yang, M. S., Lee, L., Freeman, D., Guadarrama, S., Fischer, I., Xu, W., Jang, E., Michalewski, H., et al · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
D4rl: Datasets for deep data-driven reinforcement learning
Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S · 2020
Cited alongside, same era.
Conservative q-learning for offline reinforcement learning
Kumar, A., Zhou, A., Tucker, G., and Levine, S · 2020
Cited alongside, same era.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., Kumar, A., Tucker, G., and Fu, J · 2020
Cited alongside, same era.
Critic regularized regression
Wang, Z., Novikov, A., Zolna, K., Merel, J. S., Springenberg, J. T., Reed, S. E., Shahriari, B., Siegel, N., Gulcehre, C., Heess, N., et al · 2020
Cited alongside, same era.
Deep reinforcement learning at the edge of the statistical precipice
Agarwal, R., Schwarzer, M., Castro, P. S., Courville, A., and Bellemare, M. G · 2021
Cited alongside, same era.
Offline rl without off-policy evaluation
Brandfonbrener, D., Whitney, W., Ranganath, R., and Bruna, J · 2021
Cited alongside, same era.
Decision transformer: Reinforcement learning via sequence modeling
Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., Abbeel, P., Srinivas, A., and Mordatch, I · 2021
Cited alongside, same era.
Dara: Dynamics-aware reward augmentation in offline reinforcement learning
Liu, J., Zhang, H., and Wang, D · 2022
Later among the works it cites.
CORL: Research-oriented deep offline reinforcement learning library
Tarasov, D., Nikulin, A., Akimov, D., Kurenkov, V., and Kolesnikov, S · 2022
Later among the works it cites.
Online decision transformer
Zheng, Q., Zhang, A., and Grover, A · 2022
Later among the works it cites.
Decision convformer: Local filtering in metaformer is sufficient for decision making
Kim, J., Lee, S., Kim, W., and Sung, Y · 2023
Later among the works it cites.
Ceil: Generalized contextual imitation learning
Liu, J., He, L., Kang, Y., Zhuang, Z., Wang, D., and Xu, H · 2023
Later among the works it cites.
Wu, Y.-H., Wang, X., and Hamaya, M · 2023
Later among the works it cites.
Future-conditioned unsupervised pretraining for decision transformer
Xie, Z., Lin, Z., Ye, D., Fu, Q., Wei, Y., and Li, S · 2023
Later among the works it cites.
Q-learning decision transformer: Leveraging dynamic programming for conditional sequence modelling in offline rl
Yamagata, T., Khalil, A., and Santos-Rodriguez, R · 2023
Later among the works it cites.
Behavior proximal policy optimization
Zhuang, Z., Lei, K., Liu, J., Wang, D., and Guo, Y · 2023
Later among the works it cites.