Fetching the paper…
Reading the bibliography…
Behavior regularization, which constrains the policy to stay close to some behavior policy, is widely used in offline reinforcement learning (RL) to manage the risk of hazardous exploitation of unseen actions.
Stochastic differential equations: An introduction with applications
Oksendal, B · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Earlier work this paper cites.
Reinforcement learning with deep energy-based policies
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Earlier work this paper cites.
Soft actor-critic for discrete action settings
Christodoulou, P · 2019
Earlier work this paper cites.
Off-policy deep reinforcement learning without exploration
Fujimoto, S., Meger, D., and Precup, D · 2019
Earlier work this paper cites.
Stabilizing off-policy Q-learning via bootstrapping error reduction
Kumar, A., Fu, J., Soh, M., Tucker, G., and Levine, S · 2019
Earlier work this paper cites.
Grandmaster level in starcraft II using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., Oh, J., Horgan, D., Kroiss, M., Danihelka, I., Huang, A., Sifre, L., Cai, T., Agapiou, J. P., Jaderberg, M., Vezhnevets, A. S., Leblond, R., Pohlen, T., Dalibard, V., Budden, D., Sulsky, Y., Molloy, J., Paine, T. L., Gülçehre, Ç., Wang, Z., Pfaff, T., Wu, Y., Ring, R., Yogatama, D., Wünsch, D., McKinney, K., Smith, O., Schaul, T., Lillicrap, T. P., Kavukcuoglu, K., Hassabis, D., Apps, C., and Silver, D · 2019
Earlier work this paper cites.
Behavior regularized offline reinforcement learning
Wu, Y., Tucker, G., and Nachum, O · 2019
Earlier work this paper cites.
D4RL: Datasets for deep data-driven reinforcement learning
Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., Kumar, A., Tucker, G., and Fu, J · 2020
Earlier work this paper cites.
Awac: Accelerating online reinforcement learning with offline datasets
Nair, A., Gupta, A., Dalal, M., and Levine, S · 2020
Earlier work this paper cites.
Denoising diffusion implicit models
Song, J., Meng, C., and Ermon, S · 2020
Earlier work this paper cites.
MOPO: Model-based offline policy optimization
Yu, T., Thomas, G., Yu, L., Ermon, S., Zou, J. Y., Levine, S., Finn, C., and Ma, T · 2020
Earlier work this paper cites.
Uncertainty-based offline reinforcement learning with diversified Q-ensemble
An, G., Moon, S., Kim, J., and Song, H. O · 2021
Cited alongside, same era.
Decision transformer: Reinforcement learning via sequence modeling
Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., Abbeel, P., Srinivas, A., and Mordatch, I · 2021
Cited alongside, same era.
A minimalist approach to offline reinforcement learning
Fujimoto, S. and Gu, S. S · 2021
Cited alongside, same era.
A workflow for offline model-free robotic reinforcement learning
Kumar, A., Singh, A., Tian, S., Finn, C., and Levine, S · 2021
Cited alongside, same era.
Score-based generative modeling through stochastic differential equations
Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B · 2021
Cited alongside, same era.
Pessimistic bootstrapping for uncertainty-driven offline reinforcement learning
Model-bellman inconsistency for model-based offline reinforcement learning
Sun, Y., Zhang, J., Jia, C., Lin, H., Ye, J., and Yu, Y · 2023
Later among the works it cites.
Revisiting the minimalist approach to offline reinforcement learning
Tarasov, D., Kurenkov, V., Nikulin, A., and Kolesnikov, S · 2023
Later among the works it cites.
Diffusion policies as an expressive policy class for offline reinforcement learning
Wang, Z., Hunt, J. J., and Zhou, M · 2023
Later among the works it cites.
Offline RL with no OOD actions: In-sample learning via implicit value regularization
Xu, H., Jiang, L., Li, J., Yang, Z., Wang, Z., Chan, W. K. V., and Zhan, X · 2023
Later among the works it cites.
Entropy-regularized diffusion policy with Q-ensembles for offline reinforcement learning
Zhang, R., Luo, Z., Sjölund, J., Schön, T. B., and Mattsson, P · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bai, C., Wang, L., Yang, Z., Deng, Z., Garg, A., Liu, P., and Wang, Z · 2022
Cited alongside, same era.
Planning with diffusion for flexible behavior synthesis
Janner, M., Du, Y., Tenenbaum, J. B., and Levine, S · 2022
Cited alongside, same era.
Offline reinforcement learning with implicit Q-learning
Kostrikov, I., Nair, A., and Levine, S · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P. F., Leike, J., and Lowe, R · 2022
Cited alongside, same era.
Is conditional generative modeling all you need for decision making?
Ajay, A., Du, Y., Gupta, A., Tenenbaum, J. B., Jaakkola, T. S., and Agrawal, P · 2023
Cited alongside, same era.
Offline reinforcement learning via high-fidelity generative behavior modeling
Chen, H., Lu, C., Ying, C., Su, H., and Zhu, J · 2023
Cited alongside, same era.
Extreme Q-learning: Maxent RL without entropy
Garg, D., Hejna, J., Geist, M., and Ermon, S · 2023
Cited alongside, same era.
Iterated denoising energy matching for sampling from boltzmann densities
Akhound-Sadegh, T., Rector-Brooks, J., Bose, A. J., Mittal, S., Lemos, P., Liu, C., Sendera, M., Ravanbakhsh, S., Gidel, G., Bengio, Y., Malkin, N., and Tong, A · 2024
Later among the works it cites.
Diffusion for world modeling: Visual details matter in atari
Alonso, E., Jelley, A., Micheli, V., Kanervisto, A., Storkey, A., Pearce, T., and Fleuret, F · 2024
Later among the works it cites.
Fang, L., Liu, R., Zhang, J., Wang, W., and Jing, B.-Y · 2024
Later among the works it cites.
MINDE: Mutual information neural diffusion estimation
Franzese, G., BOUNOUA, M., and Michiardi, P · 2024
Later among the works it cites.
ACT: Empowering decision transformer with dynamic programming via advantage conditioning
Gao, C.-X., Wu, C., Cao, M., Kong, R., Zhang, Z., and Yu, Y · 2024
Later among the works it cites.
Policy rehearsing: Training generalizable policies for reinforcement learning
Jia, C., Gao, C., Yin, H., Zhang, F., Chen, X.-H., Xu, T., Yuan, L., Zhang, Z., Zhou, Z.-H., and Yu, Y · 2024
Later among the works it cites.
Reward-consistent dynamics models are strongly generalizable for offline reinforcement learning
Luo, F., Xu, T., Cao, X., and Yu, Y · 2024
Later among the works it cites.
Diffusion-dice: In-sample diffusion guidance for offline reinforcement learning
Mao, L., Xu, H., Zhan, X., Zhang, W., and Zhang, A · 2024
Later among the works it cites.
Diffusion policy policy optimization
Ren, A. Z., Lidard, J., Ankile, L. L., Simeonov, A., Agrawal, P., Majumdar, A., Burchfiel, B., Dai, H., and Simchowitz, M · 2024
Later among the works it cites.
Diffusion spectral representation for reinforcement learning
Shribak, D., Gao, C.-X., Li, Y., Xiao, C., and Dai, B · 2024
Later among the works it cites.
Diffusion actor-critic with entropy regulator
Wang, Y., Wang, L., Jiang, Y., Zou, W., Liu, T., Song, X., Wang, W., Xiao, L., Wu, J., Duan, J., et al · 2024
Later among the works it cites.