Fetching the paper…
Reading the bibliography…
Constrained policy search (CPS) is a fundamental problem in offline reinforcement learning, which is generally solved by advantage weighted regression (AWR).
Convex analysis
Rockafellar, R · 1970
Earlier work this paper cites.
A convex analytic approach to markov decision processes
Borkar, V. S · 1988
Earlier work this paper cites.
Convex optimization
Boyd, S. P. and Vandenberghe, L · 2004
Earlier work this paper cites.
Relative entropy policy search
Peters, J., Mulling, K., and Altun, Y · 2010
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S · 2015
Earlier work this paper cites.
Deep reinforcement learning for automated radiation adaptation in lung cancer
Tseng, H.-H., Luo, Y., Cui, S., Chien, J.-T., Ten Haken, R. K., and Naqa, I. E · 2017
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
Fujimoto, S., Hoof, H., and Meger, D · 2018
Earlier work this paper cites.
Deep imitative models for flexible inference, planning, and control
Rhinehart, N., McAllister, R., and Levine, S · 2018
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G · 2018
Earlier work this paper cites.
Off-policy deep reinforcement learning without exploration
Fujimoto, S., Meger, D., and Precup, D · 2019
Earlier work this paper cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
Kumar, A., Fu, J., Soh, M., Tucker, G., and Levine, S · 2019
Earlier work this paper cites.
Constrained reinforcement learning has zero duality gap
Paternain, S., Chamon, L., Calvo-Fullana, M., and Ribeiro, A · 2019
Earlier work this paper cites.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
Peng, X. B., Kumar, A., Zhang, G., and Levine, S · 2019
Earlier work this paper cites.
Generative modeling by estimating gradients of the data distribution
Song, Y. and Ermon, S · 2019
Cited alongside, same era.
Behavior regularized offline reinforcement learning
Wu, Y., Tucker, G., and Nachum, O · 2019
Cited alongside, same era.
Bail: Best-action imitation learning for batch deep reinforcement learning
Chen, X., Zhou, Z., Wang, Z., Wang, C., Wu, Y., and Ross, K · 2020
Cited alongside, same era.
D4rl: Datasets for deep data-driven reinforcement learning
Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S · 2020
Cited alongside, same era.
Emaq: Expected-max q-learning operator for simple yet effective offline and online rl
Ghasemipour, S. K. S., Schuurmans, D., and Gu, S. S · 2020
Cited alongside, same era.
A minimalist approach to offline reinforcement learning
Fujimoto, S. and Gu, S. S · 2021
Later among the works it cites.
Classifier-free diffusion guidance
Ho, J. and Salimans, T · 2021
Later among the works it cites.
Offline reinforcement learning with value-based episodic memory
Ma, X., Yang, Y., Hu, H., Liu, Q., Yang, J., Zhang, C., Zhao, Q., and Liang, B · 2021
Later among the works it cites.
Learning when-to-treat policies
Nie, X., Brunskill, E., and Wager, S · 2021
Later among the works it cites.
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
Rashidinejad, P., Zhu, B., Ma, C., Jiao, J., and Russell, S · 2021
Later among the works it cites.
Tackling the generative learning trilemma with denoising diffusion gans
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Cited alongside, same era.
Morel: Model-based offline reinforcement learning
Kidambi, R., Rajeswaran, A., Netrapalli, P., and Joachims, T · 2020
Cited alongside, same era.
Conservative q-learning for offline reinforcement learning
Kumar, A., Zhou, A., Tucker, G., and Levine, S · 2020
Cited alongside, same era.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., Kumar, A., Tucker, G., and Fu, J · 2020
Cited alongside, same era.
Awac: Accelerating online reinforcement learning with offline datasets
Nair, A., Gupta, A., Dalal, M., and Levine, S · 2020
Cited alongside, same era.
Score-based generative modeling through stochastic differential equations
Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B · 2020
Cited alongside, same era.
Critic regularized regression
Wang, Z., Novikov, A., Zolna, K., Merel, J. S., Springenberg, J. T., Reed, S. E., Shahriari, B., Siegel, N., Gulcehre, C., and Heess, N · 2020
Cited alongside, same era.
Xiao, Z., Kreis, K., and Vahdat, A · 2021
Later among the works it cites.
Is conditional generative modeling all you need for decision making?
Ajay, A., Du, Y., Gupta, A., Tenenbaum, J. B., Jaakkola, T. S., and Agrawal, P · 2022
Later among the works it cites.
Offline reinforcement learning via high-fidelity generative behavior modeling
Chen, H., Lu, C., Ying, C., Su, H., and Zhu, J · 2022
Later among the works it cites.
Planning with diffusion for flexible behavior synthesis
Janner, M., Du, Y., Tenenbaum, J. B., and Levine, S · 2022
Later among the works it cites.
Behavior transformers: Cloning k k modes with one stone
Shafiullah, N. M., Cui, Z., Altanzaya, A. A., and Pinto, L · 2022
Later among the works it cites.
Diffusion policies as an expressive policy class for offline reinforcement learning
Wang, Z., Hunt, J. J., and Zhou, M · 2022
Later among the works it cites.
Idql: Implicit q-learning as an actor-critic method with diffusion policies
Hansen-Estruch, P., Kostrikov, I., Janner, M., Kuba, J. G., and Levine, S · 2023
Closest in time.
Adaptdiffuser: Diffusion models as adaptive self-evolving planners
Liang, Z., Mu, Y., Ding, M., Ni, F., Tomizuka, M., and Luo, P · 2023
Closest in time.
Lu, C., Chen, H., Chen, J., Su, H., Li, C., and Zhu, J · 2023
Closest in time.
Imitating human behaviour with diffusion models
Pearce, T., Rashid, T., Kanervisto, A., Bignell, D., Sun, M., Georgescu, R., Macua, S. V., Tan, S. Z., Momennejad, I., Hofmann, K., et al · 2023
Closest in time.
Fighting uncertainty with gradients: Offline reinforcement learning via diffusion score matching
Suh, H., Chou, G., Dai, H., Yang, L., Gupta, A., and Tedrake, R · 2023
Closest in time.