Fetching the paper…
Reading the bibliography…
Diffusion models have demonstrated their powerful generative capability in many tasks, with great potential to serve as a paradigm for offline reinforcement learning.
Learning to predict by the methods of temporal differences
Sutton, R. S · 1988
Earlier work this paper cites.
Estimation of non-normalized statistical models by score matching
Hyvärinen, A · 2005
Earlier work this paper cites.
A tutorial on energy-based learning
LeCun, Y., Chopra, S., Hadsell, R., Huang, F. J., and et al · 2006
Earlier work this paper cites.
The fast research interface for the kuka lightweight robot
Schreiber, G., Stemmer, A., and Bischoff, R · 2010
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
On the theory of stochastic processes, with particular reference to applications
Feller, W · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Ronneberger, O., Fischer, P., and Brox, T · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S · 2015
Earlier work this paper cites.
Model predictive path integral control using covariance variable importance sampling
Williams, G., Aldrich, A., and Theodorou, E · 2015
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Earlier work this paper cites.
Group normalization
Wu, Y. and He, K · 2018
Earlier work this paper cites.
Implicit generation and generalization in energy-based models
Du, Y. and Mordatch, I · 2019
Earlier work this paper cites.
Learning non-convergent non-persistent short-run MCMC toward energy-based model
Nijkamp, E., Hill, M., Zhu, S.-C., and Wu, Y. N · 2019
Earlier work this paper cites.
Generative modeling by estimating gradients of the data distribution
Song, Y. and Ermon, S · 2019
Earlier work this paper cites.
Behavior regularized offline reinforcement learning
Wu, Y., Tucker, G., and Nachum, O · 2019
Earlier work this paper cites.
An optimistic perspective on offline reinforcement learning
Agarwal, R., Schuurmans, D., and Norouzi, M · 2020
Cited alongside, same era.
Compositional visual generation with energy based models
Du, Y., Li, S., and Mordatch, I · 2020
Cited alongside, same era.
D4rl: Datasets for deep data-driven reinforcement learning
Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S · 2020
Cited alongside, same era.
Pddlstream: Integrating symbolic planners and blackbox samplers via optimistic adaptive planning
Garrett, C. R., Lozano-Pérez, T., and Kaelbling, L. P · 2020
Cited alongside, same era.
Learning the stein discrepancy for training and evaluating energy-based models without sampling
Grathwohl, W., Wang, K.-C., Jacobsen, J.-H., Duvenaud, D., and Zemel, R · 2020
Cited alongside, same era.
Offline reinforcement learning as one big sequence modeling problem
Janner, M., Li, Q., and Levine, S · 2021
Later among the works it cites.
Denoising diffusion implicit models
Song, J., Meng, C., and Ermon, S · 2021
Later among the works it cites.
Dirac delta function of matrix argument
Zhang, L · 2021
Later among the works it cites.
3D shape generation and completion through point-voxel diffusion
Zhou, L., Du, Y., and Wu, J · 2021
Later among the works it cites.
Minedojo: Building open-ended embodied agents with internet-scale knowledge
Fan, L., Wang, G., Jiang, Y., Mandlekar, A., Yang, Y., Zhu, H., Tang, A., Huang, D.-A., Zhu, Y., and Anandkumar, A · 2022
Later among the works it cites.
Planning with diffusion for flexible behavior synthesis
Janner, M., Du, Y., Tenenbaum, J., and Levine, S · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ho, J., Jain, A., and Abbeel, P · 2020
Cited alongside, same era.
Morel: Model-based offline reinforcement learning
Kidambi, R., Rajeswaran, A., Netrapalli, P., and Joachims, T · 2020
Cited alongside, same era.
Conservative q-learning for offline reinforcement learning
Kumar, A., Zhou, A., Tucker, G., and Levine, S · 2020
Cited alongside, same era.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., Kumar, A., Tucker, G., and Fu, J · 2020
Cited alongside, same era.
Mish: A self regularized non-monotonic activation function
Misra, D · 2020
Cited alongside, same era.
Keep doing what worked: Behavior modelling priors for offline reinforcement learning
Siegel, N., Springenberg, J. T., Berkenkamp, F., Abdolmaleki, A., Neunert, M., Lampe, T., Hafner, R., Heess, N., and Riedmiller, M · 2020
Cited alongside, same era.
Mopo: Model-based offline policy optimization
Yu, T., Thomas, G., Yu, L., Ermon, S., Zou, J. Y., Levine, S., Finn, C., and Ma, T · 2020
Cited alongside, same era.
Offline reinforcement learning with implicit q-learning
Kostrikov, I., Nair, A., and Levine, S · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Later among the works it cites.
A survey on offline reinforcement learning: Taxonomy, review, and open problems
Prudencio, R. F., Maximo, M. R., and Colombini, E. L · 2022
Later among the works it cites.
Hierarchical text-conditional image generation with clip latents
Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Later among the works it cites.
Bootstrapped transformer for offline reinforcement learning
Wang, K., Zhao, H., Luo, X., Ren, K., Zhang, W., and Li, D · 2022
Later among the works it cites.
A regularized implicit policy for offline reinforcement learning
Yang, S., Wang, Z., Zheng, H., Feng, Y., and Zhou, M · 2022
Later among the works it cites.
Is conditional generative modeling all you need for decision-making?
Ajay, A., Du, Y., Gupta, A., Tenenbaum, J., Jaakkola, T., and Agrawal, P · 2023
Closest in time.
Diffusion policies as an expressive policy class for offline reinforcement learning
Wang, Z., Hunt, J. J., and Zhou, M · 2023
Closest in time.