Fetching the paper…
Reading the bibliography…
The Soft Actor-Critic (SAC) algorithm with a Gaussian policy has become a mainstream implementation for realizing the Maximum Entropy Reinforcement Learning (MaxEnt RL) objective, which incorporates entropy maximization to encourage exploration and enhance policy robustness.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 1999
Earlier work this paper cites.
Pattern Recognition and Machine Learning (Information Science and Statistics)
Bishop, C. M · 2006
Earlier work this paper cites.
Robot trajectory optimization using approximate inference
Toussaint, M · 2009
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
Ziebart, B. D · 2010
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P · 2013
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Variational inference with normalizing flows
Rezende, D. and Mohamed, S · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S · 2015
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Earlier work this paper cites.
Reinforcement learning with deep energy-based policies
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S · 2017
Earlier work this paper cites.
Compiling machine learning programs via high-level tracing
Frostig, R., Johnson, M. J., and Leary, C · 2018
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
Fujimoto, S., Hoof, H., and Meger, D · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Earlier work this paper cites.
Generative modeling by estimating gradients of the data distribution
Song, Y. and Ermon, S · 2019
Earlier work this paper cites.
Generative adversarial networks
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Cited alongside, same era.
How to train your robot with deep reinforcement learning: lessons we have learned
Ibarz, J., Tan, J., Finn, C., Kalakrishnan, M., Pastor, P., and Levine, S · 2021
Cited alongside, same era.
Deep reinforcement learning for autonomous driving: A survey
Kiran, B. R., Sobh, I., Talpaert, V., Mannion, P., Al Sallab, A. A., Yogamani, S., and Pérez, P · 2021
Cited alongside, same era.
Video diffusion models
Ho, J., Salimans, T., Gritsenko, A., Chan, W., Norouzi, M., and Fleet, D. J · 2022
Cited alongside, same era.
Elucidating the design space of diffusion-based generative models
Karras, T., Aittala, M., Aila, T., and Laine, S · 2022
Cited alongside, same era.
Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps
Lu, C., Zhou, Y., Bao, F., Chen, J., Li, C., and Zhu, J · 2022
Dream3d: Zero-shot text-to-3d synthesis using 3d shape prior and text-to-image diffusion models
Xu, J., Wang, X., Cheng, W., Cao, Y.-P., Shan, Y., Qie, X., and Gao, S · 2023
Later among the works it cites.
Policy representation via diffusion probability model for reinforcement learning
Yang, L., Huang, Z., Lei, F., Zhong, Y., Yang, Y., Fang, C., Wen, S., Zhou, B., and Lin, Z · 2023
Later among the works it cites.
Dpm-solver-v3: Improved diffusion ode solver with empirical model statistics
Zheng, K., Lu, C., Chen, J., and Zhu, J · 2023
Later among the works it cites.
Iterated denoising energy matching for sampling from boltzmann densities
Akhound-Sadegh, T., Rector-Brooks, J., Bose, J., Mittal, S., Lemos, P., Liu, C.-H., Sendera, M., Ravanbakhsh, S., Gidel, G., Bengio, Y., et al · 2024
Later among the works it cites.
Crossq: Batch normalization in deep reinforcement learning for greater sample efficiency and simplicity
Bhatt, A., Palenicek, D., Belousov, B., Argus, M., Amiranashvili, A., Brox, T., and Peters, J · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Photorealistic text-to-image diffusion models with deep language understanding
Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E. L., Ghasemipour, K., Gontijo Lopes, R., Karagol Ayan, B., Salimans, T., et al · 2022
Cited alongside, same era.
Is conditional generative modeling all you need for decision making?
Ajay, A., Du, Y., Gupta, A., Tenenbaum, J. B., Jaakkola, T. S., and Agrawal, P · 2023
Cited alongside, same era.
Offline reinforcement learning via high-fidelity generative behavior modeling
Chen, H., Lu, C., Ying, C., Su, H., and Zhu, J · 2023
Cited alongside, same era.
Diffusion policy: Visuomotor policy learning via action diffusion
Chi, C., Xu, Z., Feng, S., Cousineau, E., Du, Y., Burchfiel, B., Tedrake, R., and Song, S · 2023
Cited alongside, same era.
Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models
Huang, R., Huang, J., Yang, D., Ren, Y., Liu, L., Li, M., Ye, Z., Liu, J., Yin, X., and Zhao, Z · 2023
Cited alongside, same era.
Champion-level drone racing using deep reinforcement learning
Kaufmann, E., Bauersfeld, L., Loquercio, A., Müller, M., Koltun, V., and Scaramuzza, D · 2023
Cited alongside, same era.
Later among the works it cites.
Maximum entropy reinforcement learning via energy-based normalizing flow
Chao, C.-H., Feng, C., Sun, W.-F., Lee, C.-K., See, S., and Lee, C.-Y · 2024
Later among the works it cites.
Diffusion-based reinforcement learning via q-weighted variational policy optimization
Ding, S., Hu, K., Zhang, Z., Ren, K., Zhang, W., Yu, J., Wang, J., and Shi, Y · 2024
Later among the works it cites.
Efficient diffusion policies for offline reinforcement learning
Kang, B., Ma, X., Du, C., Pang, T., and Yan, S · 2024
Later among the works it cites.
Learning multimodal behaviors from scratch with diffusion policy gradient
Li, S., Krohn, R., Chen, T., Ajay, A., Agrawal, P., and Chalvatzaki, G · 2024
Later among the works it cites.
Diffusion-dice: In-sample diffusion guidance for offline reinforcement learning
Mao, L., Xu, H., Zhan, X., Zhang, W., and Zhang, A · 2024
Later among the works it cites.
Bigger, regularized, optimistic: scaling for compute and sample efficient continuous control
Nauman, M., Ostaszewski, M., Jankowski, K., Miłoś, P., and Cygan, M · 2024
Later among the works it cites.
Model-based diffusion for trajectory optimization
Pan, C., Yi, Z., Shi, G., and Qu, G · 2024
Later among the works it cites.
Learning a diffusion model policy from rewards via q-score matching
Psenka, M., Escontrela, A., Abbeel, P., and Ma, Y · 2024
Later among the works it cites.
Improved off-policy training of diffusion samplers
Sendera, M., Kim, M., Mittal, S., Lemos, P., Scimeca, L., Rector-Brooks, J., Adam, A., Bengio, Y., and Malkin, N · 2024
Later among the works it cites.
Diffusion actor-critic with entropy regulator
Wang, Y., Wang, L., Jiang, Y., Zou, W., Liu, T., Song, X., Wang, W., Xiao, L., Wu, J., Duan, J., et al · 2024
Later among the works it cites.
Your diffusion model is secretly a noise classifier and benefits from contrastive training
Wu, Y., Luo, Y., Kong, X., Papalexakis, E. E., and Steeg, G. V · 2024
Later among the works it cites.