Fetching the paper…
Reading the bibliography…
Maximum entropy reinforcement learning (MaxEnt-RL) has become the standard approach to RL due to its beneficial exploration properties.
Reverse-time diffusion equation models
Anderson, B. D · 1982
Earlier work this paper cites.
Time reversal of diffusions
Haussmann, U. G. and Pardoux, E · 1986
Earlier work this paper cites.
A stochastic control approach to reciprocal diffusion processes
Dai Pra, P · 1991
Earlier work this paper cites.
Elements of information theory
Cover, T. M · 1999
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 1999
Earlier work this paper cites.
An auxiliary variational method
Agakov, F. V. and Barber, D · 2004
Earlier work this paper cites.
Estimation of non-normalized statistical models by score matching
Hyvärinen, A. and Dayan, P · 2005
Earlier work this paper cites.
Sequential monte carlo samplers
Del Moral, P., Doucet, A., and Jasra, A · 2006
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Ziebart, B. D., Maas, A. L., Bagnell, J. A., Dey, A. K., et al · 2008
Earlier work this paper cites.
Robot trajectory optimization using approximate inference
Toussaint, M · 2009
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
Ziebart, B. D · 2010
Earlier work this paper cites.
A connection between score matching and denoising autoencoders
Vincent, P · 2011
Earlier work this paper cites.
Bayesian learning via stochastic gradient langevin dynamics
Welling, M. and Teh, Y. W · 2011
Earlier work this paper cites.
Batch reinforcement learning
Lange, S., Gabel, T., and Riedmiller, M · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P · 2013
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S · 2015
Earlier work this paper cites.
The variational gaussian process
Tran, D., Ranganath, R., and Blei, D. M · 2015
Earlier work this paper cites.
Openai gym, 2016
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Earlier work this paper cites.
Auxiliary deep generative models
Maaløe, L., Sønderby, C. K., Sønderby, S. K., and Winther, O · 2016
Earlier work this paper cites.
Hierarchical variational models
Ranganath, R., Tran, D., and Blei, D · 2016
Earlier work this paper cites.
Learning to draw samples: With application to amortized mle for generative adversarial learning
Wang, D. and Liu, Q · 2016
Earlier work this paper cites.
A distributional perspective on reinforcement learning
Bellemare, M. G., Dabney, W., and Munos, R · 2017
Earlier work this paper cites.
Reinforcement learning with deep energy-based policies
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S · 2017
Earlier work this paper cites.
A unified view of entropy-regularized markov decision processes
Neu, G., Jonsson, A., and Gómez, V · 2017
Cited alongside, same era.
Efficient gradient-free variational inference using policy search
Arenz, O., Neumann, G., and Zhong, M · 2018
Cited alongside, same era.
Applied stochastic differential equations , volume 10
Särkkä, S. and Solin, A · 2019
Cited alongside, same era.
Theoretical guarantees for sampling and inference in generative models with latent diffusions
Tzen, B. and Raginsky, M · 2019
Cited alongside, same era.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Cited alongside, same era.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
On the theory of risk-aware agents: Bridging actor-critic and economics
Nauman, M. and Cygan, M · 2023
Later among the works it cites.
Goal conditioned imitation learning using score-based diffusion policies
Reuss, M., Li, M., Jia, X., and Lioutikov, R · 2023
Later among the works it cites.
Bayesian learning via neural schrödinger–föllmer flows
Vargas, F., Ovsianas, A., Fernandes, D., Girolami, M., Lawrence, N. D., and Nüsken, N · 2023
Later among the works it cites.
Diffusion policies as an expressive policy class for offline reinforcement learning
Wang, Z., Hunt, J. J., and Zhou, M · 2023
Later among the works it cites.
Policy representation via diffusion probability model for reinforcement learning
Yang, L., Huang, Z., Lei, F. h., Zhong, Y., Yang, Y., Fang, C., Wen, S., Zhou, B., and Lin, Z · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Levine, S., Kumar, A., Tucker, G., and Fu, J · 2020
Cited alongside, same era.
Dynamical theories of Brownian motion , volume 101
Nelson, E · 2020
Cited alongside, same era.
dm_control: Software and tasks for continuous control
Tunyasuvunakool, S., Muldal, A., Doron, Y., Liu, S., Bohez, S., Merel, J., Erez, T., Lillicrap, T., Heess, N., and Tassa, Y · 2020
Cited alongside, same era.
Deep reinforcement learning at the edge of the statistical precipice
Agarwal, R., Schwarzer, M., Castro, P. S., Courville, A. C., and Bellemare, M · 2021
Cited alongside, same era.
Schrödinger-föllmer sampler: sampling without ergodicity
Huang, J., Jiao, Y., Kang, L., Liao, X., Liu, J., and Liu, Y · 2021
Cited alongside, same era.
Score-based generative modeling through stochastic differential equations
Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B · 2021
Cited alongside, same era.
Path integral sampler: a stochastic control approach for sampling
Zhang, Q. and Chen, Y · 2021
Cited alongside, same era.
Zhang, D., Chen, R. T., Liu, C.-H., Courville, A., and Bengio, Y · 2023
Later among the works it cites.
Iterated denoising energy matching for sampling from boltzmann densities
Akhound-Sadegh, T., Rector-Brooks, J., Bose, A. J., Mittal, S., Lemos, P., Liu, C.-H., Sendera, M., Ravanbakhsh, S., Gidel, G., Bengio, Y., et al · 2024
Later among the works it cites.
Crossq: Batch normalization in deep reinforcement learning for greater sample efficiency and simplicity
Bhatt, A., Palenicek, D., Belousov, B., Argus, M., Amiranashvili, A., Brox, T., and Peters, J · 2024
Later among the works it cites.
Sequential controlled langevin diffusions
Chen, J., Richter, L., Berner, J., Blessing, D., Neumann, G., and Anandkumar, A · 2024
Later among the works it cites.
Diffusion-based reinforcement learning via q-weighted variational policy optimization
Ding, S., Hu, K., Zhang, Z., Ren, K., Zhang, W., Yu, J., Wang, J., and Shi, Y · 2024
Later among the works it cites.
Consistency models as a rich and efficient policy class for reinforcement learning
Ding, Z. and Jin, C · 2024
Later among the works it cites.
Learning multimodal behaviors from scratch with diffusion policy gradient
Li, Z., Krohn, R., Chen, T., Ajay, A., Agrawal, P., and Chalvatzaki, G · 2024
Later among the works it cites.
Diffusion-dice: In-sample diffusion guidance for offline reinforcement learning
Mao, L., Xu, H., Zhan, X., Zhang, W., and Zhang, A · 2024
Later among the works it cites.
Bigger, regularized, optimistic: scaling for compute and sample efficient continuous control
Nauman, M., Ostaszewski, M., Jankowski, K., Miłoś, P., and Cygan, M · 2024
Later among the works it cites.
Learned reference-based diffusion sampling for multi-modal distributions
Noble, M., Grenioux, L., Gabrié, M., and Durmus, A. O · 2024
Later among the works it cites.
Transport meets variational inference: Controlled monte carlo diffusions
Nusken, N., Vargas, F., Padhy, S., and Blessing, D · 2024
Later among the works it cites.
Learning a diffusion model policy from rewards via q-score matching
Psenka, M., Escontrela, A., Abbeel, P., and Ma, Y · 2024
Later among the works it cites.
Diffusion actor-critic with entropy regulator
Wang, Y., Wang, L., Jiang, Y., Zou, W., Liu, T., Song, X., Wang, W., Xiao, L., WU, J., Duan, J., and Li, S. E · 2024
Later among the works it cites.
Variational distillation of diffusion policies into mixture of experts
Zhou, H., Blessing, D., Li, G., Celik, O., Jia, X., Neumann, G., and Lioutikov, R · 2024
Later among the works it cites.
Sequential controlled langevin diffusions
Chen, J., Richter, L., Berner, J., Blessing, D., Neumann, G., and Anandkumar, A · 2025
Closest in time.
Diffusion actor-critic: Formulating constrained policy iteration as diffusion noise regression for offline reinforcement learning
Fang, L., Liu, R., Zhang, J., Wang, W., and Jing, B · 2025
Closest in time.
Langevin soft actor-critic: Efficient exploration through uncertainty-driven critic learning
Ishfaq, H., Wang, G., Islam, S. N., and Precup, D · 2025
Closest in time.
TOP-ERL: Transformer-based off-policy episodic reinforcement learning
Li, G., Tian, D., Zhou, H., Jiang, X., Lioutikov, R., and Neumann, G · 2025
Closest in time.