Fetching the paper…
Reading the bibliography…
One important property of DIstribution Correction Estimation (DICE) methods is that the solution is the optimal stationary distribution ratio between the optimized and data collection policy.
Algaedice: Policy gradient from arbitrary experience
O. Nachum, B. Dai, I. Kostrikov, Y. Chow, L. Li, and D. Schuurmans · 1912
Earlier work this paper cites.
Linear programming and sequential decisions
A. S. Manne · 1960
Earlier work this paper cites.
On measures of entropy and information
A. Rényi · 1961
Earlier work this paper cites.
Introduction to reinforcement learning
R. S. Sutton, A. G. Barto, et al · 1998
Earlier work this paper cites.
Convex optimization
S. Boyd, S. P. Boyd, and L. Vandenberghe · 2004
Earlier work this paper cites.
A connection between score matching and denoising autoencoders
P. Vincent · 2011
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli · 2015
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
S. Fujimoto, H. van Hoof, and D. Meger · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Earlier work this paper cites.
Off-policy deep reinforcement learning without exploration
S. Fujimoto, D. Meger, and D. Precup · 2019
Earlier work this paper cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
A. Kumar, J. Fu, M. Soh, G. Tucker, and S. Levine · 2019
Earlier work this paper cites.
Behavior regularized offline reinforcement learning
Y. Wu, G. Tucker, and O. Nachum · 2019
Earlier work this paper cites.
Gendice: Generalized offline estimation of stationary values
R. Zhang, B. Dai, L. Li, and D. Schuurmans · 2019
Earlier work this paper cites.
D4rl: Datasets for deep data-driven reinforcement learning
J. Fu, A. Kumar, O. Nachum, G. Tucker, and S. Levine · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Earlier work this paper cites.
Imitation learning via off-policy distribution matching
I. Kostrikov, O. Nachum, and J. Tompson · 2020
Earlier work this paper cites.
Conservative q-learning for offline reinforcement learning
A. Kumar, A. Zhou, G. Tucker, and S. Levine · 2020
Earlier work this paper cites.
Reinforcement learning via fenchel-rockafellar duality
O. Nachum and B. Dai · 2020
Cited alongside, same era.
Accelerating online reinforcement learning with offline datasets
A. Nair, M. Dalal, A. Gupta, and S. Levine · 2020
Cited alongside, same era.
Score-based generative modeling through stochastic differential equations
Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole · 2020
Cited alongside, same era.
Critic regularized regression
Z. Wang, A. Novikov, K. Zolna, J. Merel, J. T. Springenberg, S. E. Reed, B. Shahriari, N. Y. Siegel, Ç. Gülçehre, N. Heess, and N. de Freitas · 2020
Cited alongside, same era.
Latent action space for offline reinforcement learning
W. Zhou, S. Bajracharya, and D. Held · 2020
Cited alongside, same era.
Planning with diffusion for flexible behavior synthesis
M. Janner, Y. Du, J. B. Tenenbaum, and S. Levine · 2022
Later among the works it cites.
J. Lee, C. Paduraru, D. J. Mankowitz, N. Heess, D. Precup, K.-E. Kim, and A. Guez · 2022
Later among the works it cites.
When data geometry meets deep function: Generalizing offline reinforcement learning
J. Li, X. Zhan, H. Xu, X. Zhu, J. Liu, and Y.-Q. Zhang · 2022
Later among the works it cites.
Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps
C. Lu, Y. Zhou, F. Bao, J. Chen, C. Li, and J. Zhu · 2022
Later among the works it cites.
Diffusion policies as an expressive policy class for offline reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Uncertainty-based offline reinforcement learning with diversified q-ensemble
G. An, S. Moon, J.-H. Kim, and H. O. Song · 2021
Cited alongside, same era.
Decision transformer: Reinforcement learning via sequence modeling
L. Chen, K. Lu, A. Rajeswaran, K. Lee, A. Grover, M. Laskin, P. Abbeel, A. Srinivas, and I. Mordatch · 2021
Cited alongside, same era.
A minimalist approach to offline reinforcement learning
S. Fujimoto and S. S. Gu · 2021
Cited alongside, same era.
Iq-learn: Inverse soft-q learning for imitation
D. Garg, S. Chakraborty, C. Cundy, J. Song, and S. Ermon · 2021
Cited alongside, same era.
Mt-opt: Continuous multi-task robotic reinforcement learning at scale
D. Kalashnikov, J. Varley, Y. Chebotar, B. Swanson, R. Jonschkowski, C. Finn, S. Levine, and K. Hausman · 2021
Cited alongside, same era.
Demodice: Offline imitation learning with supplementary imperfect demonstrations
G.-H. Kim, S. Seo, J. Lee, W. Jeon, H. Hwang, H. Yang, and K.-E. Kim · 2021
Cited alongside, same era.
Variational diffusion models
D. Kingma, T. Salimans, B. Poole, and J. Ho · 2021
Cited alongside, same era.
Z. Wang, J. J. Hunt, and M. Zhou · 2022
Later among the works it cites.
Deepthermal: Combustion optimization for thermal power generating units using offline reinforcement learning
X. Zhan, H. Xu, Y. Zhang, X. Zhu, H. Yin, and Y. Zheng · 2022
Later among the works it cites.
Consistency models as a rich and efficient policy class for reinforcement learning
Z. Ding and C. Jin · 2023
Later among the works it cites.
Idql: Implicit q-learning as an actor-critic method with diffusion policies
P. Hansen-Estruch, I. Kostrikov, M. Janner, J. G. Kuba, and S. Levine · 2023
Later among the works it cites.
Mind the gap: Offline policy optimization for imperfect rewards
J. Li, X. Hu, H. Xu, J. Liu, X. Zhan, Q.-S. Jia, and Y.-Q. Zhang · 2023
Later among the works it cites.
Contrastive energy prediction for exact energy-guided diffusion sampling in offline reinforcement learning
C. Lu, H. Chen, J. Chen, H. Su, C. Li, and J. Zhu · 2023
Later among the works it cites.
A. Shocher, A. Dravid, Y. Gandelsman, I. Mosseri, M. Rubinstein, and A. A. Efros · 2023
Later among the works it cites.
Dual rl: Unification and new methods for reinforcement and imitation learning
H. Sikchi, Q. Zheng, A. Zhang, and S. Niekum · 2023
Later among the works it cites.
Y. Song, P. Dhariwal, M. Chen, and I. Sutskever · 2023
Later among the works it cites.
Offline multi-agent reinforcement learning with implicit global-to-local value regularization
X. Wang, H. Xu, Y. Zheng, and X. Zhan · 2023
Later among the works it cites.
Offline rl with no ood actions: In-sample learning via implicit value regularization
H. Xu, L. Jiang, J. Li, Z. Yang, Z. Wang, V. W. K. Chan, and X. Zhan · 2023
Later among the works it cites.
Madiff: Offline multi-agent learning with diffusion models
Z. Zhu, M. Liu, L. Mao, B. Kang, M. Xu, Y. Yu, S. Ermon, and W. Zhang · 2023
Later among the works it cites.
Odice: Revealing the mystery of distribution correction estimation via orthogonal-gradient update
L. Mao, H. Xu, W. Zhang, and X. Zhan · 2024
Closest in time.