Fetching the paper…
Reading the bibliography…
In this study, we investigate the DIstribution Correction Estimation (DICE) methods, an important line of work in offline reinforcement learning (RL) and imitation learning (IL).
Deep residual reinforcement learning
Shangtong Zhang, Wendelin Boehmer, and Shimon Whiteson · 1905
Earlier work this paper cites.
Linear programming and sequential decisions
Alan S Manne · 1960
Earlier work this paper cites.
Alvinn: An autonomous land vehicle in a neural network
Dean A Pomerleau · 1989
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Leemon Baird · 1995
Earlier work this paper cites.
Introduction to reinforcement learning
Richard S Sutton, Andrew G Barto, et al · 1998
Earlier work this paper cites.
Convex optimization
Stephen Boyd, Stephen P Boyd, and Lieven Vandenberghe · 2004
Earlier work this paper cites.
Numerical mathematics and computing
E Ward Cheney and David R Kincaid · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Lei Jimmy Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton · 2016
Earlier work this paper cites.
Neural photo editing with introspective adversarial networks
Andrew Brock, Theodore Lim, James M Ritchie, and Nick Weston · 2016
Earlier work this paper cites.
Learning from conditional distributions via dual embeddings
Bo Dai, Niao He, Yunpeng Pan, Byron Boots, and Le Song · 2017
Earlier work this paper cites.
Improved training of wasserstein gans
Ishaan Gulrajani, Faruk Ahmed, Martín Arjovsky, Vincent Dumoulin, and Aaron C. Courville · 2017
Earlier work this paper cites.
Reinforcement learning with deep energy-based policies
Tuomas Haarnoja, Haoran Tang, Pieter Abbeel, and Sergey Levine · 2017
Earlier work this paper cites.
Td learning with constrained gradients
Ishan Durugkar and Peter Stone · 2018
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke van Hoof, and David Meger · 2018
Earlier work this paper cites.
Off-policy deep reinforcement learning without exploration
Scott Fujimoto, David Meger, and Doina Precup · 2019
Earlier work this paper cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
Aviral Kumar, Justin Fu, Matthew Soh, George Tucker, and Sergey Levine · 2019
Earlier work this paper cites.
Behavior regularized offline reinforcement learning
Yifan Wu, George Tucker, and Ofir Nachum · 2019
Cited alongside, same era.
BAIL: best-action imitation learning for batch deep reinforcement learning
Xinyue Chen, Zijian Zhou, Zheng Wang, Che Wang, Yanqiu Wu, and Keith W. Ross · 2020
Cited alongside, same era.
D4rl: Datasets for deep data-driven reinforcement learning
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine · 2020
Cited alongside, same era.
Imitation learning via off-policy distribution matching
Ilya Kostrikov, Ofir Nachum, and Jonathan Tompson · 2020
Cited alongside, same era.
Conservative q-learning for offline reinforcement learning
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine · 2020
Cited alongside, same era.
Demodice: Offline imitation learning with supplementary imperfect demonstrations
Geon-Hyeong Kim, Seokin Seo, Jongmin Lee, Wonseok Jeon, HyeongJoo Hwang, Hongseok Yang, and Kee-Eung Kim · 2021
Later among the works it cites.
Dr3: Value-based deep reinforcement learning requires explicit regularization
Aviral Kumar, Rishabh Agarwal, Tengyu Ma, Aaron Courville, George Tucker, and Sergey Levine · 2021
Later among the works it cites.
Optidice: Offline policy optimization via stationary distribution correction estimation
Jongmin Lee, Wonseok Jeon, Byungjun Lee, Joelle Pineau, and Kee-Eung Kim · 2021
Later among the works it cites.
Model selection for offline reinforcement learning: Practical considerations for healthcare settings
Shengpu Tang and Jenna Wiens · 2021
Later among the works it cites.
Offline reinforcement learning with soft behavior regularization
Haoran Xu, Xianyuan Zhan, Jianxiong Li, and Honglei Yin · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ofir Nachum and Bo Dai · 2020
Cited alongside, same era.
Accelerating online reinforcement learning with offline datasets
Ashvin Nair, Murtaza Dalal, Abhishek Gupta, and Sergey Levine · 2020
Cited alongside, same era.
Toward the fundamental limits of imitation learning
Nived Rajaraman, Lin Yang, Jiantao Jiao, and Kannan Ramchandran · 2020
Cited alongside, same era.
Critic regularized regression
Ziyu Wang, Alexander Novikov, Konrad Zolna, Josh Merel, Jost Tobias Springenberg, Scott E. Reed, Bobak Shahriari, Noah Y. Siegel, Çaglar Gülçehre, Nicolas Heess, and Nando de Freitas · 2020
Cited alongside, same era.
Off-policy evaluation via the regularized lagrangian
Mengjiao Yang, Ofir Nachum, Bo Dai, Lihong Li, and Dale Schuurmans · 2020
Cited alongside, same era.
Gradientdice: Rethinking generalized offline estimation of stationary values
Shangtong Zhang, Bo Liu, and Shimon Whiteson · 2020
Cited alongside, same era.
Latent action space for offline reinforcement learning
Wenxuan Zhou, Sujay Bajracharya, and David Held · 2020
Cited alongside, same era.
Generalized out-of-distribution detection: A survey
Jingkang Yang, Kaiyang Zhou, Yixuan Li, and Ziwei Liu · 2021
Later among the works it cites.
DR3: value-based deep reinforcement learning requires explicit regularization
Aviral Kumar, Rishabh Agarwal, Tengyu Ma, Aaron C. Courville, George Tucker, and Sergey Levine · 2022
Later among the works it cites.
Jongmin Lee, Cosmin Paduraru, Daniel J Mankowitz, Nicolas Heess, Doina Precup, Kee-Eung Kim, and Arthur Guez · 2022
Later among the works it cites.
Smodice: Versatile offline imitation learning via state occupancy matching
Yecheng Jason Ma, Andrew Shen, Dinesh Jayaraman, and Osbert Bastani · 2022
Later among the works it cites.
When to trust your simulator: Dynamics-aware hybrid offline-and-online reinforcement learning
Haoyi Niu, Yiwen Qiu, Ming Li, Guyue Zhou, Jianming HU, Xianyuan Zhan, et al · 2022
Later among the works it cites.
Deepthermal: Combustion optimization for thermal power generating units using offline reinforcement learning
Xianyuan Zhan, Haoran Xu, Yue Zhang, Xiangyu Zhu, Honglei Yin, and Yu Zheng · 2022
Later among the works it cites.
State deviation correction for offline reinforcement learning
Hongchang Zhang, Jianzhun Shao, Yuhang Jiang, Shuncheng He, Guanwen Zhang, and Xiangyang Ji · 2022
Later among the works it cites.
Look beneath the surface: Exploiting fundamental symmetry for sample-efficient offline rl
Peng Cheng, Xianyuan Zhan, Zhihao Wu, Wenjia Zhang, Shoucheng Song, Han Wang, Youfang Lin, and Li Jiang · 2023
Later among the works it cites.
Extreme q-learning: Maxent rl without entropy
Divyansh Garg, Joey Hejna, Matthieu Geist, and Stefano Ermon · 2023
Later among the works it cites.
Harshit Sikchi, Amy Zhang, and Scott Niekum · 2023
Later among the works it cites.
Offline multi-agent reinforcement learning with implicit global-to-local value regularization
Xiangsen Wang, Haoran Xu, Yinan Zheng, and Xianyuan Zhan · 2023
Later among the works it cites.
Offline rl with no ood actions: In-sample learning via implicit value regularization
Haoran Xu, Li Jiang, Jianxiong Li, Zhuoran Yang, Zhaoran Wang, Victor Wai Kin Chan, and Xianyuan Zhan · 2023
Later among the works it cites.
Safe offline reinforcement learning with feasibility-guided diffusion model
Yinan Zheng, Jianxiong Li, Dongjie Yu, Yujie Yang, Shengbo Eben Li, Xianyuan Zhan, and Jingjing Liu · 2024
Closest in time.