Fetching the paper…
Reading the bibliography…
Most offline reinforcement learning (RL) algorithms return a target policy maximizing a trade-off between (1) the expected performance gain over the behavior policy that collected the dataset, and (2) the risk stemming from the out-of-distribution-ness of the induced state-action occupancy.
Aviral Kumar, Xue Bin Peng, and Sergey Levine · 1912
Earlier work this paper cites.
Probability inequalities for the sum in sampling without replacement
Robert J Serfling · 1974
Earlier work this paper cites.
Reconstructing human skill with machine learning
Tanja Urbancic · 1994
Earlier work this paper cites.
Constrained Markov Decision Processes
Eitan Altman · 1999
Earlier work this paper cites.
Smote: synthetic minority over-sampling technique
Nitesh V Chawla, Kevin W Bowyer, Lawrence O Hall, and W Philip Kegelmeyer · 2002
Earlier work this paper cites.
Reinforcement learning by reward-weighted regression for operational space control
Jan Peters and Stefan Schaal · 2007
Earlier work this paper cites.
Evolutionary undersampling for classification with imbalanced datasets: Proposals and taxonomy
Salvador García and Francisco Herrera · 2009
Earlier work this paper cites.
Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver · 2015
Earlier work this paper cites.
Safe reinforcement learning
Philip S Thomas · 2015
Earlier work this paper cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Earlier work this paper cites.
Safe policy improvement by minimizing robust baseline regret
Marek Petrik, Mohammad Ghavamzadeh, and Yinlam Chow · 2016
Earlier work this paper cites.
Imbalanced deep learning by minority class incremental rectification
Qi Dong, Shaogang Gong, and Xiatian Zhu · 2018
Earlier work this paper cites.
Distributed prioritized experience replay
Dan Horgan, John Quan, David Budden, Gabriel Barth-Maron, Matteo Hessel, Hado van Hasselt, and David Silver · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Earlier work this paper cites.
Behavioral cloning from observation
Faraz Torabi, Garrett Warnell, and Peter Stone · 2018
Earlier work this paper cites.
Learning imbalanced datasets with label-distribution-aware margin loss
Kaidi Cao, Colin Wei, Adrien Gaidon, Nikos Arechiga, and Tengyu Ma · 2019
Cited alongside, same era.
Information-theoretic considerations in batch reinforcement learning
Jinglin Chen and Nan Jiang · 2019
Cited alongside, same era.
Class-balanced loss based on effective number of samples
Yin Cui, Menglin Jia, Tsung-Yi Lin, Yang Song, and Serge Belongie · 2019
Cited alongside, same era.
Off-policy deep reinforcement learning without exploration
Scott Fujimoto, David Meger, and Doina Precup · 2019
Cited alongside, same era.
Safe policy improvement with baseline bootstrapping
Romain Laroche, Paul Trichelair, and Rémi Tachet des Combes · 2019
Cited alongside, same era.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
Generalized decision transformer for offline hindsight information matching
Hiroki Furuta, Yutaka Matsuo, and Shixiang Shane Gu · 2021
Later among the works it cites.
Demodice: Offline imitation learning with supplementary imperfect demonstrations
Geon-Hyeong Kim, Seokin Seo, Jongmin Lee, Wonseok Jeon, HyeongJoo Hwang, Hongseok Yang, and Kee-Eung Kim · 2021
Later among the works it cites.
Thibault Lahire, Matthieu Geist, and Emmanuel Rachelson · 2021
Later among the works it cites.
Density-based weighting for imbalanced regression
Michael Steininger, Konstantin Kobs, Padraig Davidson, Anna Krause, and Andreas Hotho · 2021
Later among the works it cites.
d3rlpy: An offline deep reinforcement library
Michita Imai Takuma Seno · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xue Bin Peng, Aviral Kumar, Grace Zhang, and Sergey Levine · 2019
Cited alongside, same era.
Reinforcement learning upside down: Don’t predict rewards–just map them to actions
Juergen Schmidhuber · 2019
Cited alongside, same era.
Training agents using upside-down reinforcement learning
Rupesh Kumar Srivastava, Pranav Shyam, Filipe Mutz, Wojciech Jaśkowski, and Jürgen Schmidhuber · 2019
Cited alongside, same era.
The importance of pessimism in fixed-dataset policy optimization
Jacob Buckman, Carles Gelada, and Marc G Bellemare · 2020
Cited alongside, same era.
D4rl: Datasets for deep data-driven reinforcement learning
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine · 2020
Cited alongside, same era.
Safe policy improvement with estimated baseline bootstrapping
Thiago D. Simão, Romain Laroche, and Rémi Tachet des Combes · 2020
Cited alongside, same era.
Deep reinforcement learning at the edge of the statistical precipice
Rishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C Courville, and Marc Bellemare · 2021
Cited alongside, same era.
David Brandfonbrener, Alberto Bietti, Jacob Buckman, Romain Laroche, and Joan Bruna · 2022
Later among the works it cites.
Offline reinforcement learning: Fundamental barriers for value function approximation
Dylan J Foster, Akshay Krishnamurthy, David Simchi-Levi, and Yunzong Xu · 2022
Later among the works it cites.
Topological experience replay
Zhang-Wei Hong, Tao Chen, Yen-Chen Lin, Joni Pajarinen, and Pulkit Agrawal · 2022
Later among the works it cites.
Offline reinforcement learning with implicit q-learning
Ilya Kostrikov, Ashvin Nair, and Sergey Levine · 2022
Later among the works it cites.
About non-Markovian policies stater-action visitation density
Romain Laroche, Jacob Buckman, and Rémi Tachet des Combes · 2022
Later among the works it cites.
Versatile offline imitation from observations and examples via regularized state-occupancy matching
Yecheng Ma, Andrew Shen, Dinesh Jayaraman, and Osbert Bastani · 2022
Later among the works it cites.
Experience replay with likelihood-free importance weights
Samarth Sinha, Jiaming Song, Animesh Garg, and Stefano Ermon · 2022
Later among the works it cites.
The curse of passive data collection in batch reinforcement learning
Chenjun Xiao, Ilbin Lee, Bo Dai, Dale Schuurmans, and Csaba Szepesvari · 2022
Later among the works it cites.
Discriminator-weighted offline imitation learning from suboptimal demonstrations
Haoran Xu, Xianyuan Zhan, Honglei Yin, and Huiling Qin · 2022
Later among the works it cites.
Delving into deep imbalanced regression
Yuzhe Yang, Kaiwen Zha, Ying-Cong Chen, Hao Wang, and Dina Katabi · 2022
Later among the works it cites.