Fetching the paper…
Reading the bibliography…
Offline reinforcement learning (RL) provides a framework for learning decision-making from offline data and therefore constitutes a promising approach for real-world applications as automated driving.
Model predictive heuristic control: Applications to industrial processes
J. Richalet, A. Rault, J.L. Testud, and J. Papon · 1978
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2005
Earlier work this paper cites.
Ngsim interstate 80 freeway dataset
John Halkias and James Colyar · 2006
Earlier work this paper cites.
Model-based reinforcement learning: A survey, 2021
Thomas M. Moerland, Joost Broekens, and Catholijn M. Jonker · 2006
Earlier work this paper cites.
Hyperparameter selection for offline reinforcement learning
Tom Le Paine, Cosmin Paduraru, Andrea Michi, Çaglar Gülçehre, Konrad Zolna, Alexander Novikov, Ziyu Wang, and Nando de Freitas · 2007
Earlier work this paper cites.
Aleatory or epistemic? does it matter?
Armen Der Kiureghian and Ove Ditlevsen · 2009
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stephane Ross, Geoffrey Gordon, and Drew Bagnell · 2011
Earlier work this paper cites.
Auto-Encoding Variational Bayes
Diederik P. Kingma and Max Welling · 2014
Earlier work this paper cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Earlier work this paper cites.
CARLA: An open urban driving simulator
Alexey Dosovitskiy, German Ros, Felipe Codevilla, Antonio Lopez, and Vladlen Koltun · 2017
Earlier work this paper cites.
Simple and scalable predictive uncertainty estimation using deep ensembles
Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell · 2017
Earlier work this paper cites.
Model predictive path integral control: From theory to parallel computation
Grady Williams, Andrew Aldrich, and Evangelos A. Theodorou · 2017
Earlier work this paper cites.
End-to-end driving via conditional imitation learning
Felipe Codevilla, Matthias Müller, Antonio López, Vladlen Koltun, and Alexey Dosovitskiy · 2018
Earlier work this paper cites.
An algorithmic perspective on imitation learning
Takayuki Osa, Joni Pajarinen, Gerhard Neumann, J. Bagnell, Pieter Abbeel, and Jan Peters · 2018
Cited alongside, same era.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 2018
Cited alongside, same era.
Exploring the limitations of behavior cloning for autonomous driving
Felipe Codevilla, Santana Eder, Antonio M. Lopez, and Adrien Gaidon · 2019
Cited alongside, same era.
Causal confusion in imitation learning
Pim de Haan, Dinesh Jayaraman, and Sergey Levine · 2019
Cited alongside, same era.
Model-predictive policy learning with uncertainty regularization for driving in dense traffic
Mikael Henaff, Alfredo Canziani, and Yann LeCun · 2019
Cited alongside, same era.
Plan online, learn offline: Efficient learning and exploration via model-based control
Kendall Lowrey, Aravind Rajeswaran, Sham Kakade, Emanuel Todorov, and Igor Mordatch · 2019
Deep dynamics models for learning dexterous manipulation
Anusha Nagabandi, Kurt Konolige, Sergey Levine, and Vikash Kumar · 2020
Later among the works it cites.
Using online verification to prevent autonomous vehicles from causing accidents
Christian Pek, Stefanie Manzinger, Markus Koschi, and Matthias Althoff · 2020
Later among the works it cites.
Mopo: Model-based offline policy optimization
Tianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon, James Y Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma · 2020
Later among the works it cites.
Model-based offline planning
Arthur Argenson and Gabriel Dulac-Arnold · 2021
Closest in time.
Lucidgames: Online unscented inverse dynamic games for adaptive trajectory prediction and planning
Simon Le Cleac’h, Mac Schwager, and Zachary Manchester · 2021
Closest in time.
Badgr: An autonomous self-supervised learning-based navigation system
Gregory Kahn, Pieter Abbeel, and Sergey Levine · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Dualdice: Behavior-agnostic estimation of discounted stationary distribution corrections
Ofir Nachum, Yinlam Chow, Bo Dai, and Lihong Li · 2019
Cited alongside, same era.
Graph neural networks: A review of methods and applications
Jie Zhou, Ganqu Cui, Shengding Hu, Zhengyan Zhang, Cheng Yang, Zhiyuan Liu, Lifeng Wang, Changcheng Li, and Maosong Sun · 2019
Cited alongside, same era.
Algames: A fast solver for constrained dynamic games
Simon Le Cleac’h, Mac Schwager, and Zachary Manchester · 2020
Cited alongside, same era.
Vectornet: Encoding hd maps and agent dynamics from vectorized representation
Jiyang Gao, Chen Sun, Hang Zhao, Yi Shen, Dragomir Anguelov, Congcong Li, and Cordelia Schmid · 2020
Cited alongside, same era.
Morel: Model-based offline reinforcement learning
Rahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, and Thorsten Joachims · 2020
Cited alongside, same era.
Conservative q-learning for offline reinforcement learning
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine · 2020
Cited alongside, same era.
Closest in time.
Reward (mis)design for autonomous driving
W. Bradley Knox, Alessandro Allievi, Holger Banzhaf, Felix Schmitt, and Peter Stone · 2021
Closest in time.
Deep structured reactive planning
Jerry Liu, Wenyuan Zeng, Raquel Urtasun, and Ersin Yumer · 2021
Closest in time.
Contingencies from observations: Tractable contingency planning with learned behavior models
Nicholas Rhinehart, Jeff He, Charles Packer, Matthew A. Wright, Rowan McAllister, Joseph E. Gonzalez, and Sergey Levine · 2021
Closest in time.
Near-optimal offline reinforcement learning via double variance reduction
Ming Yin, Yu Bai, and Yu-Xiang Wang · 2021
Closest in time.
Combo: Conservative offline model-based policy optimization, 2021
Tianhe Yu, Aviral Kumar, Rafael Rafailov, Aravind Rajeswaran, Sergey Levine, and Chelsea Finn · 2021
Closest in time.
Model-based offline planning with trajectory pruning, 2021
Xianyuan Zhan, Xiangyu Zhu, and Haoran Xu · 2021
Closest in time.