Fetching the paper…
Reading the bibliography…
Offline Reinforcement Learning (RL) enables policy learning without active interactions, making it especially appealing for self-driving tasks.
A survey of pomdp solution techniques
K. P. Murphy · 2000
Earlier work this paper cites.
A method for initialising the k-means clustering algorithm using kd-trees
S. J. Redmond and C. Heneghan · 2007
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
S. Ross, G. Gordon, and D. Bagnell · 2011
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
M. L. Puterman · 2014
Earlier work this paper cites.
Making bertha drive—an autonomous journey on a historic route
J. Ziegler, P. Bender, M. Schreiber, H. Lategahn, T. Strauss, C. Stiller, T. Dang, U. Franke, N. Appenrodt, C. G. Keller, et al · 2014
Earlier work this paper cites.
Weight uncertainty in neural network
C. Blundell, J. Cornebise, K. Kavukcuoglu, and D. Wierstra · 2015
Earlier work this paper cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Y. Gal and Z. Ghahramani · 2016
Earlier work this paper cites.
Carla: An open urban driving simulator
A. Dosovitskiy, G. Ros, F. Codevilla, A. Lopez, and V. Koltun · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
On the quantitative analysis of decoder-based generative models
Y. Wu, Y. Burda, R. Salakhutdinov, and R. Grosse · 2017
Earlier work this paper cites.
Simple and scalable predictive uncertainty estimation using deep ensembles
B. Lakshminarayanan, A. Pritzel, and C. Blundell · 2017
Earlier work this paper cites.
What uncertainties do we need in bayesian deep learning for computer vision?
A. Kendall and Y. Gal · 2017
Earlier work this paper cites.
Advances in variational inference
C. Zhang, J. Bütepage, H. Kjellström, and S. Mandt · 2018
Earlier work this paper cites.
Exponentially weighted imitation learning for batched historical data
Q. Wang, J. Xiong, L. Han, H. Liu, T. Zhang, et al · 2018
Earlier work this paper cites.
Accurate uncertainties for deep learning using calibrated regression
V. Kuleshov, N. Fenner, and S. Ermon · 2018
Earlier work this paper cites.
Model-free deep reinforcement learning for urban autonomous driving
J. Chen, B. Yuan, and M. Tomizuka · 2019
Earlier work this paper cites.
Off-policy deep reinforcement learning without exploration
S. Fujimoto, D. Meger, and D. Precup · 2019
Earlier work this paper cites.
Adversarial training for free!
A. Shafahi, M. Najibi, M. A. Ghiasi, Z. Xu, J. Dickerson, C. Studer, L. S. Davis, G. Taylor, and T. Goldstein · 2019
Cited alongside, same era.
End-to-end interpretable neural motion planner
W. Zeng, W. Luo, S. Suo, A. Sadat, B. Yang, S. Casas, and R. Urtasun · 2019
Cited alongside, same era.
A survey of autonomous driving: Common practices and emerging technologies
E. Yurtsever, J. Lambert, A. Carballo, and K. Takeda · 2020
Cited alongside, same era.
Conservative q-learning for offline reinforcement learning
A. Kumar, A. Zhou, G. Tucker, and S. Levine · 2020
Cited alongside, same era.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Cited alongside, same era.
Carla autonomous driving leaderboard
C. team · 2020
Cited alongside, same era.
Addressing optimism bias in sequence modeling for reinforcement learning
A. R. Villaflor, Z. Huang, S. Pande, J. M. Dolan, and J. Schneider · 2022
Later among the works it cites.
When does return-conditioned supervised learning work for offline reinforcement learning?
D. Brandfonbrener, A. Bietti, J. Buckman, R. Laroche, and J. Bruna · 2022
Later among the works it cites.
You can’t count on luck: Why decision transformers fail in stochastic environments
K. Paster, S. McIlraith, and J. Ba · 2022
Later among the works it cites.
Dichotomy of control: Separating what you can control from what you cannot
S. Yang, D. Schuurmans, P. Abbeel, and O. Nachum · 2022
Later among the works it cites.
M2i: From factored marginal trajectory prediction to interactive prediction
Q. Sun, X. Huang, J. Gu, B. C. Williams, and H. Zhao · 2022
Later among the works it cites.
Generalized decision transformer for offline hindsight information matching
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
nuplan: A closed-loop ml-based planning benchmark for autonomous vehicles
H. Caesar, J. Kabzan, K. S. Tan, W. K. Fong, E. Wolff, A. Lang, L. Fletcher, O. Beijbom, and S. Omari · 2021
Cited alongside, same era.
Umbrella: Uncertainty-aware model-based offline reinforcement learning leveraging planning
C. Diehl, T. Sievernich, M. Krüger, F. Hoffmann, and T. Bertran · 2021
Cited alongside, same era.
Decision transformer: Reinforcement learning via sequence modeling
L. Chen, K. Lu, A. Rajeswaran, K. Lee, A. Grover, M. Laskin, P. Abbeel, A. Srinivas, and I. Mordatch · 2021
Cited alongside, same era.
Offline reinforcement learning as one big sequence modeling problem
M. Janner, Q. Li, and S. Levine · 2021
Cited alongside, same era.
Uncertainty weighted actor-critic for offline reinforcement learning
Y. Wu, S. Zhai, N. Srivastava, J. M. Susskind, J. Zhang, R. Salakhutdinov, and H. Goh · 2021
Cited alongside, same era.
Causal influence detection for improving efficiency in reinforcement learning
M. Seitzer, B. Schölkopf, and G. Martius · 2021
Cited alongside, same era.
H. Furuta, Y. Matsuo, and S. S. Gu · 2022
Later among the works it cites.
Imitating past successes can be very suboptimal
B. Eysenbach, S. Udatha, R. R. Salakhutdinov, and S. Levine · 2022
Later among the works it cites.
Upside-down reinforcement learning can diverge in stochastic environments with episodic resets
M. Štrupl, F. Faccio, D. R. Ashley, J. Schmidhuber, and R. K. Srivastava · 2022
Later among the works it cites.
Latent discriminant deterministic uncertainty
G. Franchi, X. Yu, A. Bursuc, E. Aldea, S. Dubuisson, and D. Filliat · 2022
Later among the works it cites.
Sample efficient deep reinforcement learning via uncertainty estimation
V. Mai, K. Mani, and L. Paull · 2022
Later among the works it cites.
Constraints penalized q-learning for safe offline reinforcement learning
H. Xu, X. Zhan, and X. Zhu · 2022
Later among the works it cites.
Efficient learning of safe driving policy via human-ai copilot optimization
Q. Li, Z. Peng, and B. Zhou · 2022
Later among the works it cites.
Y.-H. Wu, X. Wang, and M. Hamaya · 2023
Closest in time.
Decision transformer under random frame dropping
K. Hu, R. C. Zheng, Y. Gao, and H. Xu · 2023
Closest in time.
Act: Empowering decision transformer with dynamic programming via advantage conditioning
C.-X. Gao, C. Wu, M. Cao, R. Kong, Z. Zhang, and Y. Yu · 2024
Closest in time.
Critic-guided decision transformer for offline reinforcement learning
Y. Wang, C. Yang, Y. Wen, Y. Liu, and Y. Qiao · 2024
Closest in time.