Fetching the paper…
Reading the bibliography…
We investigate reinforcement learning (RL) for privileged planning in autonomous driving.
Data parallel algorithms
W. D. Hillis and G. L. S. Jr · 1986
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Q-learning
C. J. Watkins and P. Dayan · 1992
Earlier work this paper cites.
Congested traffic states in empirical observations and microscopic simulations
M. Treiber, A. Hennecke, and D. Helbing · 2000
Earlier work this paper cites.
Stanley: The robot that won the DARPA grand challenge
S. Thrun, M. Montemerlo, H. Dahlkamp, D. Stavens, A. Aron, J. Diebel, P. Fong, J. Gale, M. Halpenny, G. Hoffmann, K. Lau, C. M. Oakley, M. Palatucci, V. R. Pratt, P. Stang, S. Strohband, C. Dupont, L. Jendrossek, C. Koelen, C. Markey, C. Rummel, J. van Niekerk, E. Jensen, P. Alessandrini, G. R. Bradski, B. Davies, S. Ettinger, A. Kaehler, A. V. Nefian, and P. Mahoney · 2006
Earlier work this paper cites.
Vehicle dynamics and control
R. Rajamani · 2011
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
S. Ross, G. J. Gordon, and D. Bagnell · 2011
Earlier work this paper cites.
Off-policy actor-critic
T. Degris, M. White, and R. S. Sutton · 2012
Earlier work this paper cites.
ZeroMQ: Messaging for Many Applications
P. Hintjens · 2013
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. A. Riedmiller, A. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis · 2015
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. P. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Earlier work this paper cites.
Openai gym
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
CARLA: An open urban driving simulator
A. Dosovitskiy, G. Ros, F. Codevilla, A. Lopez, and V. Koltun · 2017
Earlier work this paper cites.
Reproducibility of benchmarked deep reinforcement learning tasks for continuous control
R. Islam, P. Henderson, M. Gomrokchi, and D. Precup · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
The kinematic bicycle model: A consistent model for planning feasible trajectories for autonomous vehicles?
P. Polack, F. Altché, B. d’Andréa Novel, and A. de La Fortelle · 2017
Earlier work this paper cites.
Improving stochastic policy gradients in continuous control with deep reinforcement learning using the beta distribution
P. Chou, D. Maturana, and S. A. Scherer · 2017
Earlier work this paper cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Earlier work this paper cites.
Deep reinforcement learning that matters
P. Henderson, R. Islam, P. Bachman, J. Pineau, D. Precup, and D. Meger · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Earlier work this paper cites.
CIRL: controllable imitative reinforcement learning for vision-based self-driving
X. Liang, T. Wang, L. Yang, and E. P. Xing · 2018
Earlier work this paper cites.
Recurrent world models facilitate policy evolution
D. Ha and J. Schmidhuber · 2018
Earlier work this paper cites.
Learning to drive in a day
A. Kendall, J. Hawke, D. Janz, P. Mazur, D. Reda, J. Allen, V. Lam, A. Bewley, and A. Shah · 2019
Earlier work this paper cites.
A survey on reproducibility by evaluating deep reinforcement learning algorithms on real-world robots
N. A. Lynnerup, L. Nolling, R. Hasle, and J. Hallam · 2019
Earlier work this paper cites.
Grandmaster level in starcraft II using multi-agent reinforcement learning
O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev, J. Oh, D. Horgan, M. Kroiss, I. Danihelka, A. Huang, L. Sifre, T. Cai, J. P. Agapiou, M. Jaderberg, A. S. Vezhnevets, R. Leblond, T. Pohlen, V. Dalibard, D. Budden, Y. Sulsky, J. Molloy, T. L. Paine, Ç. Gülçehre, Z. Wang, T. Pfaff, Y. Wu, R. Ring, D. Yogatama, D. Wünsch, K. McKinney, O. Smith, T. Schaul, T. P. Lillicrap, K. Kavukcuoglu, D. Hassabis, C. Apps, and D. Silver · 2019
Earlier work this paper cites.
Dota 2 with large scale deep reinforcement learning
C. Berner, G. Brockman, B. Chan, V. Cheung, P. Debiak, C. Dennison, D. Farhi, Q. Fischer, S. Hashme, C. Hesse, R. Józefowicz, S. Gray, C. Olsson, J. Pachocki, M. Petrov, H. P. de Oliveira Pinto, J. Raiman, T. Salimans, J. Schlatter, J. Schneider, S. Sidor, I. Sutskever, J. Tang, F. Wolski, and S. Zhang · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Köpf, E. Z. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala · 2019
Earlier work this paper cites.
DD-PPO: learning near-perfect pointgoal navigators from 2.5 billion frames
E. Wijmans, A. Kadian, A. Morcos, S. Lee, I. Essa, D. Parikh, M. Savva, and D. Batra · 2020
Earlier work this paper cites.
End-to-end model-free reinforcement learning for urban driving using implicit affordances
M. Toromanoff, E. Wirbel, and F. Moutarde · 2020
Earlier work this paper cites.
Exploring data aggregation in policy learning for vision-based urban autonomous driving
A. Prakash, A. Behl, E. Ohn-Bar, K. Chitta, and A. Geiger · 2020
Earlier work this paper cites.
Label efficient visual abstractions for autonomous driving
A. Behl, K. Chitta, A. Prakash, E. Ohn-Bar, and A. Geiger · 2020
Earlier work this paper cites.
Autonomous driving: The way forward
V. Koltun · 2020
Earlier work this paper cites.
Scalability in perception for autonomous driving: Waymo open dataset
P. Sun, H. Kretzschmar, X. Dotiwalla, A. Chouard, V. Patnaik, P. Tsui, J. Guo, Y. Zhou, Y. Chai, B. Caine, V. Vasudevan, W. Han, J. Ngiam, H. Zhao, A. Timofeev, S. Ettinger, M. Krivokon, A. Gao, A. Joshi, Y. Zhang, J. Shlens, Z. Chen, and D. Anguelov · 2020
Earlier work this paper cites.
Dream to control: Learning behaviors by latent imagination
D. Hafner, T. P. Lillicrap, J. Ba, and M. Norouzi · 2020
Earlier work this paper cites.
Sample factory: Egocentric 3d control from pixels at 100000 FPS with asynchronous reinforcement learning
A. Petrenko, Z. Huang, T. Kumar, G. S. Sukhatme, and V. Koltun · 2020
Earlier work this paper cites.
Expert drivers for autonomous driving
B. Jaeger · 2021
Cited alongside, same era.
End-to-end urban driving by imitating a reinforcement learning coach
Z. Zhang, A. Liniger, D. Dai, F. Yu, and L. Van Gool · 2021
Cited alongside, same era.
Super-human performance in gran turismo sport using deep reinforcement learning
F. Fuchs, Y. Song, E. Kaufmann, D. Scaramuzza, and P. Dürr · 2021
Cited alongside, same era.
Deep reinforcement learning at the edge of the statistical precipice
R. Agarwal, M. Schwarzer, P. S. Castro, A. C. Courville, and M. G. Bellemare · 2021
Cited alongside, same era.
Simulation performance evaluation of pure pursuit, stanley, lqr, mpc controller for autonomous vehicles
J. Liu, Z. Yang, Z. Huang, W. Li, S. Dang, and H. Li · 2021
Cited alongside, same era.
Autonomous overtaking in gran turismo sport using curriculum reinforcement learning
Mastering diverse domains through world models
D. Hafner, J. Pasukonis, J. Ba, and T. Lillicrap · 2023
Later among the works it cites.
Drivelm: Driving with graph visual question answering
C. Sima, K. Renz, K. Chitta, L. Chen, H. Zhang, C. Xie, J. Beißwenger, P. Luo, A. Geiger, and H. Li · 2024
Later among the works it cites.
Rethinking imitation-based planners for autonomous driving
J. Cheng, Y. Chen, X. Mei, B. Yang, B. Li, and M. Liu · 2024
Later among the works it cites.
Differentiable integrated motion prediction and planning with learnable cost function for autonomous driving
Z. Huang, H. Liu, J. Wu, and C. Lv · 2024
Later among the works it cites.
Generalizing motion planners with mixture of experts for autonomous driving
Q. Sun, H. Wang, J. Zhan, F. Nie, X. Wen, L. Xu, K. Zhan, P. Jia, X. Lang, and H. Zhao · 2024
Later among the works it cites.
Rethinking closed-loop planning framework for imitation-based model integrating prediction and planning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Song, H. Lin, E. Kaufmann, P. Dürr, and D. Scaramuzza · 2021
Cited alongside, same era.
End-to-end urban driving by imitating a reinforcement learning coach
Z. Zhang, A. Liniger, D. Dai, F. Yu, and L. V. Gool · 2021
Cited alongside, same era.
Mastering atari with discrete world models
D. Hafner, T. P. Lillicrap, M. Norouzi, and J. Ba · 2021
Cited alongside, same era.
Proximal policy optimization with continuous bounded action space via the beta distribution
I. G. B. Petrazzini and E. A. Antonelo · 2021
Cited alongside, same era.
Urban driver: Learning to drive from real-world demonstrations using policy gradients
O. Scheel, L. Bergamini, M. Wolczyk, B. Osiński, and P. Ondruska · 2021
Cited alongside, same era.
Plant: Explainable planning transformers via object-level representations
K. Renz, K. Chitta, O.-B. Mercea, A. S. Koepke, Z. Akata, and A. Geiger · 2022
Cited alongside, same era.
Rethinking closed-loop training for autonomous driving
C. Zhang, R. Guo, W. Zeng, Y. Xiong, B. Dai, R. Hu, M. Ren, and R. Urtasun · 2022
Cited alongside, same era.
J. Guo, M. Feng, P. Zhu, C. Li, and J. Pu · 2024
Later among the works it cites.
PLUTO: pushing the limit of imitation learning-based planning for autonomous driving
J. Cheng, Y. Chen, and Q. Chen · 2024
Later among the works it cites.
An invitation to deep reinforcement learning
B. Jaeger and A. Geiger · 2024
Later among the works it cites.
Poliformer: Scaling on-policy RL with transformers results in masterful navigators
K. Zeng, Z. Zhang, K. Ehsani, R. Hendrix, J. Salvador, A. Herrasti, R. B. Girshick, A. Kembhavi, and L. Weihs · 2024
Later among the works it cites.
Think2drive: Efficient reinforcement learning by thinking with latent world model for autonomous driving (in CARLA-V2)
Q. Li, X. Jia, S. Wang, and J. Yan · 2024
Later among the works it cites.
Towards learning-based planning: The nuplan benchmark for real-world autonomous driving
N. Karnchanachari, D. Geromichalos, K. S. Tan, N. Li, C. Eriksen, S. Yaghoubi, N. Mehdipour, G. Bernasconi, W. K. Fong, Y. Guo, and H. Caesar · 2024
Later among the works it cites.
Leaderboard 1.0 to 2.0 scenario converter
CARLA · 2024
Later among the works it cites.
Carla leaderboard 2.0 metrics
C. team · 2024
Later among the works it cites.
Common mistakes in benchmarking autonomous driving
B. Jaeger, K. Chitta, D. Dauner, K. Renz, and A. Geiger · 2024
Later among the works it cites.
MBAPPE: mcts-built-around prediction for planning explicitly
R. Chekroun, T. Gilles, M. Toromanoff, S. Hornauer, and F. Moutarde · 2024
Later among the works it cites.
Diffusion-es: Gradient-free planning with diffusion for autonomous and instruction-guided driving
B. Yang, H. Su, N. Gkanatsios, T. Ke, A. Jain, J. G. Schneider, and K. Fragkiadaki · 2024
Later among the works it cites.
Easychauffeur: A baseline advancing simplicity and efficiency on waymax
L. Xiao, J. Liu, X. Ye, W. Yang, and J. Wang · 2024
Later among the works it cites.
End-to-end autonomous driving: Challenges and frontiers
L. Chen, P. Wu, K. Chitta, B. Jaeger, A. Geiger, and H. Li · 2024
Later among the works it cites.
Langprop: A code optimization framework using large language models applied to driving
S. Ishida, G. Corrado, G. Fedoseev, H. Yeo, L. Russell, J. Shotton, J. F. Henriques, and A. Hu · 2024
Later among the works it cites.
Loss of plasticity in deep continual learning
S. Dohare, J. F. Hernandez-Garcia, Q. Lan, P. Rahman, A. R. Mahmood, and R. S. Sutton · 2024
Later among the works it cites.
A study of plasticity loss in on-policy deep reinforcement learning
A. Juliani and J. Ash · 2024
Later among the works it cites.
No representation, no trust: Connecting representation, collapse, and trust issues in PPO
S. Moalla, A. Miele, D. Pyatko, R. Pascanu, and C. Gulcehre · 2024
Later among the works it cites.
Gymnasium: A standard interface for reinforcement learning environments
M. Towers, A. Kwiatkowski, J. K. Terry, J. U. Balis, G. D. Cola, T. Deleu, M. Goulão, A. Kallinteris, M. Krimmel, A. KG, R. Perez-Vicente, A. Pierré, S. Schulhoff, J. J. Tai, H. Tan, and O. G. Younis · 2024
Later among the works it cites.
Regents: Real-world safety-critical driving scenario generation made stable
Y. Yin, P. Khayatan, Éloi Zablocki, A. Boulch, , and M. Cord · 2024
Later among the works it cites.
geopandas/geopandas: v1.0.1, 2024
J. V. den Bossche, K. Jordahl, M. Fleischmann, M. Richards, J. McBride, J. Wasserman, A. G. Badaracco, A. D. Snow, B. Ward, J. Tratner, J. Gerard, M. Perry, cjqf, G. A. Hjelle, M. Taves, E. ter Hoeven, M. Cochran, R. Bell, rraymondgh, M. Bartos, P. Roggemans, L. Culbertson, G. Caria, N. Y. Tan, N. Eubank, sangarshanan, J. Flavin, S. Rey, and J. Gardiner · 2024
Later among the works it cites.
Hidden biases of end-to-end driving datasets
J. Zimmerlin, J. Beißwenger, B. Jaeger, A. Geiger, and K. Chitta · 2024
Later among the works it cites.
Int2planner: An intention-based multi-modal motion planner for integrated prediction and planning
X. Chen, J. Yan, W. Liao, T. He, and P. Peng · 2025
Closest in time.
Diffusion-based planning for autonomous driving with flexible guidance
Y. Zheng, R. Liang, K. Zheng, J. Zheng, L. Mao, J. Li, W. Gu, R. Ai, S. E. Li, X. Zhan, et al · 2025
Closest in time.
Carplanner: Consistent auto-regressive trajectory planning for large-scale reinforcement learning in autonomous driving
D. Zhang, J. Liang, K. Guo, S. Lu, Q. Wang, R. Xiong, Z. Miao, and Y. Wang · 2025
Closest in time.
Mastering diverse control tasks through world models
D. Hafner, J. Pasukonis, J. Ba, and T. Lillicrap · 2025
Closest in time.
Sim-to-real reinforcement learning for vision-based dexterous manipulation on humanoids
T. Lin, K. Sachdev, L. Fan, J. Malik, and Y. Zhu · 2025
Closest in time.
V-max: Making rl practical for autonomous driving
V. Charraut, T. Tournaire, W. Doulazmi, and T. Buhet · 2025
Closest in time.
Robust autonomy emerges from self-play
M. Cusumano-Towner, D. Hafner, A. Hertzberg, B. Huval, A. Petrenko, E. Vinitsky, E. Wijmans, T. Killian, S. Bowers, O. Sener, P. Krähenbühl, and V. Koltun · 2025
Closest in time.
Gpudrive: Data-driven, multi-agent driving simulation at 1 million FPS
S. Kazemkhani, A. Pandya, D. Cornelisse, B. Shacklett, and E. Vinitsky · 2025
Closest in time.
Pep 703 – making the global interpreter lock optional in cpython
Python · 2025
Closest in time.