Fetching the paper…
Reading the bibliography…
Reinforcement Learning (RL) can mitigate the causal confusion and distribution shift inherent to imitation learning (IL).
On information and sufficiency
Solomon Kullback and Richard A. Leibler · 1951
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih · 2013
Earlier work this paper cites.
Confusion can be beneficial for learning
Sidney D’Mello, Blair Lehman, Reinhard Pekrun, and Art Graesser · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, et al · 2017
Earlier work this paper cites.
Carla: An open urban driving simulator
Alexey Dosovitskiy, German Ros, Felipe Codevilla, Antonio Lopez, and Vladlen Koltun · 2017
Earlier work this paper cites.
Agile autonomous driving using end-to-end deep imitation learning
Yunpeng Pan, Ching-An Cheng, Kamil Saigol, Keuntaek Lee, Xinyan Yan, Evangelos Theodorou, and Byron Boots · 2017
Earlier work this paper cites.
Probabilistic recurrent state-space models
Andreas Doerr, Christian Daniel, Martin Schiegg, Nguyen-Tuong Duy, Stefan Schaal, Marc Toussaint, and Trimpe Sebastian · 2018
Earlier work this paper cites.
Model-free deep reinforcement learning for urban autonomous driving
Jianyu Chen, Bodi Yuan, and Masayoshi Tomizuka · 2019
Earlier work this paper cites.
Fighting copycat agents in behavioral cloning from observation histories
Chuan Wen, Jierui Lin, Trevor Darrell, Dinesh Jayaraman, and Yang Gao · 2020
Earlier work this paper cites.
End-to-end model-free reinforcement learning for urban driving using implicit affordances
Marin Toromanoff, Emilie Wirbel, and Fabien Moutarde · 2020
Earlier work this paper cites.
Learning by cheating
Dian Chen, Brady Zhou, Vladlen Koltun, and Philipp Krähenbühl · 2020
Earlier work this paper cites.
Dreamer: Efficient reinforcement learning through latent imagination
Danijar Hafner, Timothy Lillicrap, Jimmy Ba, and Mohammad Norouzi · 2020
Earlier work this paper cites.
Ide-net: Interactive driving event and pattern extraction from human data
Xiaosong Jia, Liting Sun, Masayoshi Tomizuka, and Wei Zhan · 2021
Earlier work this paper cites.
Multi-modal fusion transformer for end-to-end autonomous driving
Aditya Prakash, Kashyap Chitta, and Andreas Geiger · 2021
Earlier work this paper cites.
End-to-end urban driving by imitating a reinforcement learning coach
Zhejun Zhang, Alexander Liniger, Dengxin Dai, Fisher Yu, and Luc Van Gool · 2021
Earlier work this paper cites.
Multi-agent trajectory prediction by combining egocentric and allocentric views
Xiaosong Jia, Liting Sun, Hang Zhao, Masayoshi Tomizuka, and Wei Zhan · 2022
Earlier work this paper cites.
Trajectory-guided control prediction for end-to-end autonomous driving: A simple yet strong baseline
Penghao Wu, Xiaosong Jia, Li Chen, Junchi Yan, Hongyang Li, and Yu Qiao · 2022
Earlier work this paper cites.
Mobileye under the hood
Mobileye · 2022
Earlier work this paper cites.
NVIDIA DRIVE End-to-End Solutions for Autonomous Vehicles
NVIDIA · 2022
Earlier work this paper cites.
https://leaderboard.carla.org/get_started_v1/ , 2022
Get started with Leaderboard 1.0 — leaderboard.carla.org · 2022
Earlier work this paper cites.
Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers
Zhiqi Li, Wenhai Wang, Hongyang Li, Enze Xie, Chonghao Sima, Tong Lu, Yu Qiao, and Jifeng Dai · 2022
Earlier work this paper cites.
Meta-reward-net: Implicitly differentiable reward learning for preference-based reinforcement learning
Runze Liu, Fengshuo Bai, Yali Du, and Yaodong Yang · 2022
Earlier work this paper cites.
Transfuser: Imitation with transformer-based sensor fusion for autonomous driving
Kashyap Chitta, Aditya Prakash, Bernhard Jaeger, Zehao Yu, Katrin Renz, and Andreas Geiger · 2023
Earlier work this paper cites.
Policy pre-training for autonomous driving via self-supervised geometric modeling, 2023
Penghao Wu, Li Chen, Hongyang Li, Xiaosong Jia, Junchi Yan, and Yu Qiao · 2023
Cited alongside, same era.
Delving into the devils of bird’s-eye-view perception: A review, evaluation and recipe
Hongyang Li, Chonghao Sima, Jifeng Dai, Wenhai Wang, Lewei Lu, Huijie Wang, Jia Zeng, Zhiqi Li, Jiazhi Yang, Hanming Deng, et al · 2023
Cited alongside, same era.
Planning-oriented autonomous driving
Yihan Hu, Jiazhi Yang, Li Chen, Keyu Li, Chonghao Sima, Xizhou Zhu, Siqi Chai, Senyao Du, Tianwei Lin, Wenhai Wang, et al · 2023
Cited alongside, same era.
Vad: Vectorized scene representation for efficient autonomous driving
Bo Jiang, Shaoyu Chen, Qing Xu, Bencheng Liao, Jiajie Chen, Helong Zhou, Qian Zhang, Wenyu Liu, Chang Huang, and Xinggang Wang · 2023
Cited alongside, same era.
Towards capturing the temporal dynamics for trajectory prediction: a coarse-to-fine approach
Xiaosong Jia, Li Chen, Penghao Wu, Jia Zeng, Junchi Yan, Hongyang Li, and Yu Qiao · 2023
Cited alongside, same era.
Video generation models as world simulators
Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, et al · 2024
Later among the works it cites.
Rgbd objects in the wild: Scaling real-world 3d object learning from rgb-d videos
Hongchi Xia, Yang Fu, Sifei Liu, and Xiaolong Wang · 2024
Later among the works it cites.
Hidden biases of end-to-end driving datasets
Julian Zimmerlin, Jens Beißwenger, Bernhard Jaeger, Andreas Geiger, and Kashyap Chitta · 2024
Later among the works it cites.
Genad: Generative end-to-end autonomous driving
Wenzhao Zheng, Ruiqi Song, Xianda Guo, Chenming Zhang, and Long Chen · 2024
Later among the works it cites.
Amp: Autoregressive motion prediction revisited with next token prediction for autonomous driving
Xiaosong Jia, Shaoshuai Shi, Zijun Chen, Li Jiang, Wenlong Liao, Tao He, and Junchi Yan · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jiang-Tian Zhai, Ze Feng, Jihao Du, Yongqiang Mao, Jiang-Jiang Liu, Zichang Tan, Yifu Zhang, Xiaoqing Ye, and Jingdong Wang · 2023
Cited alongside, same era.
Think twice before driving: Towards scalable decoders for end-to-end autonomous driving
Xiaosong Jia, Penghao Wu, Li Chen, Jiangwei Xie, Conghui He, Junchi Yan, and Hongyang Li · 2023
Cited alongside, same era.
Driveadapter: Breaking the coupling barrier of perception and planning in end-to-end autonomous driving
Xiaosong Jia, Yulu Gao, Li Chen, Junchi Yan, Patrick Langechuan Liu, and Hongyang Li · 2023
Cited alongside, same era.
Lightzero: A unified benchmark for monte carlo tree search in general sequential decision scenarios
Yazhe Niu, Yuan Pu, Zhenjie Yang, Xueyan Li, Tong Zhou, Jiyuan Ren, Shuai Hu, Hongsheng Li, and Yu Liu · 2023
Cited alongside, same era.
https://leaderboard.carla.org/get_started/ , 2023
Get started with Leaderboard 2.0 — leaderboard.carla.org · 2023
Cited alongside, same era.
Mastering diverse domains through world models
Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap · 2023
Cited alongside, same era.
Rethinking the open-loop evaluation of end-to-end autonomous driving in nuscenes
Jiang-Tian Zhai, Ze Feng, Jinhao Du, Yongqiang Mao, Jiang-Jiang Liu, Zichang Tan, Yifu Zhang, Xiaoqing Ye, and Jingdong Wang · 2023
Cited alongside, same era.
Yutao Zhu, Xiaosong Jia, Xinyu Yang, and Junchi Yan · 2024
Later among the works it cites.
Activead: Planning-oriented active learning for end-to-end autonomous driving
Han Lu, Xiaosong Jia, Yichen Xie, Wenlong Liao, Xiaokang Yang, and Junchi Yan · 2024
Later among the works it cites.
Junqi You, Xiaosong Jia, Zhiyuan Zhang, Yutao Zhu, and Junchi Yan · 2024
Later among the works it cites.
Efficient preference-based reinforcement learning via aligned experience estimation
Fengshuo Bai, Rui Zhao, Hongming Zhang, Sijia Cui, Ying Wen, Yaodong Yang, Bo Xu, and Lei Han · 2024
Later among the works it cites.
Rat: Adversarial attacks on deep reinforcement agents for targeted behaviors
Fengshuo Bai, Runze Liu, Yali Du, Ying Wen, and Yaodong Yang · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al · 2025
Closest in time.
Enhancing end-to-end autonomous driving with latent world model
Yingyan Li, Lue Fan, Jiawei He, Yuqi Wang, Yuntao Chen, Zhaoxiang Zhang, and Tieniu Tan · 2025
Closest in time.
Adawm: Adaptive world model based planning for autonomous driving, 2025
Hang Wang, Xin Ye, Feng Tao, Chenbin Pan, Abhirup Mallik, Burhaneddin Yaman, Liu Ren, and Junshan Zhang · 2025
Closest in time.
Drivetransformer: Unified transformer for scalable end-to-end autonomous driving
Xiaosong Jia, Junqi You, Zhiyuan Zhang, and Junchi Yan · 2025
Closest in time.
Carl: Learning scalable planning policies with simple rewards
Bernhard Jaeger, Daniel Dauner, Jens Beißwenger, Simon Gerstenecker, Kashyap Chitta, and Andreas Geiger · 2025
Closest in time.
Don’t shake the wheel: Momentum-aware planning in end-to-end autonomous driving
Ziying Song, Caiyan Jia, Lin Liu, Hongyu Pan, Yongchang Zhang, Junming Wang, Xingyu Zhang, Shaoqing Xu, Lei Yang, and Yadan Luo · 2025
Closest in time.
https://github.com/Thinklab-SJTU/Bench2Drive
GitHub - Thinklab-SJTU/Bench2Drive: [NeurIPS 2024 Datasets and Benchmarks Track] Closed-Loop E2E-AD Benchmark Enhanced by World Model RL Expert — github.com · 2025
Closest in time.
Béziergs: Dynamic urban scene reconstruction with bézier curve gaussian splatting
Zipei Ma, Junzhe Jiang, Yurui Chen, and Li Zhang · 2025
Closest in time.
A focused human body model for accurate anthropometric measurements extraction
Shuhang Chen, Xianliang Huang, Zhizhou Zhong, Juhong Guan, and Shuigeng Zhou · 2025
Closest in time.
Resim: Reliable world simulation for autonomous driving
Jiazhi Yang, Kashyap Chitta, Shenyuan Gao, Long Chen, Yuqian Shao, Xiaosong Jia, Hongyang Li, Andreas Geiger, Xiangyu Yue, and Li Chen · 2025
Closest in time.
Trajectory-llm: A language-based data generator for trajectory prediction in autonomous driving
Kairui Yang, Zihao Guo, Gengjie Lin, Haotian Dong, Zhao Huang, Yipeng Wu, Die Zuo, Jibin Peng, Ziyuan Zhong, Xin Wang, et al · 2025
Closest in time.
Drivemoe: Mixture-of-experts for vision-language-action model in end-to-end autonomous driving
Zhenjie Yang, Yilin Chai, Xiaosong Jia, Qifeng Li, Yuqian Shao, Xuekai Zhu, Haisheng Su, and Junchi Yan · 2025
Closest in time.
Retrieval dexterity: Efficient object retrieval in clutters with dexterous hand
Fengshuo Bai, Yu Li, Jie Chu, Tawei Chou, Runchuan Zhu, Ying Wen, Yaodong Yang, and Yuanpei Chen · 2025
Closest in time.
Amulet: Realignment during test time for personalized preference adaptation of LLMs
Zhaowei Zhang, Fengshuo Bai, Qizhi Chen, Chengdong Ma, Mingzhi Wang, Haoran Sun, Zilong Zheng, and Yaodong Yang · 2025
Closest in time.
Flowrl: Matching reward distributions for llm reasoning
Xuekai Zhu, Daixuan Cheng, Dinghuai Zhang, Hengli Li, Kaiyan Zhang, Che Jiang, Youbang Sun, Ermo Hua, Yuxin Zuo, Xingtai Lv, et al · 2025
Closest in time.