Fetching the paper…
Reading the bibliography…
Autonomous driving (AD) agents generate driving policies based on online perception results, which are obtained at multiple levels of abstraction, e.g., behavior planning, motion planning and control.
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu, “Asynchronous methods for deep reinforcement learning,” in
1937
Earlier work this paper cites.
R. Bellman and R. Kalaba, “On the role of dynamic programming in statistical communication theory,”
1957
Earlier work this paper cites.
R. Bellman, “Dynamic programming,”
1966
Earlier work this paper cites.
C. A. Holloway,
1979
Earlier work this paper cites.
E. D. Dickmanns and A. Zapp, “Autonomous high speed road vehicle guidance by computer vision,”
1987
Earlier work this paper cites.
C. Thorpe, M. H. Hebert, T. Kanade, and S. A. Shafer, “Vision and navigation for the carnegie-mellon navlab,”
1988
Earlier work this paper cites.
D. A. Pomerleau, “Alvinn: An autonomous land vehicle in a neural network,” in
1989
Earlier work this paper cites.
G. P. Papavassilopoulos and M. G. Safonov, “Robust control design via game theoretic methods,” in
1989
Earlier work this paper cites.
W. S. Lovejoy, “A survey of algorithmic methods for partially observed markov decision processes,”
1991
Earlier work this paper cites.
D. A. Pomerleau, “Efficient training of artificial neural networks for autonomous navigation,”
1991
Earlier work this paper cites.
S. B. Thrun, “Efficient exploration in reinforcement learning,” 1992
1992
Earlier work this paper cites.
C. J. C. H. Watkins and P. Dayan, “Technical note q-learning,”
1992
Earlier work this paper cites.
L.-J. Lin, “Self-improving reactive agents based on reinforcement learning, planning and teaching,”
1992
Earlier work this paper cites.
R. J. Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,”
1992
Earlier work this paper cites.
B. Efron and R. Tibshirani,
1993
Earlier work this paper cites.
G. A. Rummery and M. Niranjan,
1994
Earlier work this paper cites.
L. C. Baird, “Reinforcement learning in continuous time: Advantage updating,” in
1994
Earlier work this paper cites.
G. Yu and I. K. Sethi, “Road-following with continuous learning,” in
1995
Earlier work this paper cites.
P. Wilson and D. Stahl, “On players’ models of other players: Theory and experimental evidence,”
1995
Earlier work this paper cites.
R. Camacho and D. Michie, “Behavioral cloning A correction,”
1995
Earlier work this paper cites.
Hong Zhang, V. Kumar, and J. Ostrowski, “Motion planning with uncertainty,” in
1998
Earlier work this paper cites.
A. Y. Ng and S. J. Russell, “Algorithms for inverse reinforcement learning,” in
2000
Earlier work this paper cites.
T. G. Dietterich, “Ensemble methods in machine learning,” in
2000
Earlier work this paper cites.
C. Guestrin, M. G. Lagoudakis, and R. Parr, “Coordinated reinforcement learning,” in
2002
Earlier work this paper cites.
Z. Wang, T. Schaul, M. Hessel, H. Hasselt, M. Lanctot, and N. Freitas, “Dueling network architectures for deep reinforcement learning,” in
2003
Earlier work this paper cites.
A. Y. Ng, A. Coates, M. Diel, V. Ganapathi, J. Schulte, B. Tse, E. Berger, and E. Liang, “Autonomous inverted helicopter flight via reinforcement learning,” in
2004
Earlier work this paper cites.
M. Coggan, “Exploration and exploitation in reinforcement learning,”
2004
Earlier work this paper cites.
P. Abbeel and A. Y. Ng, “Apprenticeship learning via inverse reinforcement learning,” in
2004
Earlier work this paper cites.
S. Thrun, M. Montemerlo, H. Dahlkamp, D. Stavens, A. Aron, J. Diebel, P. Fong, J. Gale, M. Halpenny, G. Hoffmann
2006
Earlier work this paper cites.
H. John and C. James, “Ngsim interstate 80 freeway dataset,” US Fedeal Highway Administration, FHWA-HRT-06-137, Washington, DC, USA, Tech. Rep., 2006
2006
Earlier work this paper cites.
N. D. Ratliff, J. A. Bagnell, and M. Zinkevich, “Maximum margin planning,” in
2006
Earlier work this paper cites.
N. D. Ratliff, D. M. Bradley, J. A. Bagnell, and J. E. Chestnutt, “Boosting structured prediction for imitation learning,” in
2006
Earlier work this paper cites.
U. Muller, J. Ben, E. Cosatto, B. Flepp, and Y. L. Cun, “Off-road obstacle avoidance through end-to-end learning,” in
2006
Earlier work this paper cites.
U. Branch, S. Ganebnyi, S. Kumkov, V. Patsko, and S. Pyatko, “Robust control in game problems with linear dynamics,”
2007
Earlier work this paper cites.
C. Urmson and W. Whittaker, “Self-driving cars and the urban challenge,”
2008
Earlier work this paper cites.
B. D. Ziebart, A. L. Maas, J. A. Bagnell, and A. K. Dey, “Maximum entropy inverse reinforcement learning,” in
2008
Earlier work this paper cites.
K. Leyton-Brown and Y. Shoham,
2008
Earlier work this paper cites.
M. Buehler, K. Iagnemma, and S. Singh,
2009
Earlier work this paper cites.
B. D. Argall, S. Chernova, M. M. Veloso, and B. Browning, “A survey of robot learning from demonstration,”
2009
Earlier work this paper cites.
N. D. Ratliff, D. Silver, and J. A. Bagnell, “Learning to search: Functional gradient techniques for imitation learning,”
2009
Earlier work this paper cites.
S. Thrun, “Toward robotic cars,”
2010
Earlier work this paper cites.
S. Ross and D. Bagnell, “Efficient reductions for imitation learning,” in
2010
Earlier work this paper cites.
S. Ross, G. J. Gordon, and D. Bagnell, “A reduction of imitation learning and structured prediction to no-regret online learning,” in
2011
Earlier work this paper cites.
S. Levine, Z. Popovic, and V. Koltun, “Nonlinear inverse reinforcement learning with gaussian processes,” in
2011
Earlier work this paper cites.
C. Desjardins and B. Chaib-Draa, “Cooperative adaptive cruise control: A reinforcement learning approach,”
2011
Earlier work this paper cites.
A. Eskandarian,
2012
Earlier work this paper cites.
G. Shani, J. Pineau, and R. Kaplow, “A survey of point-based pomdp solvers,”
2013
Earlier work this paper cites.
D. Zhao, B. Wang, and D. Liu, “A supervised actor–critic approach for adaptive cruise control,”
2013
Earlier work this paper cites.
W. Payre, J. Cestac, and P. Delhomme, “Intention to use a fully automated car: Attitudes and a priori acceptability,”
2014
Earlier work this paper cites.
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. A. Riedmiller, “Deterministic policy gradient algorithms,” in
2014
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in
2014
Earlier work this paper cites.
M. Zhu, M. Otte, P. Chaudhari, and E. Frazzoli, “Game theoretic controller synthesis for multi-robot motion planning part i: Trajectory based algorithms,” in
2014
Earlier work this paper cites.
Y. LeCun, Y. Bengio, and G. E. Hinton, “Deep learning,”
2015
Earlier work this paper cites.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. A. Riedmiller, A. Fidjeland, G. Ostrovski, and et al., “Human-level control through deep reinforcement learning,”
2015
Earlier work this paper cites.
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz, “Trust region policy optimization,” in
2015
Earlier work this paper cites.
M. Wulfmeier, P. Ondruska, and I. Posner, “Maximum entropy deep inverse reinforcement learning,”
2015
Earlier work this paper cites.
M. Kuderer, S. Gulati, and W. Burgard, “Learning driving styles for autonomous vehicles from demonstration,” in
2015
Earlier work this paper cites.
J. García, Fern, and o Fernández, “A comprehensive survey on safe reinforcement learning,”
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
H. Bai, S. Cai, N. Ye, D. Hsu, and W. S. Lee, “Intention-aware online pomdp planning for autonomous driving in a crowd,” in
2015
Earlier work this paper cites.
M. J. Kochenderfer,
2015
Earlier work this paper cites.
A. Talebpour and H. S. Mahmassani, “Influence of connected and autonomous vehicles on traffic flow stability and throughput,”
2016
Earlier work this paper cites.
D. González, J. Pérez, V. Milanés, and F. Nashashibi, “A review of motion planning techniques for automated vehicles,”
2016
Earlier work this paper cites.
X. Li, Z. Sun, D. Cao, Z. He, and Q. Zhu, “Real-time trajectory planning for autonomous urban driving: Framework, algorithms, and verifications,”
2016
Earlier work this paper cites.
B. Paden, M. Čáp, S. Z. Yong, D. Yershov, and E. Frazzoli, “A survey of motion planning and control techniques for self-driving urban vehicles,”
2016
Earlier work this paper cites.
I. J. Goodfellow, Y. Bengio, and A. C. Courville,
2016
Earlier work this paper cites.
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra, “Continuous control with deep reinforcement learning,” in
2016
Earlier work this paper cites.
T. Schaul, J. Quan, I. Antonoglou, and D. Silver, “Prioritized experience replay,” in
2016
Earlier work this paper cites.
S. Gu, T. P. Lillicrap, I. Sutskever, and S. Levine, “Continuous deep q-learning with model-based acceleration,” in
2016
Earlier work this paper cites.
J. Schulman, P. Moritz, S. Levine, M. I. Jordan, and P. Abbeel, “High-dimensional continuous control using generalized advantage estimation,” in
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
J. Zhang and K. Cho, “Query-efficient imitation learning for end-to-end autonomous driving,”
2016
Earlier work this paper cites.
C. Finn, S. Levine, and P. Abbeel, “Guided cost learning: Deep inverse optimal control via policy optimization,” in
2016
Earlier work this paper cites.
J. Ho and S. Ermon, “Generative adversarial imitation learning,” in
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2016
Cited alongside, same era.
M. J. Hausknecht and P. Stone, “Deep reinforcement learning in parameterized action space,” in
2016
Cited alongside, same era.
K. Min, H. Kim, and K. Huh, “Deep q learning based high level driving policy determination,” in
2018
Later among the works it cites.
R. P. Bhattacharyya, D. J. Phillips, B. Wulfe, J. Morton, A. Kuefler, and M. J. Kochenderfer, “Multi-agent imitation learning for driving simulation,” in
2018
Later among the works it cites.
M. Everett, Y. F. Chen, and J. P. How, “Motion planning among dynamic, decision-making agents with deep reinforcement learning,” in
2018
Later among the works it cites.
S. Qi and S.-C. Zhu, “Intent-aware multi-agent reinforcement learning,” in
2018
Later among the works it cites.
N. Jansen, B. Könighofer, S. Junges, and R. Bloem, “Shielded decision-making in mdps,”
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
2016
Cited alongside, same era.
A. Alahi, K. Goel, V. Ramanathan, A. Robicquet, F. Li, and S. Savarese, “Social LSTM: human trajectory prediction in crowded spaces,” in
2016
Cited alongside, same era.
Y. Gal, “Uncertainty in deep learning,”
2016
Cited alongside, same era.
A. Turnwald, D. Althoff, D. Wollherr, and M. Buss, “Understanding human avoidance behavior: interaction-aware decision making based on game theory,”
2016
Cited alongside, same era.
2017
Cited alongside, same era.
Y. Zhu, R. Mottaghi, E. Kolve, J. J. Lim, A. Gupta, L. Fei-Fei, and A. Farhadi, “Target-driven visual navigation in indoor scenes using deep reinforcement learning,” in
2017
Cited alongside, same era.
N. Fulton and A. Platzer, “Safe reinforcement learning via formal methods: Toward safe control through proof and learning,” in
2018
Later among the works it cites.
D. Isele, A. Nakhaei, and K. Fujimura, “Safe reinforcement learning on autonomous vehicles,” in
2018
Later among the works it cites.
G. Ding, S. Aghli, C. Heckman, and L. Chen, “Game-theoretic cooperative lane changing using data-driven models,” in
2018
Later among the works it cites.
A. Gupta, J. Johnson, L. Fei-Fei, S. Savarese, and A. Alahi, “Social GAN: socially acceptable trajectories with generative adversarial networks,” in
2018
Later among the works it cites.
A. Vemula, K. Muelling, and J. Oh, “Social attention: Modeling attention in human crowds,” in
2018
Later among the works it cites.
L. Sun, W. Zhan, and M. Tomizuka, “Probabilistic prediction of interactive driving behavior via hierarchical inverse reinforcement learning,” in
2018
Later among the works it cites.
S. Choi, K. Lee, S. Lim, and S. Oh, “Uncertainty-aware learning from demonstration using mixture density networks with sampling-free variance modeling,” in
2018
Later among the works it cites.
2018
Later among the works it cites.
W. Dabney, G. Ostrovski, D. Silver, and R. Munos, “Implicit quantile networks for distributional reinforcement learning,” in
2018
Later among the works it cites.
J. Lin and Z. Zhang, “Acgail: Imitation learning about multiple intentions with auxiliary classifier gans,” in
2018
Later among the works it cites.
E. Schmerling, K. Leung, W. Vollprecht, and M. Pavone, “Multimodal probabilistic model-based planning for human-robot interaction,” in
2018
Later among the works it cites.
R. Tian, S. Li, N. Li, I. Kolmanovsky, A. Girard, and Y. Yildiz, “Adaptive game-theoretic decision making for autonomous vehicle control at roundabouts,” in
2018
Later among the works it cites.
C. Hubmann, J. Schulz, G. Xu, D. Althoff, and C. Stiller, “A belief state planner for interactive merge maneuvers in congested traffic,” in
2018
Later among the works it cites.
Y. Chen, C. Dong, P. Palanisamy, P. Mudalige, K. Muelling, and J. M. Dolan, “Attention-based hierarchical deep reinforcement learning for lane change behaviors in autonomous driving,” in
2019
Later among the works it cites.
A. Kendall, J. Hawke, D. Janz, P. Mazur, D. Reda, J.-M. Allen, V.-D. Lam, A. Bewley, and A. Shah, “Learning to drive in a day,” in
2019
Later among the works it cites.
2019
Later among the works it cites.
J. Chen, B. Yuan, and M. Tomizuka, “Deep imitation learning for autonomous driving in generic urban scenarios with enhanced safety,” in
2019
Later among the works it cites.
D. Wang, C. Devin, Q.-Z. Cai, F. Yu, and T. Darrell, “Deep object-centric policies for autonomous driving,” in
2019
Later among the works it cites.
M. Bouton, A. Nakhaei, K. Fujimura, and M. J. Kochenderfer, “Safe reinforcement learning with scene decomposition for navigating complex urban environments,” in
2019
Later among the works it cites.
M. Bouton, A. Nakhaei, K. Fujimura, and M. J. Kochenderfer, “Cooperation-aware reinforcement learning for merging in dense traffic,” in
2019
Later among the works it cites.
Y. Tang, “Towards learning multi-agent negotiations via self-play,” in
2019
Later among the works it cites.
P. Wang, H. Li, and C.-Y. Chan, “Continuous control for automated lane change behavior based on deep deterministic policy gradient algorithm,” in
2019
Later among the works it cites.
R. Vasquez and B. Farooq, “Multi-objective autonomous braking system using naturalistic dataset,” in
2019
Later among the works it cites.
M. Abdou, H. Kamal, S. El-Tantawy, A. Abdelkhalek, O. Adel, K. Hamdy, and M. Abaas, “End-to-end deep conditional imitation learning for autonomous driving,” in
2019
Later among the works it cites.
Y. Cui, D. Isele, S. Niekum, and K. Fujimura, “Uncertainty-aware data aggregation for deep imitation learning,” in
2019
Later among the works it cites.
2019
Later among the works it cites.
J. Chen, B. Yuan, and M. Tomizuka, “Model-free deep reinforcement learning for urban autonomous driving,” in
2019
Later among the works it cites.
P. Hart, L. Rychly, and A. Knoll, “Lane-merging using policy-based reinforcement learning and post-optimization,” in
2019
Later among the works it cites.
P. Wang, D. Liu, J. Chen, H. Li, and C.-Y. Chan, “Human-like decision making for autonomous driving via adversarial inverse reinforcement learning,”
2019
Later among the works it cites.
A. Alizadeh, M. Moghadam, Y. Bicer, N. K. Ure, U. Yavas, and C. Kurtulus, “Automated lane change decision making using deep reinforcement learning in dynamic and uncertain highway environment,” in
2019
Later among the works it cites.
T. Tram, I. Batkovic, M. Ali, and J. Sjöberg, “Learning when to drive in intersections by combining reinforcement learning and model predictive control,” in
2019
Later among the works it cites.
C. Li and K. Czarnecki, “Urban driving with multi-objective deep reinforcement learning,” in
2019
Later among the works it cites.
M. P. Ronecker and Y. Zhu, “Deep q-network based decision making for autonomous driving,” in
2019
Later among the works it cites.
W. Yuan, M. Yang, Y. He, C. Wang, and B. Wang, “Multi-reward architecture based reinforcement learning for highway driving policies,” in
2019
Later among the works it cites.
J. Lee and J. W. Choi, “May i cut into your lane?: A policy network to learn interactive lane change behavior for autonomous driving,” in
2019
Later among the works it cites.
L. Wang, F. Ye, Y. Wang, J. Guo, I. Papamichail, M. Papageorgiou, S. Hu, and L. Zhang, “A q-learning foresighted approach to ego-efficient lane changes of connected and automated vehicles on freeways,” in
2019
Later among the works it cites.
D. Liu, M. Brännstrom, A. Backhouse, and L. Svensson, “Learning faster to perform autonomous lane changes by constructing maneuvers from shielded semantic actions,” in
2019
Later among the works it cites.
K. Min, H. Kim, and K. Huh, “Deep distributional reinforcement learning based high-level driving policy determination,”
2019
Later among the works it cites.
K. Rezaee, P. Yadmellat, M. S. Nosrati, E. A. Abolfathi, M. Elmahgiubi, and J. Luo, “Multi-lane cruising using hierarchical planning and reinforcement learning,” in
2019
Later among the works it cites.
T. Shi, P. Wang, X. Cheng, C.-Y. Chan, and D. Huang, “Driving decision and control for automated lane change behavior based on deep reinforcement learning,” in
2019
Later among the works it cites.
F. Behbahani, K. Shiarlis, X. Chen, V. Kurin, S. Kasewa, C. Stirbu, J. Gomes, S. Paul, F. A. Oliehoek, J. Messias
2019
Later among the works it cites.
L. Chen, Y. Chen, X. Yao, Y. Shan, and L. Chen, “An adaptive path tracking controller based on reinforcement learning with urban driving application,” in
2019
Later among the works it cites.
M. Huegle, G. Kalweit, B. Mirchevska, M. Werling, and J. Boedecker, “Dynamic input for deep reinforcement learning in autonomous driving,” in
2019
Later among the works it cites.
R. P. Bhattacharyya, D. J. Phillips, C. Liu, J. K. Gupta, K. Driggs-Campbell, and M. J. Kochenderfer, “Simulating emergent properties of human driving behavior using multi-agent reward augmented imitation learning,” in
2019
Later among the works it cites.
D. Hayashi, Y. Xu, T. Bando, and K. Takeda, “A predictive reward function for human-like driving based on a transition model of surrounding environment,” in
2019
Later among the works it cites.
C. Chen, Y. Liu, S. Kreiss, and A. Alahi, “Crowd-robot interaction: Crowd-aware robot navigation with attention-based deep reinforcement learning,” in
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
F. Pusse and M. Klusch, “Hybrid online pomdp planning and deep reinforcement learning for safer self-driving cars,” in
2019
Later among the works it cites.
A. Mohseni-Kabir, D. Isele, and K. Fujimura, “Interaction-aware multi-agent reinforcement learning for mobile agents with individual goals,” in
2019
Later among the works it cites.
J. F. Fisac, E. Bronstein, E. Stefansson, D. Sadigh, S. S. Sastry, and A. D. Dragan, “Hierarchical game-theoretic planning for autonomous vehicles,” in
2019
Later among the works it cites.
2019
Later among the works it cites.
M. Henaff, A. Canziani, and Y. LeCun, “Model-predictive policy learning with uncertainty regularization for driving in dense traffic,” in
2019
Later among the works it cites.
C. Yu, X. Wang, X. Xu, M. Zhang, H. Ge, J. Ren, L. Sun, B. Chen, and G. Tan, “Distributed multiagent coordinated learning for autonomous driving in highways based on dynamic coordination graphs,”
2019
Later among the works it cites.
P. Wang, Y. Li, S. Shekhar, and W. F. Northrop, “Uncertainty estimation with distributional reinforcement learning for applications in intelligent transportation systems: A case study,” in
2019
Later among the works it cites.
J. Bernhard, S. Pollok, and A. Knoll, “Addressing inherent uncertainty: Risk-sensitive behavior generation for automated driving using distributional reinforcement learning,” in
2019
Later among the works it cites.
J. Li, H. Ma, and M. Tomizuka, “Interaction-aware multi-agent tracking and probabilistic behavior prediction via adversarial learning,” in
2019
Later among the works it cites.
H. Ma, J. Li, W. Zhan, and M. Tomizuka, “Wasserstein generative learning with kinematic constraints for probabilistic interactive driving behavior prediction,” in
2019
Later among the works it cites.
D. Isele, “Interactive decision making for autonomous vehicles in dense traffic,” in
2019
Later among the works it cites.
S. M. Grigorescu, B. Trasnea, T. T. Cocias, and G. Macesanu, “A survey of deep learning techniques for autonomous driving,”
2020
Later among the works it cites.
2020
Later among the works it cites.
J. Jiang, C. Dun, T. Huang, and Z. Lu, “Graph convolutional reinforcement learning,” in
2020
Later among the works it cites.
C. Fei, B. Wang, Y. Zhuang, Z. Zhang, J. Hao, H. Zhang, X. Ji, and W. Liu, “Triple-gail: A multi-modal imitation learning framework with generative adversarial nets,” in
2020
Later among the works it cites.
A. Folkers, M. Rick, and C. Büskens, “Controlling an autonomous vehicle with deep reinforcement learning,” in
2031
Closest in time.
D. Isele, R. Rahimi, A. Cosgun, K. Subramanian, and K. Fujimura, “Navigating occluded intersections with autonomous vehicles using deep reinforcement learning,” in
2039
Closest in time.
N. Deshpande and A. Spalanzani, “Deep reinforcement learning based vehicle navigation amongst pedestrians using a grid-based state representation,” in
2086
Closest in time.
H. v. Hasselt, A. Guez, and D. Silver, “Deep reinforcement learning with double q-learning,” in
2094
Closest in time.