Fetching the paper…
Reading the bibliography…
Efficient and effective exploration in continuous space is a central problem in applying reinforcement learning (RL) to autonomous driving.
R. S. Sutton, D. Precup, and S. Singh, “Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning,” Artificial intelligence , vol. 112, no. 1-2, pp. 181–211, 1999
1999
Earlier work this paper cites.
J. Jing, E. C. Cropper, I. Hurwitz, and K. R. Weiss, “The construction of movement with behavior-specific and behavior-independent modules,” Journal of Neuroscience , vol. 24, no. 28, pp. 6315–6325, 2004
2004
Earlier work this paper cites.
T. M. Howard and A. Kelly, “Optimal rough terrain trajectory generation for wheeled mobile robots,” The International Journal of Robotics Research , vol. 26, no. 2, pp. 141–166, 2007
2007
Earlier work this paper cites.
J. Peters and S. Schaal, “Reinforcement learning of motor skills with policy gradients,” Neural networks , vol. 21, no. 4, pp. 682–697, 2008
2008
Earlier work this paper cites.
B. D. Ziebart, Modeling purposeful adaptive behavior with the principle of maximum causal entropy . Carnegie Mellon University, 2010
2010
Earlier work this paper cites.
M. W. Mueller, M. Hehn, and R. D’Andrea, “A computationally efficient motion primitive for quadrocopter trajectory generation,” IEEE transactions on robotics , vol. 31, no. 6, pp. 1294–1310, 2015
2015
Earlier work this paper cites.
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot et al. , “Mastering the game of go with deep neural networks and tree search,” nature , vol. 529, no. 7587, pp. 484–489, 2016
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
T. Gu, “Improved trajectory planning for on-road self-driving vehicles via combined graph search, optimization & topology analysis,” Ph.D. dissertation, Carnegie Mellon University, 2017
2017
Earlier work this paper cites.
P.-L. Bacon, J. Harb, and D. Precup, “The option-critic architecture,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 31, no. 1, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
A. Nair, B. McGrew, M. Andrychowicz, W. Zaremba, and P. Abbeel, “Overcoming exploration in reinforcement learning with demonstrations,” in 2018 IEEE international conference on robotics and automation (ICRA) . IEEE, 2018, pp. 6292–6299
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
T. Hester, M. Vecerik, O. Pietquin, M. Lanctot, T. Schaul, B. Piot, D. Horgan, J. Quan, A. Sendonaris, I. Osband et al. , “Deep q-learning from demonstrations,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 32, no. 1, 2018
2018
Cited alongside, same era.
N. Deo, A. Rangesh, and M. M. Trivedi, “How would surround vehicles move? a unified framework for maneuver classification and motion prediction,” IEEE Transactions on Intelligent Vehicles , vol. 3, no. 2, pp. 129–140, 2018
2018
Cited alongside, same era.
Y. Lee, S.-H. Sun, S. Somasundaram, E. S. Hu, and J. J. Lim, “Composing complex skills by learning transition policies,” in International Conference on Learning Representations , 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
Z. Li, W. Zhan, L. Sun, C.-Y. Chan, and M. Tomizuka, “Adaptive sampling-based motion planning with a non- conservatively defensive strategy for autonomous driving,” in The 21st IFAC World Congress , 2020
2020
Later among the works it cites.
D. M. Saxena, S. Bae, A. Nakhaei, K. Fujimura, and M. Likhachev, “Driving in dense traffic with model-free reinforcement learning,” in 2020 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2020, pp. 5385–5392
2020
Later among the works it cites.
2020
Later among the works it cites.
J. Merel, S. Tunyasuvunakool, A. Ahuja, Y. Tassa, L. Hasenclever, V. Pham, T. Erez, G. Wayne, and N. Heess, “Catch & carry: reusable neural controllers for vision-guided whole-body tasks,” ACM Transactions on Graphics (TOG) , vol. 39, no. 4, pp. 39–1, 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
2018
Cited alongside, same era.
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, “Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,” in International conference on machine learning . PMLR, 2018, pp. 1861–1870
2018
Cited alongside, same era.
2018
Cited alongside, same era.
O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev et al. , “Grandmaster level in starcraft ii using multi-agent reinforcement learning,” Nature , vol. 575, no. 7782, pp. 350–354, 2019
2019
Cited alongside, same era.
T. Kipf, Y. Li, H. Dai, V. Zambaldi, A. Sanchez-Gonzalez, E. Grefenstette, P. Kohli, and P. Battaglia, “Compile: Compositional imitation learning and execution,” in International Conference on Machine Learning . PMLR, 2019, pp. 3418–3428
2019
Cited alongside, same era.
T. Shankar, S. Tulsiani, L. Pinto, and A. Gupta, “Discovering motor programs by recomposing demonstrations,” in International Conference on Learning Representations , 2019
2019
Cited alongside, same era.
Y. Pan, J. Xue, P. Zhang, W. Ouyang, J. Fang, and X. Chen, “Navigation command matching for vision-based autonomous driving,” in 2020 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2020, pp. 4343–4349
2020
Cited alongside, same era.
2020
Later among the works it cites.
K. Pertsch, O. Rybkin, J. Yang, S. Zhou, K. Derpanis, K. Daniilidis, J. Lim, and A. Jaegle, “Keyframing the future: Keyframe discovery for visual prediction and planning,” in Learning for Dynamics and Control . PMLR, 2020, pp. 969–979
2020
Later among the works it cites.
2021
Later among the works it cites.
Y. Wu, M. Mozifian, and F. Shkurti, “Shaping rewards for reinforcement learning with imperfect demonstrations using generative models,” in 2021 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2021, pp. 6628–6634
2021
Later among the works it cites.
2021
Later among the works it cites.
M. Dalal, D. Pathak, and R. R. Salakhutdinov, “Accelerating robotic reinforcement learning via parameterized action primitives,” Advances in Neural Information Processing Systems , vol. 34, 2021
2021
Later among the works it cites.
Z. Wang, C. R. Garrett, L. P. Kaelbling, and T. Lozano-Pérez, “Learning compositional models of robot skills for task and motion planning,” The International Journal of Robotics Research , vol. 40, no. 6-7, pp. 866–894, 2021
2021
Later among the works it cites.
L. Wang, L. Sun, M. Tomizuka, and W. Zhan, “Socially-compatible behavior design of autonomous vehicles with verification on real human data,” IEEE Robotics and Automation Letters , vol. 6, no. 2, pp. 3421–3428, 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.