Fetching the paper…
Reading the bibliography…
With the development of deep representation learning, the domain of reinforcement learning (RL) has become a powerful learning framework now capable of learning complex policies in high dimensional environments.
V. Pareto, Manual of political economy . OUP Oxford, 1906
1906
Earlier work this paper cites.
B. F. Skinner, The behavior of organisms: An experimental analysis. Appleton-Century, 1938
1938
Earlier work this paper cites.
R. Bellman, Dynamic Programming . Princeton, NJ, USA: Princeton University Press, 1957
1957
Earlier work this paper cites.
C. J. C. H. Watkins, “Learning from delayed rewards,” Ph.D. dissertation, King’s College, Cambridge, 1989
1989
Earlier work this paper cites.
D. A. Pomerleau, “Alvinn: An autonomous land vehicle in a neural network,” in Advances in neural information processing systems , 1989
1989
Earlier work this paper cites.
R. S. Sutton, “Integrated architectures for learning, planning, and reacting based on approximating dynamic programming,” in Machine Learning Proceedings 1990 . Elsevier, 1990
1990
Earlier work this paper cites.
D. Pomerleau, “Efficient training of artificial neural networks for autonomous navigation,” Neural Computation , vol. 3, no. 1, 1991
1991
Earlier work this paper cites.
C. J. Watkins and P. Dayan, “Technical note: Q-learning,” Machine Learning , vol. 8, no. 3-4, 1992
1992
Earlier work this paper cites.
R. J. Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” Machine Learning , vol. 8, pp. 229–256, 1992
1992
Earlier work this paper cites.
M. L. Puterman, Markov Decision Processes: Discrete Stochastic Dynamic Programming , 1st ed. New York, NY, USA: John Wiley & Sons, Inc., 1994
1994
Earlier work this paper cites.
G. A. Rummery and M. Niranjan, “On-line Q-learning using connectionist systems,” Cambridge University Engineering Department, Cambridge, England, Tech. Rep. TR 166, 1994
1994
Earlier work this paper cites.
G. Tesauro, “Td-gammon, a self-teaching backgammon program, achieves master-level play,” Neural Computing , vol. 6, no. 2, Mar. 1994
1994
Earlier work this paper cites.
T. M. Mitchell, Machine learning , ser. McGraw-Hill series in computer science. Boston (Mass.), Burr Ridge (Ill.), Dubuque (Iowa): McGraw-Hill, 1997
1997
Earlier work this paper cites.
J. Randløv and P. Alstrøm, “Learning to drive a bicycle using reinforcement learning and shaping,” in Proceedings of the Fifteenth International Conference on Machine Learning , ser. ICML ’98. San Francisco, CA, USA: Morgan Kaufmann Publishers Inc., 1998, pp. 463–471
1998
Earlier work this paper cites.
A. Y. Ng, D. Harada, and S. J. Russell, “Policy invariance under reward transformations: Theory and application to reward shaping,” in Proceedings of the Sixteenth International Conference on Machine Learning , ser. ICML ’99, 1999, pp. 278–287
1999
Earlier work this paper cites.
R. S. Sutton, D. Precup, and S. Singh, “Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning,” Artificial intelligence , vol. 112, no. 1-2, pp. 181–211, 1999
1999
Earlier work this paper cites.
A. Y. Ng, D. Harada, and S. Russell, “Policy invariance under reward transformations: Theory and application to reward shaping,” in ICML , vol. 99, 1999, pp. 278–287
1999
Earlier work this paper cites.
D. H. Wolpert, K. R. Wheeler, and K. Tumer, “Collective intelligence for control of distributed dynamical systems,” EPL (Europhysics Letters) , vol. 49, no. 6, p. 708, 2000
2000
Earlier work this paper cites.
A. Y. Ng, S. J. Russell et al. , “Algorithms for inverse reinforcement learning.” in ICML , 2000
2000
Earlier work this paper cites.
B. Wymann, E. Espié, C. Guionneau, C. Dimitrakakis, R. Coulom, and A. Sumner, “Torcs, the open racing car simulator,” Software available at http://torcs. sourceforge. net , vol. 4, 2000
2000
Earlier work this paper cites.
S. M. LaValle and J. James J. Kuffner, “Randomized kinodynamic planning,” The International Journal of Robotics Research , vol. 20, no. 5, pp. 378–400, 2001
2001
Earlier work this paper cites.
R. I. Brafman and M. Tennenholtz, “R-max-a general polynomial time algorithm for near-optimal reinforcement learning,” Journal of Machine Learning Research , vol. 3, no. Oct, 2002
2002
Earlier work this paper cites.
P. Abbeel and A. Y. Ng, “Apprenticeship learning via inverse reinforcement learning,” in Proceedings of the twenty-first international conference on Machine learning . ACM, 2004, p. 1
2004
Earlier work this paper cites.
N. Koenig and A. Howard, “Design and use paradigms for gazebo, an open-source multi-robot simulator,” in 2004 International Conference on Intelligent Robots and Systems (IROS) , vol. 3. IEEE, 2004, pp. 2149–2154
2004
Earlier work this paper cites.
P. Abbeel and A. Y. Ng, “Apprenticeship learning via inverse reinforcement learning,” in Proceedings of the Twenty-first International Conference on Machine Learning , ser. ICML ’04. ACM, 2004
2004
Earlier work this paper cites.
P. Abbeel and A. Y. Ng, “Exploration and apprenticeship learning in reinforcement learning,” in Proceedings of the 22nd international conference on Machine learning . ACM, 2005, pp. 1–8
2005
Earlier work this paper cites.
N. Chentanez, A. G. Barto, and S. P. Singh, “Intrinsically motivated reinforcement learning,” in Advances in neural information processing systems , 2005, pp. 1281–1288
2005
Earlier work this paper cites.
S. M. LaValle, Planning Algorithms . New York, NY, USA: Cambridge University Press, 2006
2006
Earlier work this paper cites.
W. G. Najm, J. D. Smith, M. Yanagisawa et al. , “Pre-crash scenario typology for crash avoidance research,” United States. National Highway Traffic Safety Administration, Tech. Rep., 2007
2007
Earlier work this paper cites.
D. Silver, R. S. Sutton, and M. Müller, “Sample-based learning and search with permanent and transient memories,” in Proceedings of the 25th international conference on Machine learning . ACM, 2008, pp. 968–975
2008
Earlier work this paper cites.
B. Recht, “A tour of reinforcement learning: The view from continuous control,” Annual Review of Control, Robotics, and Autonomous Systems , 2008
2008
Earlier work this paper cites.
Y. Kuwata, J. Teo, G. Fiore, S. Karaman, E. Frazzoli, and J. P. How, “Real-time motion planning with applications to autonomous urban driving,” IEEE Transactions on Control Systems Technology , vol. 17, no. 5, pp. 1105–1118, 2009
2009
Earlier work this paper cites.
S. J. Russell and P. Norvig, Artificial intelligence: a modern approach (3rd edition) . Prentice Hall, 2009
2009
Earlier work this paper cites.
M. E. Taylor and P. Stone, “Transfer learning for reinforcement learning domains: A survey,” Journal of Machine Learning Research , vol. 10, no. Jul, pp. 1633–1685, 2009
2009
Earlier work this paper cites.
L. Buşoniu, R. Babuška, and B. Schutter, “Multi-agent reinforcement learning: An overview,” in Innovations in Multi-Agent Systems and Applications - 1 , ser. Studies in Computational Intelligence, D. Srinivasan and L. Jain, Eds. Springer Berlin Heidelberg, 2010, vol. 310
2010
Earlier work this paper cites.
S. Ross and D. Bagnell, “Efficient reductions for imitation learning,” in Proceedings of the thirteenth international conference on artificial intelligence and statistics , 2010, pp. 661–668
2010
Earlier work this paper cites.
S. Devlin and D. Kudenko, “Theoretical considerations of potential-based reward shaping for multi-agent systems,” in Proceedings of the 10th International Conference on Autonomous Agents and Multiagent Systems (AAMAS) , 2011
2011
Earlier work this paper cites.
D. C. K. Ngai and N. H. C. Yung, “A multiple-goal reinforcement learning method for complex vehicle overtaking maneuvers,” IEEE Transactions on Intelligent Transportation Systems , vol. 12, no. 2, pp. 509–522, 2011
2011
Earlier work this paper cites.
M. Wiering and M. van Otterlo, Eds., Reinforcement Learning: State-of-the-Art . Springer, 2012
2012
Earlier work this paper cites.
D. M. Roijers, P. Vamplew, S. Whiteson, and R. Dazeley, “A survey of multi-objective sequential decision-making,” Journal of Artificial Intelligence Research , vol. 48, pp. 67–113, 2013
2013
Earlier work this paper cites.
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller, “Deterministic policy gradient algorithms,” in ICML , 2014
2014
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in Advances in Neural Information Processing Systems 27 , 2014, pp. 2672–2680
2014
Earlier work this paper cites.
M. Cutler, T. J. Walsh, and J. P. How, “Reinforcement learning with multi-fidelity simulators,” in 2014 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2014, pp. 3888–3895
2014
Earlier work this paper cites.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski et al. , “Human-level control through deep reinforcement learning,” Nature , vol. 518, 2015
2015
Earlier work this paper cites.
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz, “Trust region policy optimization,” in International Conference on Machine Learning , 2015, pp. 1889–1897
2015
Earlier work this paper cites.
R. S. Sutton and A. G. Barto, “Reinforcement learning an introduction–second edition, in progress (draft),” 2015
2015
Earlier work this paper cites.
K. Narasimhan, T. Kulkarni, and R. Barzilay, “Language understanding for text-based games using deep reinforcement learning,” in Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing . Association for Computational Linguistics, 2015, pp. 1–11
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
M. Colby and K. Tumer, “An evolutionary game theoretic analysis of difference evaluation functions,” in Proceedings of the 2015 Annual Conference on Genetic and Evolutionary Computation . ACM, 2015, pp. 1391–1398
2015
Earlier work this paper cites.
W. Böhmer, J. T. Springenberg, J. Boedecker, M. Riedmiller, and K. Obermayer, “Autonomous learning of state representations for control: An emerging field aims to autonomously learn state representations for reinforcement learning agents from their real-world sensor observations,” KI-Künstliche Intelligenz , vol. 29, no. 4, pp. 353–362, 2015
2015
Earlier work this paper cites.
M. Watter, J. Springenberg, J. Boedecker, and M. Riedmiller, “Embed to control: A locally linear latent dynamics model for control from raw images,” in Advances in neural information processing systems , 2015
2015
Cited alongside, same era.
N. Wahlström, T. B. Schön, and M. P. Deisenroth, “Learning deep dynamical models from image pixels,” IFAC-PapersOnLine , vol. 48, no. 28, pp. 1059–1064, 2015
2015
Cited alongside, same era.
M. Kuderer, S. Gulati, and W. Burgard, “Learning driving styles for autonomous vehicles from demonstration,” in Robotics and Automation (ICRA), 2015 IEEE International Conference on . IEEE, 2015, pp. 2641–2646
2015
Cited alongside, same era.
J. Garcıa and F. Fernández, “A comprehensive survey on safe reinforcement learning,” Journal of Machine Learning Research , vol. 16, no. 1, pp. 1437–1480, 2015
2015
Cited alongside, same era.
T. Lesort, N. Diaz-Rodriguez, J.-F. Goudou, and D. Filliat, “State representation learning for control: An overview,” Neural Networks , vol. 108, pp. 379 – 392, 2018
2018
Later among the works it cites.
B. Kang, Z. Jie, and J. Feng, “Policy optimization with demonstrations,” in International Conference on Machine Learning , 2018, pp. 2474–2483
2018
Later among the works it cites.
T. Hester, M. Vecerik, O. Pietquin, M. Lanctot, T. Schaul, B. Piot, D. Horgan, J. Quan, A. Sendonaris, I. Osband et al. , “Deep q-learning from demonstrations,” in Thirty-Second AAAI Conference on Artificial Intelligence , 2018
2018
Later among the works it cites.
S. Ibrahim and D. Nevin, “End-to-end framework for fast learning asynchronous agents,” in the 32nd Conference on Neural Information Processing Systems, Imitation Learning and its Challenges in Robotics workshop , 2018
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
B. Paden, M. Čáp, S. Z. Yong, D. Yershov, and E. Frazzoli, “A survey of motion planning and control techniques for self-driving urban vehicles,” IEEE Transactions on intelligent vehicles , vol. 1, no. 1, pp. 33–55, 2016
2016
Cited alongside, same era.
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra, “Continuous control with deep reinforcement learning.” in 4th International Conference on Learning Representations, ICLR 2016, San Juan, Puerto Rico, May 2-4, 2016, Conference Track Proceedings , Y. Bengio and Y. LeCun, Eds., 2016
2016
Cited alongside, same era.
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu, “Asynchronous methods for deep reinforcement learning,” in International Conference on Machine Learning , 2016
2016
Cited alongside, same era.
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot et al. , “Mastering the game of go with deep neural networks and tree search,” nature , vol. 529, no. 7587, pp. 484–489, 2016
2016
Cited alongside, same era.
P. Mannion, K. Mason, S. Devlin, J. Duggan, and E. Howley, “Multi-objective dynamic dispatch optimisation using multi-agent reinforcement learning,” in Proceedings of the 15th International Conference on Autonomous Agents and Multiagent Systems (AAMAS) , 2016
2016
Cited alongside, same era.
K. Mason, P. Mannion, J. Duggan, and E. Howley, “Applying multi-agent reinforcement learning to watershed management,” in Proceedings of the Adaptive and Learning Agents workshop (at AAMAS 2016) , 2016
2016
Cited alongside, same era.
J. Ho and S. Ermon, “Generative adversarial imitation learning,” in Advances in Neural Information Processing Systems , 2016, pp. 4565–4573
2016
Cited alongside, same era.
A. E. Sallab, M. Abdou, E. Perot, and S. Yogamani, “End-to-end deep reinforcement learning for lane keeping assist,” in MLITS, NIPS Workshop , vol. 2, 2016
2016
Cited alongside, same era.
E. Leurent, Y. Blanco, D. Efimov, and O.-A. Maillard, “A survey of state-action representations for autonomous driving,” HAL archives , 2018
2018
Later among the works it cites.
P. Wang, C.-Y. Chan, and A. de La Fortelle, “A reinforcement learning based approach for automated lane change maneuvers,” in 2018 IEEE Intelligent Vehicles Symposium (IV) . IEEE, 2018, pp. 1379–1384
2018
Later among the works it cites.
2018
Later among the works it cites.
H. Mania, A. Guy, and B. Recht, “Simple random search of static linear policies is competitive for reinforcement learning,” in Advances in Neural Information Processing Systems 31 , S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, Eds., 2018, pp. 1800–1809
2018
Later among the works it cites.
S. Shah, D. Dey, C. Lovett, and A. Kapoor, “Airsim: High-fidelity visual and physical simulation for autonomous vehicles,” in Field and Service Robotics . Springer, 2018, pp. 621–635
2018
Later among the works it cites.
P. A. Lopez, M. Behrisch, L. Bieker-Walz, J. Erdmann, Y.-P. Flötteröd, R. Hilbrich, L. Lücken, J. Rummel, P. Wagner, and E. Wießner, “Microscopic traffic simulation using sumo,” in The 21st IEEE International Conference on Intelligent Transportation Systems . IEEE, 2018
2018
Later among the works it cites.
C. Quiter and M. Ernst, “deepdrive/deepdrive: 2.0,” Mar. 2018. [Online]. Available: https://doi.org/10.5281/zenodo.1248998
2018
Later among the works it cites.
P. Henderson, R. Islam, P. Bachman, J. Pineau, D. Precup, and D. Meger, “Deep reinforcement learning that matters,” in Thirty-Second AAAI Conference on Artificial Intelligence , 2018
2018
Later among the works it cites.
K. Bousmalis, A. Irpan, P. Wohlhart, Y. Bai, M. Kelcey, M. Kalakrishnan, L. Downs, J. Ibarz, P. Pastor, K. Konolige et al. , “Using simulation and domain adaptation to improve efficiency of deep robotic grasping,” in 2018 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2018, pp. 4243–4250
2018
Later among the works it cites.
X. B. Peng, M. Andrychowicz, W. Zaremba, and P. Abbeel, “Sim-to-real transfer of robotic control with dynamics randomization,” in 2018 IEEE international conference on robotics and automation (ICRA) . IEEE, 2018, pp. 1–8
2018
Later among the works it cites.
2018
Later among the works it cites.
M. Al-Shedivat, T. Bansal, Y. Burda, I. Sutskever, I. Mordatch, and P. Abbeel, “Continuous adaptation via meta-learning in nonstationary and competitive environments,” in 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings . OpenReview.net, 2018
2018
Later among the works it cites.
D. Ha and J. Schmidhuber, “Recurrent world models facilitate policy evolution,” in Advances in Neural Information Processing Systems , 2018
2018
Later among the works it cites.
M. Bansal, A. Krizhevsky, and A. Ogale, “Chauffeurnet: Learning to drive by imitating the best and synthesizing the worst,” in Robotics: Science and Systems XV , 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
V. Talpaert., I. Sobh., B. R. Kiran., P. Mannion., S. Yogamani., A. El-Sallab., and P. Perez., “Exploring applications of deep reinforcement learning for real-world autonomous driving systems,” in Proceedings of the 14th International Joint Conference on Computer Vision, Imaging and Computer Graphics Theory and Applications - Volume 5 VISAPP: VISAPP, , INSTICC. SciTePress, 2019, pp. 564–572
2019
Later among the works it cites.
K. El Madawi, H. Rashed, A. El Sallab, O. Nasr, H. Kamel, and S. Yogamani, “Rgb and lidar fusion based 3d semantic segmentation for autonomous driving,” in 2019 IEEE Intelligent Transportation Systems Conference (ITSC) . IEEE, 2019, pp. 7–12
2019
Later among the works it cites.
M. Uřičář, P. Křížek, G. Sistu, and S. Yogamani, “Soilingnet: Soiling detection on automotive surround-view cameras,” in 2019 IEEE Intelligent Transportation Systems Conference (ITSC) . IEEE, 2019, pp. 67–72
2019
Later among the works it cites.
G. Sistu, I. Leang, S. Chennupati, S. Yogamani, C. Hughes, S. Milz, and S. Rawashdeh, “Neurall: Towards a unified visual perception model for automated driving,” in 2019 IEEE Intelligent Transportation Systems Conference (ITSC) . IEEE, 2019, pp. 796–803
2019
Later among the works it cites.
S. Yogamani, C. Hughes, J. Horgan, G. Sistu, P. Varley, D. O’Dea, M. Uricár, S. Milz, M. Simon, K. Amende et al. , “Woodscape: A multi-task, multi-camera fisheye dataset for autonomous driving,” in Proceedings of the IEEE International Conference on Computer Vision , 2019, pp. 9308–9318
2019
Later among the works it cites.
2019
Later among the works it cites.
M. Uřičář, P. Křížek, D. Hurych, I. Sobh, S. Yogamani, and P. Denny, “Yes, we gan: Applying adversarial techniques for autonomous driving,” Electronic Imaging , vol. 2019, no. 15, pp. 48–1, 2019
2019
Later among the works it cites.
C. Li and K. Czarnecki, “Urban driving with multi-objective deep reinforcement learning,” in Proceedings of the 18th International Conference on Autonomous Agents and MultiAgent Systems . International Foundation for Autonomous Agents and Multiagent Systems, 2019, pp. 359–367
2019
Later among the works it cites.
J. Chen, B. Yuan, and M. Tomizuka, “Model-free deep reinforcement learning for urban autonomous driving,” in 2019 IEEE Intelligent Transportation Systems Conference (ITSC) . IEEE, 2019, pp. 2765–2771
2019
Later among the works it cites.
2019
Later among the works it cites.
A. Kendall, J. Hawke, D. Janz, P. Mazur, D. Reda, J.-M. Allen, V.-D. Lam, A. Bewley, and A. Shah, “Learning to drive in a day,” in 2019 International Conference on Robotics and Automation (ICRA) . IEEE, 2019, pp. 8248–8254
2019
Later among the works it cites.
Nvidia, “Drive Constellation now available,” https://blogs.nvidia.com/blog/2019/03/18/drive-constellation-now-available/ , 2019, [accessed 14-April-2019]
2019
Later among the works it cites.
A. S. et al., “Multi-Agent Autonomous Driving Simulator built on top of TORCS,” https://github.com/madras-simulator/MADRaS , 2019, [Online; accessed 14-April-2019]
2019
Later among the works it cites.
E. Leurent, “A collection of environments for autonomous driving and tactical decision-making tasks,” https://github.com/eleurent/highway-env , 2019, [Online; accessed 14-April-2019]
2019
Later among the works it cites.
F. Rosique, P. J. Navarro, C. Fernández, and A. Padilla, “A systematic review of perception system and simulators for autonomous vehicles research,” Sensors , vol. 19, no. 3, p. 648, 2019
2019
Later among the works it cites.
F. C. German Ros, Vladlen Koltun and A. M. Lopez, “Carla autonomous driving challenge,” https://carlachallenge.org/ , 2019, [Online; accessed 14-April-2019]
2019
Later among the works it cites.
Y. Abeysirigoonawardena, F. Shkurti, and G. Dudek, “Generating adversarial driving scenarios in high-fidelity simulators,” in 2019 IEEE International Conference on Robotics and Automation (ICRA) . ICRA, 2019
2019
Later among the works it cites.
A. Bewley, J. Rigley, Y. Liu, J. Hawke, R. Shen, V.-D. Lam, and A. Kendall, “Learning to drive from simulation without real world labels,” in 2019 International Conference on Robotics and Automation (ICRA) . IEEE, 2019, pp. 4818–4824
2019
Later among the works it cites.
J. Zhang, L. Tai, P. Yun, Y. Xiong, M. Liu, J. Boedecker, and W. Burgard, “Vr-goggles for robots: Real-to-sim domain adaptation for visual control,” IEEE Robotics and Automation Letters , vol. 4, no. 2, pp. 1148–1155, 2019
2019
Later among the works it cites.
T. Buhet, E. Wirbel, and X. Perrotton, “Conditional vehicle trajectories prediction in carla urban environment,” in Proceedings of the IEEE International Conference on Computer Vision Workshops , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
Sergio Guadarrama, Anoop Korattikara, Oscar Ramirez, Pablo Castro, Ethan Holly, Sam Fishman, Ke Wang, Ekaterina Gonina, Neal Wu, Chris Harris, Vincent Vanhoucke, Eugene Brevdo, “TF-Agents: A library for reinforcement learning in tensorflow,” https://github.com/tensorflow/agents , 2018, [Online; accessed 25-June-2019]. [Online]. Available: https://github.com/tensorflow/agents
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
S. Kuutti, R. Bowden, Y. Jin, P. Barber, and S. Fallah, “A survey of deep learning applications to autonomous vehicle control,” IEEE Transactions on Intelligent Transportation Systems , 2020
2020
Closest in time.
R. Rădulescu, P. Mannion, D. M. Roijers, and A. Nowé, “Multi-objective multi-agent decision making: a utility-based analysis and survey,” Autonomous Agents and Multi-Agent Systems , vol. 34, no. 1, p. 10, 2020
2020
Closest in time.
P. Palanisamy, “Multi-agent connected autonomous driving using deep reinforcement learning,” in 2020 International Joint Conference on Neural Networks (IJCNN) . IEEE, 2020, pp. 1–7
2020
Closest in time.
S. Bhalla, S. Ganapathi Subramanian, and M. Crowley, “Deep multi agent reinforcement learning for autonomous driving,” in Advances in Artificial Intelligence , C. Goutte and X. Zhu, Eds. Cham: Springer International Publishing, 2020, pp. 67–78
2020
Closest in time.
C. Yu, X. Wang, X. Xu, M. Zhang, H. Ge, J. Ren, L. Sun, B. Chen, and G. Tan, “Distributed multiagent coordinated learning for autonomous driving in highways based on dynamic coordination graphs,” IEEE Transactions on Intelligent Transportation Systems , vol. 21, no. 2, pp. 735–748, 2020
2020
Closest in time.
D. Isele, R. Rahimi, A. Cosgun, K. Subramanian, and K. Fujimura, “Navigating occluded intersections with autonomous vehicles using deep reinforcement learning,” in 2018 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2018, pp. 2034–2039
2039
Closest in time.
H. Van Hasselt, A. Guez, and D. Silver, “Deep reinforcement learning with double q-learning.” in AAAI , vol. 16, 2016, pp. 2094–2100
2094
Closest in time.