Fetching the paper…
Reading the bibliography…
Autonomous driving is a multi-agent setting where the host vehicle must apply sophisticated negotiation skills with other road users when overtaking, giving way, merging, taking left and right turns and while pushing ahead in unstructured urban roadways.
Iterative solution of games by fictitious play
George W Brown · 1951
Earlier work this paper cites.
Dynamic programming and lagrange multipliers
Richard Bellman · 1956
Earlier work this paper cites.
Stochastic optimization of systems
VM Aleksandrov, VI Sysoev, and VV Shemeneva · 1968
Earlier work this paper cites.
Introduction to the mathematical theory of control processes , volume 2
Richard Bellman · 1971
Earlier work this paper cites.
Likelilood ratio gradient estimation: an overview
Peter W Glynn · 1987
Earlier work this paper cites.
A survey of solution techniques for the partially observed markov decision process
Chelsea C White III · 1991
Earlier work this paper cites.
Q-learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Michael L Littman · 1994
Earlier work this paper cites.
Dynamic programming and optimal control , volume 1
Dimitri P Bertsekas · 1995
Earlier work this paper cites.
Learning to play the game of chess
S. Thrun · 1995
Earlier work this paper cites.
Reinforcement learning: A survey
Leslie Pack Kaelbling, Michael L Littman, and Andrew W Moore · 1996
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Adversarial reinforcement learning
William Uther and Manuela Veloso · 1997
Cited alongside, same era.
Reinforcement learning: An introduction , volume 1
Richard S Sutton and Andrew G Barto · 1998
Cited alongside, same era.
A simple adaptive procedure leading to correlated equilibrium
S. Hart and A. Mas-Colell · 2000
Cited alongside, same era.
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder Singh · 2002
Cited alongside, same era.
R-max–a general polynomial time algorithm for near-optimal reinforcement learning
Ronen I Brafman and Moshe Tennenholtz · 2003
Cited alongside, same era.
Nash q-learning for general-sum stochastic games
Junling Hu and Michael P Wellman · 2003
Cited alongside, same era.
Online learning and online convex optimization
Shai Shalev-Shwartz · 2011
Later among the works it cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Later among the works it cites.
Reinforcement learning in robotics: A survey
Jens Kober, J Andrew Bagnell, and Jan Peters · 2013
Later among the works it cites.
Stochastic dual coordinate ascent methods for regularized loss
Shai Shalev-Shwartz and Tong Zhang · 2013
Later among the works it cites.
Stochastic gradient descent, weighted sampling, and the randomized kaczmarz algorithm
Deanna Needell, Rachel Ward, and Nati Srebro · 2014
Later among the works it cites.
Safe policy search for lifelong reinforcement learning with sublinear regret
Haitham Bou Ammar, Rasul Tutunov, and Eric Eaton · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Variance reduction techniques for gradient estimates in reinforcement learning
Evan Greensmith, Peter L Bartlett, and Jonathan Baxter · 2004
Cited alongside, same era.
Learning structured prediction models: A large margin approach
Ben Taskar, Vassil Chatalbashev, Daphne Koller, and Carlos Guestrin · 2005
Cited alongside, same era.
Prediction, learning, and games
N. Cesa-Bianchi and G. Lugosi · 2006
Cited alongside, same era.
If multi-agent learning is the answer, what is the question?
Yoav Shoham, Rob Powers, and Trond Grenager · 2007
Cited alongside, same era.
Reinforcement learning of motor skills with policy gradients
Jan Peters and Stefan Schaal · 2008
Cited alongside, same era.
Algorithms for reinforcement learning
Csaba Szepesvári · 2010
Cited alongside, same era.
Approximate value iteration with temporally extended actions
Timothy A. Mann, Doina Precup, and Shie Mannor · 2015
Later among the works it cites.
Human drivers are bumping into driverless cars and exposing a key flaw
Keith Naughton · 2015
Later among the works it cites.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel · 2015
Later among the works it cites.
End to end learning for self-driving cars
Mariusz Bojarski, Davide Del Testa, Daniel Dworakowski, Bernhard Firner, Beat Flepp, Prasoon Goyal, Lawrence D Jackel, Mathew Monfort, Urs Muller, Jiakai Zhang, et al · 2016
Closest in time.
Sdca without duality, regularization, and individual convexity
Shai Shalev-Shwartz · 2016
Closest in time.
A deep hierarchical approach to lifelong learning in minecraft
Chen Tessler, Shahar Givony, Tom Zahavy, Daniel J Mankowitz, and Shie Mannor · 2016
Closest in time.