Fetching the paper…
Reading the bibliography…
Traffic signal control is an important problem in urban mobility with a significant potential of economic and environmental impact.
Soft actor-critic for discrete action settings, 2019
Petros Christodoulou · 1910
Earlier work this paper cites.
Benchmarking batch deep reinforcement learning algorithms
Scott Fujimoto, Edoardo Conti, Mohammad Ghavamzadeh, and Joelle Pineau · 1910
Earlier work this paper cites.
Dynamic programming
Richard Bellman · 1966
Earlier work this paper cites.
Optimal control of oversaturated store-and-forward transportation networks
GC D’ans and DC Gazis · 1976
Earlier work this paper cites.
The sydney coordinated adaptive traffic (scat) system philosophy and benefits
A.G. Sims and K.W. Dobinson · 1980
Earlier work this paper cites.
Alvinn: An autonomous land vehicle in a neural network
Dean A Pomerleau · 1988
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
Richard S Sutton · 1991
Earlier work this paper cites.
Estimating the mean and variance of the target probability distribution
David A Nix and Andreas S Weigend · 1994
Earlier work this paper cites.
Stable function approximation in dynamic programming
Geoffrey J Gordon · 1995
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Y Ng, Daishi Harada, and Stuart Russell · 1999
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Damien Ernst, Pierre Geurts, and Louis Wehenkel · 2005
Earlier work this paper cites.
Metric Spaces
M. O’Searcoid · 2006
Earlier work this paper cites.
The origin-destination matrix estimation problem: analysis and computations
Anders Peterson · 2007
Earlier work this paper cites.
Batch reinforcement learning
Sascha Lange, Thomas Gabel, and Martin Riedmiller · 2012
Earlier work this paper cites.
Pac optimal exploration in continuous space markov decision processes
Jason Pazis and Ronald Parr · 2013
Earlier work this paper cites.
Max pressure control of a network of signalized intersections
Pravin Varaiya · 2013
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Handbook of discrete and computational geometry
Csaba D Toth, Joseph O’Rourke, and Jacob E Goodman · 2017
Cited alongside, same era.
A survey on reinforcement learning models and algorithms for traffic signal control
Kok-Lim Alvin Yau, Junaid Qadir, Hooi Ling Khoo, Mee Hong Ling, and Peter Komisarczuk · 2017
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke Hoof, and David Meger · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Microscopic traffic simulation using sumo
Pablo Alvarez Lopez, Michael Behrisch, Laura Bieker-Walz, Jakob Erdmann, Yun-Pang Flötteröd, Robert Hilbrich, Leonhard Lücken, Johannes Rummel, Peter Wagner, and Evamarie Wießner · 2018
Cited alongside, same era.
Arthur Argenson and Gabriel Dulac-Arnold · 2020
Later among the works it cites.
Toward a thousand lights: Decentralized deep reinforcement learning for large-scale traffic signal control
Chacha Chen, Hua Wei, Nan Xu, Guanjie Zheng, Ming Yang, Yuanhao Xiong, Kai Xu, and Zhenhui Li · 2020
Later among the works it cites.
Morel: Model-based offline reinforcement learning
Rahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, and Thorsten Joachims · 2020
Later among the works it cites.
Conservative q-learning for offline reinforcement learning, 2020
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tomoki Nishi, Keisuke Otaki, Keiichiro Hayakawa, and Takayoshi Yoshimura · 2018
Cited alongside, same era.
An algorithmic perspective on imitation learning
Takayuki Osa, Joni Pajarinen, Gerhard Neumann, J Andrew Bagnell, Pieter Abbeel, Jan Peters, et al · 2018
Cited alongside, same era.
Dissipation of stop-and-go waves via control of autonomous vehicles: Field experiments
Raphael E. Stern, Shumo Cui, Maria Laura Delle Monache, Rahul Bhadani, Matt Bunting, Miles Churchill, Nathaniel Hamilton, R’mani Haulcy, Hannah Pohlmann, Fangyu Wu, Benedetto Piccoli, Benjamin Seibold, Jonathan Sprinkle, and Daniel B. Work · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Intellilight: A reinforcement learning approach for intelligent traffic light control
Hua Wei, Guanjie Zheng, Huaxiu Yao, and Zhenhui Li · 2018
Cited alongside, same era.
Building a large-scale microscopic road network traffic simulator in apache spark
Zishan Fu, Jia Yu, and Mohamed Sarwat · 2019
Cited alongside, same era.
When to trust your model: Model-based policy optimization
Michael Janner, Justin Fu, Marvin Zhang, and Sergey Levine · 2019
Cited alongside, same era.
Later among the works it cites.
A game theoretic framework for model based reinforcement learning
Aravind Rajeswaran, Igor Mordatch, and Vikash Kumar · 2020
Later among the works it cites.
Deepaveragers: Offline reinforcement learning by solving derived non-parametric mdps
Aayam Shrestha, Stefan Lee, Prasad Tadepalli, and Alan Fern · 2020
Later among the works it cites.
Mopo: Model-based offline policy optimization, 2020
Tianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon, James Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma · 2020
Later among the works it cites.
A minimalist approach to offline reinforcement learning
Scott Fujimoto and Shixiang Shane Gu · 2021
Later among the works it cites.
Revisiting design choices in model-based offline reinforcement learning
Cong Lu, Philip J Ball, Jack Parker-Holder, Michael A Osborne, and Stephen J Roberts · 2021
Later among the works it cites.
Offline reinforcement learning for autonomous driving with safety and exploration enhancement
Tianyu Shi, Dong Chen, Kaian Chen, and Zhaojian Li · 2021
Later among the works it cites.
Behrad Toghi, Rodolfo Valiente, Dorsa Sadigh, Ramtin Pedarsani, and Yaser P Fallah · 2021
Later among the works it cites.
Flow: A modular learning framework for mixed autonomy traffic
Cathy Wu, Abdul Rahman Kreidieh, Kanaad Parvate, Eugene Vinitsky, and Alexandre M Bayen · 2021
Later among the works it cites.
Combo: Conservative offline model-based policy optimization
Tianhe Yu, Aviral Kumar, Rafael Rafailov, Aravind Rajeswaran, Sergey Levine, and Chelsea Finn · 2021
Later among the works it cites.
https://github.com/spotify/annoy
ANNOY: Approximate nearest neighbors in c++/python · 2022
Closest in time.
https://github.com/hiive/hiivemdptoolbox
Markov decision process (MDP) toolbox for python · 2022
Closest in time.
Deploying traffic smoothing cruise controllers learned from trajectory data
Nathan Lichtlé, Eugene Vinitsky, Matthew Nice, Benjamin Seibold, Dan Work, and Alexandre M Bayen · 2022
Closest in time.
Reinforcement learning (RL) algorithms are quite finicky
Andrew Ng · 2022
Closest in time.