Fetching the paper…
Reading the bibliography…
Evaluations of Deep Reinforcement Learning (DRL) methods are an integral part of scientific progress of the field.
On method overfitting
Emanuel Falkenauer · 1998
Earlier work this paper cites.
Towards a universal test suite for combinatorial auction algorithms
Kevin Leyton-Brown, Mark Pearson, and Yoav Shoham · 2000
Earlier work this paper cites.
Benchmarking optimization software with performance profiles
Elizabeth D. Dolan and Jorge J. Moré · 2002
Earlier work this paper cites.
Incremental natural actor-critic algorithms
Shalabh Bhatnagar, Mohammad Ghavamzadeh, Mark Lee, and Richard S Sutton · 2007
Earlier work this paper cites.
An empirical analysis of value function-based and policy search reinforcement learning
Shivaram Kalyanakrishnan and Peter Stone · 2009
Earlier work this paper cites.
Reinforcement learning design for cancer clinical trials
Yufan Zhao, Michael R. Kosorok, and Donglin Zeng · 2009
Earlier work this paper cites.
Protecting against evaluation overfitting in empirical reinforcement learning
Shimon Whiteson, Brian Tanner, Matthew E. Taylor, and Peter Stone · 2011
Earlier work this paper cites.
Max pressure control of a network of signalized intersections
Pravin Varaiya · 2013
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, P. Abbeel, Michael I. Jordan, and Philipp Moritz · 2015
Earlier work this paper cites.
Variability of capacity and traffic performance at urban and rural signalised intersections
Janusz Chodur, Krzysztof Ostrowski, and Marian Tracz · 2016
Earlier work this paper cites.
Hidden parameter markov decision processes: A semiparametric regression approach for discovering latent task parametrizations
Finale Doshi-Velez and George Dimitri Konidaris · 2016
Earlier work this paper cites.
Machine learning in genomic medicine: A review of computational problems and data sets
Michael K. K. Leung, Andrew Delong, Babak Alipanahi, and Brendan J. Frey · 2016
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Optimizing chemical reactions with deep reinforcement learning
Zhenpeng Zhou, Xiaocheng Li, and Richard N. Zare · 2017
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke Hoof, and David Meger · 2018
Earlier work this paper cites.
Deep reinforcement learning that matters
Peter Henderson, Riashat Islam, Philip Bachman, Joelle Pineau, Doina Precup, and David Meger · 2018
Cited alongside, same era.
Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, Timothy P. Lillicrap, and Martin A. Riedmiller · 2018
Cited alongside, same era.
Benchmarks for reinforcement learning in mixed-autonomy traffic
Eugene Vinitsky, Aboudy Kreidieh, Luc Le Flem, Nishant Kheterpal, Kathy Jang, Cathy Wu, Fangyu Wu, Richard Liaw, Eric Liang, and Alexandre M Bayen · 2018
Cited alongside, same era.
Intellilight: A reinforcement learning approach for intelligent traffic light control
Hua Wei, Guanjie Zheng, Huaxiu Yao, and Zhenhui Li · 2018
Cited alongside, same era.
Transfer learning in biomedical natural language processing: an evaluation of bert and elmo on ten benchmarking datasets
Yifan Peng, Shankai Yan, and Zhiyong Lu · 2019
Cited alongside, same era.
Model-based reinforcement learning for atari
Lukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski, Roy H. Campbell, K. Czechowski, D. Erhan, Chelsea Finn, Piotr Kozakowski, Sergey Levine, Ryan Sepassi, G. Tucker, and Henryk Michalewski · 2020
Later among the works it cites.
Automated traffic signal performance measures component details
Utah Department of Transportation · 2020
Later among the works it cites.
Attendlight: Universal attention-based reinforcement learning model for traffic signal control
Afshin Oroojlooy, Mohammadreza Nazari, Davood Hajinezhad, and Jorge Silva · 2020
Later among the works it cites.
Learning to locomote: Understanding how environment design matters for deep reinforcement learning
Daniele Reda, Tianxin Tao, and Michiel van de Panne · 2020
Later among the works it cites.
Deep reinforcement learning at the edge of the statistical precipice
Rishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C Courville, and Marc Bellemare · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Aadt unrounded, 2019
UDOT · 2019
Cited alongside, same era.
Learning phase competition for traffic signal control
Guanjie Zheng, Yuanhao Xiong, Xinshi Zang, J. Feng, Hua Wei, Huichu Zhang, Yong Li, Kai Xu, and Zhenhui Jessie Li · 2019
Cited alongside, same era.
What matters in on-policy reinforcement learning? a large-scale empirical study
Marcin Andrychowicz, Anton Raichuk, Piotr Stańczyk, Manu Orsini, Sertan Girgin, Raphael Marinier, Léonard Hussenot, Matthieu Geist, Olivier Pietquin, Marcin Michalski, et al · 2020
Cited alongside, same era.
Learning an interpretable traffic signal control policy
James Ault, Josiah P. Hanna, and Guni Sharon · 2020
Cited alongside, same era.
Measuring the reliability of reinforcement learning algorithms
Stephanie CY Chan, Samuel Fishman, John Canny, Anoop Korattikara, and Sergio Guadarrama · 2020
Cited alongside, same era.
Toward a thousand lights: Decentralized deep reinforcement learning for large-scale traffic signal control
Chacha Chen, Hua Wei, Nan Xu, Guanjie Zheng, Ming Yang, Yuanhao Xiong, Kai Xu, and Zhenhui Li · 2020
Cited alongside, same era.
Mot20: A benchmark for multi object tracking in crowded scenes
Patrick Dendorfer, Hamid Rezatofighi, Anton Milan, Javen Shi, Daniel Cremers, Ian Reid, Stefan Roth, Konrad Schindler, and Laura Leal-Taixé · 2020
Cited alongside, same era.
Reinforcement learning benchmarks for traffic signal control
James Ault and Guni Sharon · 2021
Later among the works it cites.
Carl: A benchmark for contextual and adaptive reinforcement learning
Carolin Benjamins, Theresa Eimer, Frederik Schubert, André Biedenkapp, Bodo Rosenhahn, Frank Hutter, and Marius Lindauer · 2021
Later among the works it cites.
Hyperparameters in contextual rl are highly situational
Theresa Eimer, Carolin Benjamins, and Marius Lindauer · 2021
Later among the works it cites.
Why generalization in rl is difficult: Epistemic pomdps and implicit partial observability
Dibya Ghosh, Jad Rahme, Aviral Kumar, Amy Zhang, Ryan P. Adams, and Sergey Levine · 2021
Later among the works it cites.
Mixed autonomous supervision in traffic signal control
Vindula Jayawardana, Anna Landler, and Cathy Wu · 2021
Later among the works it cites.
Contextualize me - the case for context in reinforcement learning
Caroline Benjamins, Theresa Eimer, Frederik Schubert, Aditya Mohan, Andr’e Biedenkapp, Bodo Rosenhahn, Frank Hutter, and Marius Thomas Lindauer · 2022
Closest in time.
Is high variance unavoidable in rl? a case study in continuous control
Johan Bjorck, Carla P Gomes, and Kilian Q Weinberger · 2022
Closest in time.
Udot data portal, 2022
Utah DOT · 2022
Closest in time.
Learning eco-driving strategies at signalized intersections
Vindula Jayawardana and Cathy Wu · 2022
Closest in time.
Open street map, 2022
Open Street Map · 2022
Closest in time.