Double q-learning
Hado V Hasselt · 2010
Cited alongside, same era.
Adaptive algorithms and stochastic approximations
Albert Benveniste, Michel Métivier, and Pierre Priouret · 2012
Cited alongside, same era.
A tutorial on linear function approximators for dynamic programming and reinforcement learning
Alborz Geramifard, Thomas J Walsh, Stefanie Tellex, Girish Chowdhary, Nicholas Roy, and Jonathan P How · 2013
Cited alongside, same era.
Openai gym
Original
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Deep reinforcement learning with double q-learning
Hado Van Hasselt, Arthur Guez, and David Silver · 2016
Cited alongside, same era.
Averaged-dqn: Variance reduction and stabilization for deep reinforcement learning
Oron Anschel, Nir Baram, and Nahum Shimkin · 2017
Cited alongside, same era.
Fastest convergence for q-learning
Original
Adithya M Devraj and Sean P Meyn · 2017
Cited alongside, same era.
Weighted double q-learning
Zongzhang Zhang, Zhiyuan Pan, and Mykel J Kochenderfer · 2017
Cited alongside, same era.
A finite time analysis of temporal difference learning with linear function approximation
Jalaj Bhandari, Daniel Russo, and Raghav Singal · 2018
Cited alongside, same era.
Finite sample analyses for td(0) with function approximation
Gal Dalal, Balázs Szörényi, Gugan Thoppe, and Shie Mannor · 2018
Cited alongside, same era.
Finite sample analysis of two-timescale stochastic approximation with applications to reinforcement learning
Gal Dalal, Balázs Szörényi, Gugan Thoppe, and Shie Mannor · 2018
Cited alongside, same era.
Linear stochastic approximation: How far does constant step-size and iterate averaging go?
Chandrashekar Lakshminarayanan and Csaba Szepesvari · 2018
Cited alongside, same era.