Fetching the paper…
Reading the bibliography…
Value-based methods constitute a fundamental methodology in planning and deep reinforcement learning (RL).
Neuronlike adaptive elements that can solve difficult learning control problems
Andrew G Barto, Richard S Sutton, and Charles W Anderson · 1983
Earlier work this paper cites.
Numerical linear algebra
Lloyd N Trefethen and III David Bau · 1997
Earlier work this paper cites.
Unscented filtering and nonlinear estimation
Simon J Julier and Jeffrey K Uhlmann · 2004
Earlier work this paper cites.
Bounded real-time dynamic programming: Rtdp with monotone upper bounds and performance guarantees
H Brendan McMahan, Maxim Likhachev, and Geoffrey J Gordon · 2005
Earlier work this paper cites.
Focused real-time dynamic programming for mdps: Squeezing more out of a heuristic
Trey Smith and Reid Simmons · 2006
Earlier work this paper cites.
Consensus algorithms for double-integrator dynamics
Wei Ren and Randal W Beard · 2008
Earlier work this paper cites.
Exact matrix completion via convex optimization
Emmanuel J Candès and Benjamin Recht · 2009
Earlier work this paper cites.
Spectral regularization algorithms for learning large incomplete matrices
Rahul Mazumder, Trevor Hastie, and Robert Tibshirani · 2010
Earlier work this paper cites.
Closing the learning-planning loop with predictive state representations
Byron Boots, Sajid M Siddiqi, and Geoffrey J Gordon · 2011
Earlier work this paper cites.
Goal-directed online learning of predictive models
Sylvie CW Ong, Yuri Grinberg, and Joelle Pineau · 2011
Earlier work this paper cites.
A survey of monte carlo tree search methods
Cameron B Browne, Edward Powley, Daniel Whitehouse, Simon M Lucas, Peter I Cowling, Philipp Rohlfshagen, Stephen Tavener, Diego Perez, Spyridon Samothrakis, and Simon Colton · 2012
Earlier work this paper cites.
Predictive state representations: A new theory for modeling dynamical systems
Satinder Singh, Michael James, and Matthew Rudary · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Sparse multi-task reinforcement learning
Daniele Calandriello, Alessandro Lazaric, and Marcello Restelli · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Community detection in general stochastic block models: Fundamental limits and efficient algorithms for recovery
Emmanuel Abbe and Colin Sandon · 2015
Cited alongside, same era.
Learning and planning with timing information in markov decision processes
Pierre-Luc Bacon, Borja Balle, and Doina Precup · 2015
Cited alongside, same era.
Matrix estimation by universal singular value thresholding
Sourav Chatterjee et al · 2015
Cited alongside, same era.
Yudong Chen and Martin J Wainwright · 2015
Cited alongside, same era.
Efficient high-dimensional stochastic optimal motion control using tensor-train decomposition
Alex A Gorodetsky, Sertac Karaman, and Youssef M Marzouk · 2015
Cited alongside, same era.
Deep reinforcement learning with double q-learning
Hado Van Hasselt, Arthur Guez, and David Silver · 2016
Later among the works it cites.
Thy friend is my friend: Iterative collaborative filtering for sparse matrix estimation
Christian Borgs, Jennifer Chayes, Christina E Lee, and Devavrat Shah · 2017
Later among the works it cites.
Contextual decision processes with low bellman rank are pac-learnable
Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2017
Later among the works it cites.
Count-based exploration with neural density models
Georg Ostrovski, Marc G Bellemare, Aäron van den Oord, and Rémi Munos · 2017
Later among the works it cites.
Harnessing structures in big data via guaranteed low-rank matrix estimation
Yudong Chen and Yuejie Chi · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Value function approximation via low-rank models
Hao Yi Ong · 2015
Cited alongside, same era.
Dueling network architectures for deep reinforcement learning
Ziyu Wang, Tom Schaul, Matteo Hessel, Hado Van Hasselt, Marc Lanctot, and Nando De Freitas · 2015
Cited alongside, same era.
Automated synthesis of low-rank control systems from sc-ltl specifications using tensor-train decompositions
John Irvin Alora, Alex Gorodetsky, Sertac Karaman, Youssef Marzouk, and Nathan Lowry · 2016
Cited alongside, same era.
An overview of low-rank matrix recovery from incomplete observations
Mark A Davenport and Justin Romberg · 2016
Cited alongside, same era.
Jianqing Fan, Weichen Wang, and Yiqiao Zhong · 2016
Cited alongside, same era.
Matrix completion has no spurious local minimum
Rong Ge, Jason D Lee, and Tengyu Ma · 2016
Cited alongside, same era.
Will Dabney, Georg Ostrovski, David Silver, and Rémi Munos · 2018
Later among the works it cites.
The leave-one-out approach for matrix completion: Primal and dual analysis
Lijun Ding and Yudong Chen · 2018
Later among the works it cites.
High-dimensional stochastic optimal control using continuous tensor decompositions
Alex Gorodetsky, Sertac Karaman, and Youssef Marzouk · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, et al · 2018
Later among the works it cites.
A continuous analogue of the tensor-train decomposition
Alex Gorodetsky, Sertac Karaman, and Youssef Marzouk · 2019
Closest in time.
Devavrat Shah, Qiaomin Xie, and Zhi Xu · 2019
Closest in time.
Underactuated robotics: Algorithms for walking, running, swimming, flying, and manipulation
Russ Tedrake · 2019
Closest in time.
ME-Net: Towards effective adversarial robustness with matrix estimation
Yuzhe Yang, Guo Zhang, Dina Katabi, and Zhi Xu · 2019
Closest in time.