Fetching the paper…
Reading the bibliography…
Value-based reinforcement learning (RL) methods like Q-learning have shown success in a variety of domains.
Memory approaches to reinforcement learning in non-Markovian domains
L. Lin and T. Mitchell · 1992
Earlier work this paper cites.
Q-learning
C. Watkins and P. Dayan · 1992
Earlier work this paper cites.
Reinforcement learning with high-dimensional, continuous actions
L. Baird III and A. Klopf · 1993
Earlier work this paper cites.
Combining trust region and line search techniques
J. Nocedal and Y. Yuan · 1998
Earlier work this paper cites.
Tree based discretization for continuous state space reinforcement learning
W. Uther and M. Veloso · 1998
Earlier work this paper cites.
Q-learning in continuous state and action spaces
C. Gaskett, D. Wettergreen, and A. Zelinsky · 1999
Earlier work this paper cites.
A computational study of search strategies for mixed integer programming
J. Linderoth and M. Savelsbergh · 1999
Earlier work this paper cites.
The cross-entropy method for combinatorial and continuous optimization
R. Rubinstein · 1999
Earlier work this paper cites.
Practical reinforcement learning in continuous spaces
W. Smart and L. Kaelbling · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. Sutton, D. McAllester, S. Singh P, and Y. Mansour · 2000
Earlier work this paper cites.
Online facility location
A. Meyerson · 2001
Earlier work this paper cites.
Envelope theorems for arbitrary choice sets
P. Milgrom and I. Segal · 2002
Earlier work this paper cites.
Continuous-action Q-learning
J. Millán, D. Posenato, and E. Dedieu · 2002
Earlier work this paper cites.
The cross entropy method for fast policy search
S. Mannor, R. Rubinstein, and Y. Gat · 2003
Earlier work this paper cites.
Numerical optimization
J. Nocedal and S. Wright · 2006
Earlier work this paper cites.
Policy gradient methods for robotics
J. Peters and S. Schaal · 2006
Earlier work this paper cites.
Q-learning with linear function approximation
F. Melo and M. Ribeiro · 2007
Earlier work this paper cites.
Reinforcement learning in continuous action spaces through sequential Monte Carlo methods
A. Lazaric, M. Restelli, and A. Bonarini · 2008
Earlier work this paper cites.
Playing Atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller · 2013
Earlier work this paper cites.
Convergence of a Q-learning variant for continuous states and actions
S. Carden · 2014
Cited alongside, same era.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
M. Puterman · 2014
Cited alongside, same era.
Deterministic policy gradient algorithms
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller · 2014
Cited alongside, same era.
Deep reinforcement learning in large discrete action spaces
G. Dulac-Arnold, R. Evans, H. van Hasselt, P. Sunehag, T. Lillicrap, J. Hunt, T. Mann, T. Weber, T. Degris, and B. Coppin · 2015
Cited alongside, same era.
Benchmarking deep reinforcement learning for continuous control
Y. Duan, X. Chen, R. Houthooft, J. Schulman, and P. Abbeel · 2016
Cited alongside, same era.
Continuous deep Q-learning with model-based acceleration
Output range analysis for deep feedforward neural networks
S. Dutta, S. Jha, S. Sankaranarayanan, and A. Tiwari · 2018
Later among the works it cites.
Global convergence of policy gradient methods for the linear quadratic regulator
M. Fazel, R. Ge, S. Kakade, and M. Mesbahi · 2018
Later among the works it cites.
Deep neural networks and mixed integer linear optimization
M. Fischetti and J. Jo · 2018
Later among the works it cites.
Addressing function approximation error in actor-critic methods
S. Fujimoto, H. van Hoof, and D. Meger · 2018
Later among the works it cites.
Horizon: Facebook’s open source applied reinforcement learning platform
J. Gauci, E. Conti, Y. Liang, K. Virochsiri, Y. He, Z. Kaden, V. Narayanan, and X. Ye · 2018
Later among the works it cites.
The SCIP Optimization Suite 6.0
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Gu, T. Lillicrap, I. Sutskever, and S. Levine · 2016
Cited alongside, same era.
Deep reinforcement learning with double Q-learning
H. Van Hasselt, A. Guez, and D. Silver · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Cited alongside, same era.
Input convex neural networks
B. Amos, L. Xu, and Z. Kolter · 2017
Cited alongside, same era.
Maximum resilience of artificial neural networks
C.-H. Cheng, G. Nührenberg, and H. Ruess · 2017
Cited alongside, same era.
Reluplex: An efficient SMT solver for verifying deep neural networks
G. Katz, C. Barrett, D. L. Dill, K. Julian, and M. J. Kochenderfer · 2017
Cited alongside, same era.
An approach to reachability analysis for feed-forward ReLU neural networks
A. Lomuscio and L. Maganti · 2017
Cited alongside, same era.
A. Gleixner, M. Bastubbe, L. Eifler, T. Gally, G. Gamrath, R. L. Gottwald, G. Hendel, C. Hojny, T. Koch, M. E. Lübbecke, S. J. Maher, M. Miltenberger, B. Müller, M. E. Pfetsch, C. Puchert, D. Rehfeldt, F. Schlösser, C. Schubert, F. Serrano, Y. Shinano, J. M. Viernickel, M. Walter, F. Wegscheider, J. T. Witt, and J. Witzig · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Later among the works it cites.
QT-Opt: Scalable deep reinforcement learning for vision-based robotic manipulation
D. Kalashnikov, A. Irpan, P. Pastor, J. Ibarz, A. Herzog, E. Jang, D. Quillen, E. Holly, M. Kalakrishnan, V. Vanhoucke, and S. Levine · 2018
Later among the works it cites.
Actor-Expert: A framework for using action-value methods in continuous action spaces
S. Lim, A. Joseph, L. Le, Y. Pan, and M. White · 2018
Later among the works it cites.
D. Quillen, E. Jang, O. Nachum, C. Finn, J. Ibarz, and S. Levine · 2018
Later among the works it cites.
Bounding and counting linear regions of deep neural networks
T. Serra, C. Tjandraatmadja, and S. Ramalingam · 2018
Later among the works it cites.
Towards fast computation of certified robustness for ReLU networks
T. Weng, H. Zhang, H. Chen, Z. Song, C. Hsieh, D. Boning, I. Dhillon, and L. Daniel · 2018
Later among the works it cites.
Strong mixed-integer programming formulations for trained neural networks
R. Anderson, J. Huchette, W. Ma, C. Tjandraatmadja, and J. P. Vielma · 2019
Closest in time.
V12.1: Users manual for CPLEX
IBM ILOG CPLEX · 2019
Closest in time.
Challenges of real-world reinforcement learning
G. Dulac-Arnold, D. Mankowitz, and T. Hester · 2019
Closest in time.
Gurobi optimizer reference manual, 2019
Gurobi · 2019
Closest in time.
Equivalent and approximate transformations of deep neural networks
A. Kumar, T. Serra, and S. Ramalingam · 2019
Closest in time.
Evaluating robustness of neural networks with mixed integer programming
V. Tjeng, K. Xiao, and R. Tedrake · 2019
Closest in time.