Fetching the paper…
Reading the bibliography…
Reinforcement learning constantly deals with hard integrals, for example when computing expectations in policy evaluation and policy iteration.
Asynchronous Methods for Deep Reinforcement Learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T. P., Harley, T., Silver, D., and Kavukcuoglu, K. (2016) · 1937
Earlier work this paper cites.
The Monte Carlo Method
Metropolis, N. and Ulam, S. (1949) · 1949
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
Polyak, B. T. (1964) · 1964
Earlier work this paper cites.
Monte Carlo analysis of reactivity coefficients in fast reactors general theory and applications
Miller, L. B. (1967) · 1967
Earlier work this paper cites.
Some problems in Monte Carlo optimization
Rubinstein, R. Y. (1969) · 1969
Earlier work this paper cites.
Linear Optimal Control Systems
Kwakernaak, H. and Sivan, R. (1972) · 1972
Earlier work this paper cites.
Uniformly distributed sequences with an additional uniform property
Sobol, I. M. (1976) · 1976
Earlier work this paper cites.
Quasi-Monte Carlo methods and pseudo-random numbers
Niederreiter, H. (1978) · 1978
Earlier work this paper cites.
Random Number Generation and Quasi-Monte Carlo Methods
Niederreiter, H. (1992) · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J. (1992) · 1992
Earlier work this paper cites.
Quasi-Monte Carlo Methods in Numerical Finance
Joy, C., Boyle, P., and Tan, K. (1996) · 1996
Earlier work this paper cites.
Biped dynamic walking using reinforcement learning
Benbrahim, H. and Franklin, J. A. (1997) · 1997
Earlier work this paper cites.
Scrambled net variance for integrals of smooth functions
Owen, A. B. (1997) · 1997
Earlier work this paper cites.
Reinforcement learning for continuous action using stochastic gradient ascent
Kimura, H. (1998) · 1998
Earlier work this paper cites.
On the L2-discrepancy for anchored boxes
Matoušek, J. (1998) · 1998
Earlier work this paper cites.
Scrambling Sobol’ and Niederreiter–Xing Points
Owen, A. B. (1998) · 1998
Earlier work this paper cites.
When are Quasi-Monte carlo algorithms efficient for high dimensional integrals?
Sloan, I. H. and Woźniakowski, H. (1998) · 1998
Earlier work this paper cites.
Some New Perspectives on the Method of Control Variates
Glynn, P. W. and Szechtman, R. (2002) · 2000
Earlier work this paper cites.
Extensible Lattice Sequences for Quasi-Monte Carlo Quadrature
Hickernell, F. J., Hong, H. S., L’Écuyer, P., and Lemieux, C. (2000) · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y. (2000) · 2000
Earlier work this paper cites.
Variance Reduction Techniques for Gradient Estimates in Reinforcement Learning
Greensmith, E., Bartlett, P. L., and Baxter, J. (2001a) · 2001
Earlier work this paper cites.
Variance Reduction Techniques for Gradient Estimates in Reinforcement Learning
Greensmith, E., Bartlett, P. L., and Baxter, J. (2001b) · 2001
Earlier work this paper cites.
Efficient Gradient Estimation for Motor Control Learning
Lawrence, G., Cowan, N., and Russell, S. (2002) · 2002
Earlier work this paper cites.
Sufficient conditions for fast quasi-Monte Carlo convergence
Papageorgiou, A. (2003) · 2003
Earlier work this paper cites.
Reinforcement learning for humanoid robotics
Peters, J., Vijayakumar, S., and Schaal, S. (2003) · 2003
Earlier work this paper cites.
Monte Carlo Methods in Financial Engineering
Glasserman, P. (2004) · 2004
Earlier work this paper cites.
Policy gradient reinforcement learning for fast quadrupedal locomotion
Kohl, N. and Stone, P. (2004) · 2004
Cited alongside, same era.
Monte Carlo Statistical Methods
Robert, C. P. and Casella, G. (2004) · 2004
Cited alongside, same era.
Solving deep memory POMDPs with recurrent policy gradients
Wierstra, D., Foerster, A., Peters, J., and Schmidhuber, J. (2007) · 2007
Cited alongside, same era.
Constructing sobol sequences with better Two-Dimensional projections
Joe, S. and Kuo, F. Y. (2008) · 2008
Cited alongside, same era.
On array-RQMC for Markov chains: Mapping alternatives and convergence rates
L’Ecuyer, P., Lécot, C., and L’Archevêque-Gaudet, A. (2009) · 2008
Cited alongside, same era.
A Randomized Quasi-Monte Carlo Simulation Method for Markov Chains
L’Ecuyer, P., Lécot, C., and Tuffin, B. (2008) · 2008
Cited alongside, same era.
Interpolated Policy Gradient: Merging On-Policy and Off-Policy Gradient Estimation for Deep Reinforcement Learning
Gu, S., Lillicrap, T., Turner, R. E., Ghahramani, Z., Schölkopf, B., and Levine, S. (2017) · 2017
Later among the works it cites.
Spinning Up in Deep Reinforcement Learning
Achiam, J. (2018) · 2018
Later among the works it cites.
Quasi-Monte Carlo Variational Inference
Buchholz, A., Wenzel, F., and Mandt, S. (2018) · 2018
Later among the works it cites.
Global Convergence of Policy Gradient Methods for the Linear Quadratic Regulator
Fazel, M., Ge, R., Kakade, S. M., and Mesbahi, M. (2018) · 2018
Later among the works it cites.
Addressing Function Approximation Error in Actor-Critic Methods
Fujimoto, S., van Hoof, H., and Meger, D. (2018) · 2018
Later among the works it cites.
Backpropagation through the Void: Optimizing control variates for black-box gradient estimation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Reinforcement learning of motor skills with policy gradients
Peters, J. and Schaal, S. (2008) · 2008
Cited alongside, same era.
Quasi-Monte Carlo Methods with Applications in Finance
L’Ecuyer, P. (2009) · 2009
Cited alongside, same era.
A Stochastic Gradient Method with an Exponential Convergence Rate for Finite Training Sets
Le Roux, N., Schmidt, M., and Bach, F. R. (2012) · 2012
Cited alongside, same era.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y. (2012) · 2012
Cited alongside, same era.
Accelerating Stochastic Gradient Descent using Predictive Variance Reduction
Johnson, R. and Zhang, T. (2013) · 2013
Cited alongside, same era.
Stochastic dual coordinate ascent methods for regularized loss minimization
Shalev-Shwartz, S. and Zhang, T. (2013) · 2013
Cited alongside, same era.
Grathwohl, W., Choi, D., Wu, Y., Roeder, G., and Duvenaud, D. (2018) · 2018
Later among the works it cites.
Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S. (2018) · 2018
Later among the works it cites.
Accelerating Stochastic Gradient Descent for Least Squares Regression
Jain, P., Kakade, S. M., Kidambi, R., Netrapalli, P., and Sidford, A. (2018) · 2018
Later among the works it cites.
On the insufficiency of existing momentum schemes for Stochastic Optimization
Kidambi, R., Netrapalli, P., Jain, P., and Kakade, S. M. (2018) · 2018
Later among the works it cites.
Action-dependent Control Variates for Policy Optimization via Stein Identity
Liu, H., Feng, Y., Mao, Y., Zhou, D., Peng, J., and Liu, Q. (2018) · 2018
Later among the works it cites.
Stochastic Variance-Reduced Policy Gradient
Papini, M., Binaghi, D., Canonaco, G., Pirotta, M., and Restelli, M. (2018) · 2018
Later among the works it cites.
Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations
Rajeswaran, A., Kumar, V., Gupta, A., Vezzani, G., Schulman, J., Todorov, E., and Levine, S. (2018) · 2018
Later among the works it cites.
Reward Estimation for Variance Reduction in Deep Reinforcement Learning
Romoff, J., Henderson, P., Piche, A., Francois-Lavet, V., and Pineau, J. (2018) · 2018
Later among the works it cites.
The Mirage of Action-Dependent Baselines in Reinforcement Learning
Tucker, G., Bhupatiraju, S., Gu, S., Turner, R. E., Ghahramani, Z., and Levine, S. (2018) · 2018
Later among the works it cites.
Quasi-Monte Carlo Methods Applied to Tau-Leaping in Stochastic Biological Systems
Beentjes, C. H. L. and Baker, R. E. (2019) · 2019
Later among the works it cites.
Monte Carlo Gradient Estimation in Machine Learning
Mohamed, S., Rosca, M., Figurnov, M., and Mnih, A. (2019) · 2019
Later among the works it cites.
PyTorch: An Imperative Style, High-Performance Deep Learning Library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S. (2019) · 2019
Later among the works it cites.
A tour of reinforcement learning: The view from continuous control
Recht, B. (2019) · 2019
Later among the works it cites.
SVRG for Policy Evaluation with Fewer Gradient Evaluations
Peng, Z., Touati, A., Vincent, P., and Precup, D. (2020) · 2020
Later among the works it cites.
Sample Factory: Egocentric 3D Control from Pixels at 100000 FPS with Asynchronous Reinforcement Learning
Petrenko, A., Huang, Z., Kumar, T., Sukhatme, G., and Koltun, V. (2020) · 2020
Later among the works it cites.
Advances in Importance Sampling
Elvira, V. and Martino, L. (2021) · 2021
Later among the works it cites.
Brax - A Differentiable Physics Engine for Large Scale Rigid Body Simulation
Freeman, C. D., Frey, E., Raichuk, A., Girgin, S., Mordatch, I., and Bachem, O. (2021) · 2021
Later among the works it cites.
Quasi-Monte Carlo Quasi-Newton in Variational Bayes
Liu, S. and Owen, A. B. (2021) · 2021
Later among the works it cites.
Isaac gym: High performance gpu-based physics simulation for robot learning
Makoviychuk, V., Wawrzyniak, L., Guo, Y., Lu, M., Storey, K., Macklin, M., Hoeller, D., Rudin, N., Allshire, A., Handa, A., and State, G. (2021) · 2021
Later among the works it cites.
Variance Reduction with Array-RQMC for Tau-Leaping Simulation of Stochastic Biological and Chemical Reaction Networks
Puchhammer, F., Ben Abdellah, A., and L’Ecuyer, P. (2021) · 2021
Later among the works it cites.