Fetching the paper…
Reading the bibliography…
Estimates of predictive uncertainty are important for accurate model-based planning and reinforcement learning.
A new vector partition of the probability score
Murphy, A. H · 1973
Earlier work this paper cites.
Present position and potential developments: Some personal views: Statistical theory: The prequential approach
Dawid, A. P · 1984
Earlier work this paper cites.
A neuro-dynamic programming approach to retailer inventory management
Van Roy, B., Bertsekas, D. P., Lee, Y., and Tsitsiklis, J. N · 1997
Earlier work this paper cites.
Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods
Platt, J. et al · 1999
Earlier work this paper cites.
Reinforcement learning for spoken dialogue systems
Singh, S. P., Kearns, M. J., Litman, D. J., and Walker, M. A · 2000
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Auer, P., Cesa-Bianchi, N., and Fischer, P · 2002
Earlier work this paper cites.
Transforming classifier scores into accurate multiclass probability estimates
Zadrozny, B. and Elkan, C · 2002
Earlier work this paper cites.
Weather forecasting with ensemble methods
Gneiting, T. and Raftery, A. E · 2005
Earlier work this paper cites.
Predicting good probabilities with supervised learning
Niculescu-Mizil, A. and Caruana, R · 2005
Earlier work this paper cites.
Using bayesian model averaging to calibrate forecast ensembles
Raftery, A. E., Gneiting, T., Balabdaoui, F., and Polakowski, M · 2005
Earlier work this paper cites.
Strictly proper scoring rules, prediction, and estimation
Gneiting, T. and Raftery, A. E · 2007
Earlier work this paper cites.
Probabilistic forecasts, calibration and sharpness
Gneiting, T., Balabdaoui, F., and Raftery, A. E · 2007
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Li, L., Chu, W., Langford, J., and Schapire, R. E · 2010
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
Deisenroth, M. and Rasmussen, C. E · 2011
Cited alongside, same era.
Playing games against nature: optimal policies for renewable resource allocation
Ermon, S., Conrad, J., Gomes, C. P., and Selman, B · 2012
Cited alongside, same era.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Cited alongside, same era.
Climate change and food systems
Vermeulen, S. J., Campbell, B. M., and Ingram, J. S · 2012
Cited alongside, same era.
Calibrated structured prediction
Kuleshov, V. and Liang, P · 2015
Cited alongside, same era.
Novel decompositions of proper scoring rules for classification: Score adjustment as precursor to calibration
Kull, M. and Flach, P · 2015
Cited alongside, same era.
Concrete dropout
Gal, Y., Hron, J., and Kendall, A · 2017
Later among the works it cites.
Densely connected convolutional networks
Huang, G., Liu, Z., Van Der Maaten, L., and Weinberger, K. Q · 2017
Later among the works it cites.
Estimating uncertainty online against an adversary
Kuleshov, V. and Ermon, S · 2017
Later among the works it cites.
Sample-efficient reinforcement learning with stochastic ensemble value expansion
Buckman, J., Hafner, D., Tucker, G., Brevdo, E., and Lee, H · 2018
Later among the works it cites.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Chua, K., Calandra, R., McAllister, R., and Levine, S · 2018
Later among the works it cites.
Model-based reinforcement learning via meta-policy optimization
Clavera, I., Rothfuss, J., Schulman, J., Fujita, Y., Asfour, T., and Abbeel, P · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
From predictive models to instructional policies
Rollinson, J. and Brunskill, E · 2015
Cited alongside, same era.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Cited alongside, same era.
Learning and policy search in stochastic dynamical systems with bayesian neural networks
Depeweg, S., Hernández-Lobato, J. M., Doshi-Velez, F., and Udluft, S · 2016
Cited alongside, same era.
Sparse gaussian processes for bayesian optimization
McIntire, M., Ratner, D., and Ermon, S · 2016
Cited alongside, same era.
Epopt: Learning robust neural network policies using model ensembles
Rajeswaran, A., Ghotra, S., Ravindran, B., and Levine, S · 2016
Cited alongside, same era.
Safe model-based reinforcement learning with stability guarantees
Berkenkamp, F., Turchetta, M., Schoellig, A., and Krause, A · 2017
Cited alongside, same era.
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Later among the works it cites.
Synthesizing neural network controllers with probabilistic model based reinforcement learning
Higuera, J. C. G., Meger, D., and Dudek, G · 2018
Later among the works it cites.
Accurate uncertainties for deep learning using calibrated regression
Kuleshov, V., Fenner, N., and Ermon, S · 2018
Later among the works it cites.
Model-ensemble trust-region policy optimization
Kurutach, T., Clavera, I., Duan, Y., Tamar, A., and Abbeel, P · 2018
Later among the works it cites.
Individualized sepsis treatment using reinforcement learning
Saria, S · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Later among the works it cites.