Fetching the paper…
Reading the bibliography…
We introduce a unified probabilistic framework for solving sequential decision making problems ranging from Bayesian optimisation to contextual bandits and reinforcement learning.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
W. R. Thompson · 1933
Earlier work this paper cites.
The application of bayesian methods for seeking the extremum. vol. 2, 1978
J. Moćkus, V. Tiesis, and A. Źilinskas · 1978
Earlier work this paper cites.
Evolutionary principles in self-referential learning, or on learning how to learn: the meta-meta-… hook
J. Schmidhuber · 1987
Earlier work this paper cites.
Efficient learning and planning within the dyna framework
J. Peng and R. J. Williams · 1993
Earlier work this paper cites.
Robot see, robot do: An overview of robot imitation
P. Bakker and Y. Kuniyoshi · 1996
Earlier work this paper cites.
Global versus local search in constrained optimization of computer models
M. Schonlau, W. J. Welch, and D. R. Jones · 1998
Earlier work this paper cites.
Gaussian processes in machine learning
C. E. Rasmussen · 2003
Earlier work this paper cites.
Prediction, learning, and games
N. Cesa-Bianchi and G. Lugosi · 2006
Earlier work this paper cites.
Multi-task gaussian process prediction
E. V. Bonilla, K. M. A. Chai, and C. K. Williams · 2008
Earlier work this paper cites.
Mapreduce: simplified data processing on large clusters
J. Dean and S. Ghemawat · 2008
Earlier work this paper cites.
Factorization meets the neighborhood: a multifaceted collaborative filtering model
Y. Koren · 2008
Earlier work this paper cites.
Bayesian probabilistic matrix factorization using markov chain monte carlo
R. Salakhutdinov and A. Mnih · 2008
Earlier work this paper cites.
Large-scale parallel collaborative filtering for the netflix prize
Y. Zhou, D. Wilkinson, R. Schreiber, and R. Pan · 2008
Earlier work this paper cites.
Gaussian process optimization in the bandit setting: No regret and experimental design
N. Srinivas, A. Krause, S. M. Kakade, and M. Seeger · 2009
Earlier work this paper cites.
Variational learning of inducing variables in sparse gaussian processes
M. Titsias · 2009
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
M. Deisenroth and C. E. Rasmussen · 2011
Earlier work this paper cites.
Contextual gaussian process bandit optimization
A. Krause and C. S. Ong · 2011
Earlier work this paper cites.
A survey of monte carlo tree search methods
C. B. Browne, E. Powley, D. Whitehouse, S. M. Lucas, P. I. Cowling, P. Rohlfshagen, S. Tavener, D. Perez, S. Samothrakis, and S. Colton · 2012
Earlier work this paper cites.
Svdfeature: a toolkit for feature-based collaborative filtering
T. Chen, W. Zhang, Q. Lu, K. Chen, Z. Zheng, and Y. Yu · 2012
Earlier work this paper cites.
Practical bayesian optimization of machine learning algorithms
J. Snoek, H. Larochelle, and R. P. Adams · 2012
Earlier work this paper cites.
Structure discovery in nonparametric regression through compositional kernel search
D. Duvenaud, J. R. Lloyd, R. Grosse, J. B. Tenenbaum, and Z. Ghahramani · 2013
Earlier work this paper cites.
Local low-rank matrix approximation
J. Lee, S. Kim, G. Lebanon, and Y. Singer · 2013
Earlier work this paper cites.
Intriguing properties of neural networks
C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, D. Erhan, I. Goodfellow, and R. Fergus · 2013
Cited alongside, same era.
Explaining and harnessing adversarial examples
I. J. Goodfellow, J. Shlens, and C. Szegedy · 2014
Cited alongside, same era.
Weight uncertainty in neural networks
C. Blundell, J. Cornebise, K. Kavukcuoglu, and D. Wierstra · 2015
Cited alongside, same era.
Learning continuous control policies by stochastic value gradients
N. Heess, G. Wayne, D. Silver, T. Lillicrap, T. Erez, and Y. Tassa · 2015
Cited alongside, same era.
Siamese neural networks for one-shot image recognition
G. Koch, R. Zemel, and R. Salakhutdinov · 2015
Cited alongside, same era.
Multiplicative normalizing flows for variational bayesian neural networks
C. Louizos and M. Welling · 2017
Later among the works it cites.
Vfunc: a deep generative model for functions
P. Bachman, R. Islam, A. Sordoni, and Z. Ahmed · 2018
Later among the works it cites.
Distributed distributional deterministic policy gradients
G. Barth-Maron, M. W. Hoffman, D. Budden, W. Dabney, D. Horgan, A. Muldal, N. Heess, and T. Lillicrap · 2018
Later among the works it cites.
Federated meta-learning for recommendation
F. Chen, Z. Dong, Z. Li, and X. He · 2018
Later among the works it cites.
Avoiding latent variable collapse with generative skip models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Autorec: Autoencoders meet collaborative filtering
S. Sedhain, A. K. Menon, S. Sanner, and L. Xie · 2015
Cited alongside, same era.
Learning to learn by gradient descent by gradient descent
M. Andrychowicz, M. Denil, S. Gomez, M. W. Hoffman, D. Pfau, T. Schaul, B. Shillingford, and N. De Freitas · 2016
Cited alongside, same era.
C. Beattie, J. Z. Leibo, D. Teplyashin, T. Ward, M. Wainwright, H. Küttler, A. Lefrancq, S. Green, V. Valdés, A. Sadik, J. Schrittwieser, K. Anderson, S. York, M. Cant, A. Cain, A. Bolton, S. Gaffney, H. King, D. Hassabis, S. Legg, and S. Petersen · 2016
Cited alongside, same era.
H. Edwards and A. Storkey · 2016
Cited alongside, same era.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Y. Gal and Z. Ghahramani · 2016
Cited alongside, same era.
The movielens datasets: History and context
F. M. Harper and J. A. Konstan · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
A. B. Dieng, Y. Kim, A. M. Rush, and D. M. Blei · 2018
Later among the works it cites.
Scalable distributed deep-rl with importance weighted actor-learner architectures
L. Espeholt, H. Soyer, R. Munos, K. Simonyan, V. Mnih, T. Ward, Y. Doron, V. Firoiu, T. Harley, I. Dunning, S. Legg, and K. Kavukcuoglu · 2018
Later among the works it cites.
Probabilistic model agnostic meta-learning
C. Finn, K. Xu, and S. Levine · 2018
Later among the works it cites.
Recasting gradient-based meta-learning as hierarchical bayes
E. Grant, C. Finn, S. Levine, T. Darrell, and T. Griffiths · 2018
Later among the works it cites.
Multi-task deep reinforcement learning with popart
M. Hessel, H. Soyer, L. Espeholt, W. Czarnecki, S. Schmitt, and H. v. Hasselt · 2018
Later among the works it cites.
The variational homoencoder: Learning to learn high capacity generative models from few examples
L. B. Hewitt, M. I. Nye, A. Gane, T. Jaakkola, and J. B. Tenenbaum · 2018
Later among the works it cites.
Empirical evaluation of neural process objectives
T. A. Le, H. Kim, M. Garnelo, D. Rosenbaum, J. Schwarz, and Y. W. Teh · 2018
Later among the works it cites.
Deep online learning via meta-learning: Continual adaptation for model-based rl
A. Nagabandi, C. Finn, and S. Levine · 2018
Later among the works it cites.
Deep bayesian bandits showdown: An empirical comparison of bayesian deep networks for thompson sampling
C. Riquelme, G. Tucker, and J. Snoek · 2018
Later among the works it cites.
Uncovering surprising behaviors in reinforcement learning via worst-case analysis
A. Ruderman, R. Everett, B. Sikder, H. Soyer, J. Uesato, A. Kumar, C. Beattie, and P. Kohli · 2018
Later among the works it cites.
Meta reinforcement learning with latent variable gaussian processes
S. Sæmundsson, K. Hofmann, and M. P. Deisenroth · 2018
Later among the works it cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Later among the works it cites.
Unsupervised predictive memory in a goal-directed agent
G. Wayne, C. Hung, D. Amos, M. Mirza, A. Ahuja, A. Grabska-Barwinska, J. W. Rae, P. Mirowski, J. Z. Leibo, A. Santoro, M. Gemici, M. Reynolds, T. Harley, J. Abramson, S. Mohamed, D. J. Rezende, D. Saxton, A. Cain, C. Hillier, D. Silver, K. Kavukcuoglu, M. Botvinick, D. Hassabis, and T. P. Lillicrap · 2018
Later among the works it cites.
Recurrent experience replay in distributed reinforcement learning
S. Kapturowski, G. Ostrovski, W. Dabney, J. Quan, and R. Munos · 2019
Closest in time.
Attentive neural processes
H. Kim, A. Mnih, J. Schwarz, M. Garnelo, A. Eslami, D. Rosenbaum, O. Vinyals, and Y. W. Teh · 2019
Closest in time.
R. Mendonca, A. Gupta, R. Kralev, P. Abbeel, S. Levine, and C. Finn · 2019
Closest in time.
Amortized bayesian meta-learning
S. Ravi and A. Beatson · 2019
Closest in time.