Fetching the paper…
Reading the bibliography…
We present a new type of acquisition functions for online decision making in multi-armed and contextual bandit problems with extreme payoffs.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William Thompson · 1933
Earlier work this paper cites.
Design and analysis of computer experiments
Jerome Sacks, William Welch, Toby Mitchell, and Henry Wynn · 1989
Earlier work this paper cites.
On the limited memory bfgs method for large scale optimization
Dong Liu and Jorge Nocedal · 1989
Earlier work this paper cites.
Sample mean based index policies with O(log n) regret for the multi-armed bandit problem
Rajeev Agrawal · 1995
Earlier work this paper cites.
Bayesian Q-learning
Richard Dearden, Nir Friedman, and Stuart Russell · 1998
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Peter Auer · 2002
Earlier work this paper cites.
Gaussian processes for machine learning
Carl Edward Rasmussen and Christopher Williams · 2006
Earlier work this paper cites.
Multi-fidelity optimization via surrogate modelling
Alexander Forrester, András Sóbester, and Andy Keane · 2007
Earlier work this paper cites.
Optimum design configuration of Savonius rotor through wind tunnel experiments
U.K. Saha, S. Thotla, and D. Maity · 2008
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
Varsha Dani, Thomas Hayes, and Sham Kakade · 2008
Earlier work this paper cites.
Gaussian process optimization in the bandit setting: No regret and experimental design
Niranjan Srinivas, Andreas Krause, Sham Kakade, and Matthias Seeger · 2009
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Lihong Li, Wei Chu, John Langford, and Robert Schapire · 2010
Earlier work this paper cites.
Survey of modeling and optimization strategies to solve high-dimensional design problems with computationally-expensive black-box functions
Songqing Shan and G Gary Wang · 2010
Earlier work this paper cites.
Convergence properties of the expected improvement algorithm with fixed mean and covariance functions
Emmanuel Vazquez and Julien Bect · 2010
Earlier work this paper cites.
Batch Bayesian optimization via simulation matching
Javad Azimi, Alan Fern, and Xiaoli Fern · 2010
Earlier work this paper cites.
Contextual Gaussian process bandit optimization
Andreas Krause and Cheng Soon Ong · 2011
Earlier work this paper cites.
Contextual bandits with linear payoff functions
Wei Chu, Lihong Li, Lev Reyzin, and Robert Schapire · 2011
Earlier work this paper cites.
An empirical evaluation of Thompson sampling
Olivier Chapelle and Lihong Li · 2011
Cited alongside, same era.
Practical variational inference for neural networks
Alex Graves · 2011
Cited alongside, same era.
Thompson sampling: An asymptotically optimal finite-time analysis
Emilie Kaufmann, Nathaniel Korda, and Rémi Munos · 2012
Cited alongside, same era.
Analysis of Thompson sampling for the multi-armed bandit problem
Shipra Agrawal and Navin Goyal · 2012
Cited alongside, same era.
Stochastic variational inference
Matthew Hoffman, David Blei, Chong Wang, and John Paisley · 2013
Cited alongside, same era.
Efficient Thompson sampling for online matrix-factorization recommendation
Jaya Kawale, Hung Bui, Branislav Kveton, Long Tran-Thanh, and Sanjay Chawla · 2015
Cited alongside, same era.
Sequential sampling strategy for extreme event statistics in nonlinear dynamical systems
Mustafa Mohamad and Themistoklis Sapsis · 2018
Later among the works it cites.
Parallelised Bayesian optimisation via Thompson sampling
Kirthevasan Kandasamy, Akshay Krishnamurthy, Jeff Schneider, and Barnabás Póczos · 2018
Later among the works it cites.
Carlos Riquelme, George Tucker, and Jasper Snoek · 2018
Later among the works it cites.
JAX: composable transformations of Python+NumPy programs, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Richard Sutton and Andrew Barto · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver · 2015
Cited alongside, same era.
Automatic differentiation in machine learning: a survey
Atilim Gunes Baydin, Barak A Pearlmutter, Alexey Andreyevich Radul, and Jeffrey Mark Siskind · 2015
Cited alongside, same era.
Deep exploration via bootstrapped dqn
Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy · 2016
Cited alongside, same era.
Multifidelity information fusion algorithms for high-dimensional systems and massive data sets
Paris Perdikaris, Daniele Venturi, and George Em Karniadakis · 2016
Cited alongside, same era.
Improving kriging surrogates of high-dimensional design models by partial least squares dimension reduction
Mohamed Amine Bouhlel, Nathalie Bartoli, Abdelkader Otsmane, and Joseph Morlier · 2016
Cited alongside, same era.
Thompson sampling is asymptotically optimal in general environments
Jan Leike, Tor Lattimore, Laurent Orseau, and Marcus Hutter · 2016
Cited alongside, same era.
Adversarial learning of a sampler based on an unnormalized distribution
Chunyuan Li, Ke Bai, Jianqiao Li, Guoyin Wang, Changyou Chen, and Lawrence Carin · 2019
Later among the works it cites.
Multi-fidelity classification using gaussian processes: Accelerating the prediction of large-scale computational models
Francisco Sahli Costabal, Paris Perdikaris, Ellen Kuhl, and Daniel Hurtado · 2019
Later among the works it cites.
Multifidelity and multiscale Bayesian framework for high-dimensional engineering design and calibration
Soumalya Sarkar, Sudeepta Mondal, Michael Joly, Matthew Lynch, Shaunak Bopardikar, Ranadip Acharya, and Paris Perdikaris · 2019
Later among the works it cites.
Physics-informed neural networks for cardiac activation mapping
Francisco Sahli Costabal, Yibo Yang, Paris Perdikaris, Daniel Hurtado, and Ellen Kuhl · 2020
Later among the works it cites.
Output-weighted optimal sampling for Bayesian regression and rare event statistics using few samples
Themistoklis Sapsis · 2020
Later among the works it cites.
Output-weighted importance sampling for Bayesian experimental design and uncertainty quantification
Antoine Blanchard and Themistoklis Sapsis · 2020
Later among the works it cites.
Informative path planning for anomaly detection in environment exploration and monitoring
Antoine Blanchard and Themistoklis Sapsis · 2020
Later among the works it cites.
Efficiently sampling functions from gaussian process posteriors
James Wilson, Viacheslav Borovitskiy, Alexander Terenin, Peter Mostowsky, and Marc Deisenroth · 2020
Later among the works it cites.
Using smart city tools to evaluate the effectiveness of a low emissions zone in spain: Madrid central
Irene Lebrusán and Jamal Toutouh · 2020
Later among the works it cites.
Bayesian optimization with output-weighted importance sampling
Antoine Blanchard and Themistoklis Sapsis · 2021
Closest in time.
Periodic-gp: Learning periodic world with gaussian process bandits
Hengrui Cai, Zhihao Cen, Ling Leng, and Rui Song · 2021
Closest in time.