Fetching the paper…
Reading the bibliography…
In this work, we initiate the idea of using denoising diffusion models to learn priors for online decision making problems.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R Thompson · 1933
Earlier work this paper cites.
Estimation of the mean of a multivariate normal distribution
Charles M Stein · 1981
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin Puterman · 1994
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Richard Sutton and Andrew Barto · 1998
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Peter Auer · 2002
Earlier work this paper cites.
Convergence of the monte carlo expectation maximization for curved exponential families
Gersende Fort and Eric Moulines · 2003
Earlier work this paper cites.
Action elimination and stopping conditions for the multi-armed bandit and reinforcement learning problems
Eyal Even-Dar, Shie Mannor, and Yishay Mansour · 2006
Earlier work this paper cites.
Generalized SURE for exponential families: Applications to regularization
Yonina C Eldar · 2008
Earlier work this paper cites.
Monte-Carlo SURE: A black-box optimization of regularization parameters for general denoising algorithms
Sathish Ramani, Thierry Blu, and Michael Unser · 2008
Earlier work this paper cites.
Pure exploration in multi-armed bandits problems
Sebastien Bubeck, Remi Munos, and Gilles Stoltz · 2009
Earlier work this paper cites.
Best arm identification in multi-armed bandits
Jean-Yves Audibert, Sebastien Bubeck, and Remi Munos · 2010
Earlier work this paper cites.
Gaussian sampling by local perturbations
George Papandreou and Alan L Yuille · 2010
Earlier work this paper cites.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Sébastien Bubeck and Nicolò Cesa-Bianchi · 2012
Earlier work this paper cites.
An empirical evaluation of Thompson sampling
Olivier Chapelle and Lihong Li · 2012
Earlier work this paper cites.
Combinatorial multi-armed bandit: General framework and applications
Wei Chen, Yajun Wang, and Yang Yuan · 2013
Earlier work this paper cites.
Bounded regret for finite-armed structured bandits
Tor Lattimore and Remi Munos · 2014
Earlier work this paper cites.
ipinyou global rtb bidding algorithm competition dataset
Hairen Liao, Lingxiao Peng, Zhenchuan Liu, and Xuehua Shen · 2014
Earlier work this paper cites.
Latent bandits
Odalric-Ambrym Maillard and Shie Mannor · 2014
Earlier work this paper cites.
Learning to optimize via posterior sampling
Daniel Russo and Benjamin Van Roy · 2014
Earlier work this paper cites.
Spectral bandits for smooth graph functions
Michal Valko, Remi Munos, Branislav Kveton, and Tomas Kocak · 2014
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli · 2015
Earlier work this paper cites.
Stochastic online shortest path routing: The value of feedback
Mohammad Sadegh Talebi, Zhenhua Zou, Richard Combes, Alexandre Proutiere, and Mikael Johansson · 2017
Earlier work this paper cites.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms
Han Xiao, Kashif Rasul, and Roland Vollgraf · 2017
Earlier work this paper cites.
Unsupervised learning with Stein’s unbiased risk estimator
Christopher A Metzler, Ali Mousavi, Reinhard Heckel, and Richard G Baraniuk · 2018
Cited alongside, same era.
Deep Bayesian bandits showdown: An empirical comparison of Bayesian deep networks for Thompson sampling
Carlos Riquelme, George Tucker, and Jasper Snoek · 2018
Cited alongside, same era.
A tutorial on thompson sampling
Daniel J Russo, Benjamin Van Roy, Abbas Kazerouni, Ian Osband, Zheng Wen, et al · 2018
Cited alongside, same era.
Thompson sampling for combinatorial semi-bandits
Siwei Wang and Wei Chen · 2018
Cited alongside, same era.
Information-theoretic confidence bounds for reinforcement learning
Xiuyuan Lu and Benjamin Van Roy · 2019
Cited alongside, same era.
Generative modeling by estimating gradients of the data distribution
Progressive distillation for fast sampling of diffusion models
Tim Salimans and Jonathan Ho · 2021
Later among the works it cites.
Top- k k extreme contextual bandits with arm hierarchy
Rajat Sen, Alexander Rakhlin, Lexing Ying, Rahul Kidambi, Dean Foster, Daniel Hill, and Inderjit Dhillon · 2021
Later among the works it cites.
Bayesian decision-making under misspecified priors with applications to meta-learning
Max Simchowitz, Christopher Tosh, Akshay Krishnamurthy, Daniel J Hsu, Thodoris Lykouris, Miro Dudik, and Robert E Schapire · 2021
Later among the works it cites.
Score-based generative modeling through stochastic differential equations
Yang Song, Jascha Sohfl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole · 2021
Later among the works it cites.
Protein sequence design with deep generative models
Zachary Wu, Kadina E Johnston, Frances H Arnold, and Kevin K Yang · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yang Song and Stefano Ermon · 2019
Cited alongside, same era.
Extending stein’s unbiased risk estimator to train deep denoisers with correlated pairs of noisy images
Magauiya Zhussip, Shakarim Soltanayev, and Se Young Chun · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Meta-learning with stochastic linear bandits
Leonardo Cella, Alessandro Lazaric, and Massimiliano Pontil · 2020
Cited alongside, same era.
A unified approach to translate classical bandit algorithms to the structured bandit setting
Samarth Gupta, Shreyas Chaudhari, Subhojyoti Mukherjee, Gauri Joshi, and Osman Yağan · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Cited alongside, same era.
Latent bandits revisited
Joey Hong, Branislav Kveton, Manzil Zaheer, Yinlam Chow, Amr Ahmed, and Craig Boutilier · 2020
Cited alongside, same era.
Anurag Ajay, Yilun Du, Abhi Gupta, Joshua Tenenbaum, Tommi Jaakkola, and Pulkit Agrawal · 2022
Later among the works it cites.
Estimating the optimal covariance with imperfect mean in diffusion probabilistic models
Fan Bao, Chongxuan Li, Jiacheng Sun, Jun Zhu, and Bo Zhang · 2022
Later among the works it cites.
GENIE: Higher-Order Denoising Diffusion Solvers
Tim Dockhorn, Arash Vahdat, and Karsten Kreis · 2022
Later among the works it cites.
Vincent Dutordoir, Alan Saul, Zoubin Ghahramani, and Fergus Simpson · 2022
Later among the works it cites.
Diffusion models as plug-and-play priors
Alexandros Graikos, Nikolay Malkin, Nebojsa Jojic, and Dimitris Samaras · 2022
Later among the works it cites.
Improving conditional score-based generation with calibrated classification and joint training
Paul Kuo-Ming Huang, Si-An Chen, and Hsuan-Tien Lin · 2022
Later among the works it cites.
Planning with diffusion for flexible behavior synthesis
Michael Janner, Yilun Du, Joshua B Tenenbaum, and Sergey Levine · 2022
Later among the works it cites.
Elucidating the design space of diffusion-based generative models
Tero Karras, Miika Aittala, Timo Aila, and Samuli Laine · 2022
Later among the works it cites.
Denoising diffusion restoration models
Bahjat Kawar, Michael Elad, Stefano Ermon, and Jiaming Song · 2022
Later among the works it cites.
Convergence of score-based generative modeling for general data distributions
Holden Lee, Jianfeng Lu, and Yixin Tan · 2022
Later among the works it cites.
Repaint: Inpainting using denoising diffusion probabilistic models
Andreas Lugmayr, Martin Danelljan, Andres Romero, Fisher Yu, Radu Timofte, and Luc Van Gool · 2022
Later among the works it cites.
Probabilistic machine learning: an introduction
Kevin P Murphy · 2022
Later among the works it cites.
Meta-learning via classifier (-free) guidance
Elvis Nava, Seijin Kobayashi, Yifei Yin, Robert K Katzschmann, and Benjamin F Grewe · 2022
Later among the works it cites.
Metalearning linear bandits by prior update
Amit Peleg, Naama Pearl, and Ron Meir · 2022
Later among the works it cites.
Photorealistic text-to-image diffusion models with deep language understanding
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S Sara Mahdavi, Rapha Gontijo Lopes, et al · 2022
Later among the works it cites.
Solving inverse problems in medical imaging with score-based generative models
Yang Song, Liyue Shen, Lei Xing, and Stefano Ermon · 2022
Later among the works it cites.
Conditioning and sampling in variational diffusion models for speech super-resolution
Chin-Yun Yu, Sung-Lin Yeh, György Fazekas, and Hao Tang · 2022
Later among the works it cites.
Fast sampling of diffusion models via operator learning
Hongkai Zheng, Weili Nie, Arash Vahdat, Kamyar Azizzadenesheli, and Anima Anandkumar · 2022
Later among the works it cites.