Fetching the paper…
Reading the bibliography…
We study online learning in contextual pay-per-click auctions where at each of the $T$ rounds, the learner receives some context along with a set of ads and needs to make an estimate on their click-through rate (CTR) in order to run a second-price pay-per-click auction.
Asymptotically efficient adaptive allocation rules
Tze Leung Lai, Herbert Robbins, et al · 1985
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, Yoav Freund, and Robert E Schapire · 2002
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
Martin Zinkevich · 2003
Earlier work this paper cites.
Truthful auctions for pricing search keywords
Gagan Aggarwal, Ashish Goel, and Rajeev Motwani · 2006
Earlier work this paper cites.
Logarithmic regret algorithms for online convex optimization
Elad Hazan, Amit Agarwal, and Satyen Kale · 2007
Earlier work this paper cites.
The epoch-greedy algorithm for multi-armed bandits with side information
John Langford and Tong Zhang · 2007
Earlier work this paper cites.
The price of truthfulness for pay-per-click auctions
Nikhil R Devanur and Sham M Kakade · 2009
Earlier work this paper cites.
Efficient optimal learning for contextual bandits
Miroslav Dudik, Daniel Hsu, Satyen Kale, Nikos Karampatziakis, John Langford, Lev Reyzin, and Tong Zhang · 2011
Earlier work this paper cites.
Bayesian learning via stochastic gradient langevin dynamics
Max Welling and Yee W Teh · 2011
Earlier work this paper cites.
A truthful learning mechanism for contextual multi-slot sponsored search auctions with externalities
Nicola Gatti, Alessandro Lazaric, and Francesco Trovò · 2012
Earlier work this paper cites.
Taming the monster: A fast and simple algorithm for contextual bandits
Alekh Agarwal, Daniel Hsu, Satyen Kale, John Langford, Lihong Li, and Robert Schapire · 2014
Earlier work this paper cites.
Characterizing truthful multi-armed bandit mechanisms
Moshe Babaioff, Yogeshwer Sharma, and Aleksandrs Slivkins · 2014
Earlier work this paper cites.
Stochastic bandits with side observations on networks
Swapna Buccapatnam, Atilla Eryilmaz, and Ness B Shroff · 2014
Cited alongside, same era.
Truthful mechanisms with implicit payment computation
Moshe Babaioff, Robert D Kleinberg, and Aleksandrs Slivkins · 2015
Cited alongside, same era.
Practical contextual bandits with regression oracles
Dylan Foster, Alekh Agarwal, Miroslav Dudik, Haipeng Luo, and Robert Schapire · 2018
Cited alongside, same era.
Learning in repeated auctions with budgets: Regret minimization and equilibrium
Santiago R Balseiro and Yonatan Gur · 2019
Cited alongside, same era.
Beyond ucb: Optimal and efficient contextual bandits with regression oracles
Dylan Foster and Alexander Rakhlin · 2020
Cited alongside, same era.
Learning to bid optimally and efficiently in adversarial first-price auctions
The statistical complexity of interactive decision making
Dylan J Foster, Sham M Kakade, Jian Qian, and Alexander Rakhlin · 2021
Later among the works it cites.
Budget pacing in repeated auctions: Regret and efficiency without convergence
Jason Gaitonde, Yingkai Li, Bar Light, Brendan Lucier, and Aleksandrs Slivkins · 2022
Later among the works it cites.
Bypassing the monster: A faster and simpler optimal algorithm for contextual bandits under realizability
David Simchi-Levi and Yunzong Xu · 2022
Later among the works it cites.
Feel-good thompson sampling for contextual bandits and reinforcement learning
Tong Zhang · 2022
Later among the works it cites.
Contextual bandits with smooth regret: Efficient learning in continuous action spaces
Yinglun Zhu and Paul Mineiro · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yanjun Han, Zhengyuan Zhou, Aaron Flores, Erik Ordentlich, and Tsachy Weissman · 2020
Cited alongside, same era.
Bandit algorithms
Tor Lattimore and Csaba Szepesvári · 2020
Cited alongside, same era.
Taking a hint: How to leverage loss predictors in contextual bandits?
Chen-Yu Wei, Haipeng Luo, and Alekh Agarwal · 2020
Cited alongside, same era.
Upper counterfactual confidence bounds: a new optimism principle for contextual bandits
Yunbei Xu and Assaf Zeevi · 2020
Cited alongside, same era.
Robust auction design in the auto-bidding world
Santiago Balseiro, Yuan Deng, Jieming Mao, Vahab Mirrokni, and Song Zuo · 2021
Cited alongside, same era.
Efficient first-order contextual bandits: Prediction, allocation, and triangular discrimination
Dylan J Foster and Akshay Krishnamurthy · 2021
Cited alongside, same era.
The best of many worlds: Dual mirror descent for online allocation problems
Santiago R Balseiro, Haihao Lu, and Vahab Mirrokni · 2023
Closest in time.
Improved online learning algorithms for ctr prediction in ad auctions
Zhe Feng, Christopher Liaw, and Zixin Zhou · 2023
Closest in time.
Vcg mechanism design with unknown agent values under stochastic bandit feedback
Kirthevasan Kandasamy, Joseph E Gonzalez, Michael I Jordan, and Ion Stoica · 2023
Closest in time.
Autobidders with budget and roi constraints: Efficiency, regret, and pacing dynamics
Brendan Lucier, Sarath Pattathil, Aleksandrs Slivkins, and Mengxiao Zhang · 2023
Closest in time.
Learning to bid in repeated first-price auctions with budgets
Qian Wang, Zongjun Yang, Xiaotie Deng, and Yuqing Kong · 2023
Closest in time.
On the robustness of epoch-greedy in multi-agent contextual bandit mechanisms
Yinglun Xu, Bhuvesh Kumar, and Jacob Abernethy · 2023
Closest in time.