Fetching the paper…
Reading the bibliography…
Restless multi-armed bandits (RMAB) has been widely used to model constrained sequential decision making problems, where the state of each restless arm evolves according to a Markov chain and each state transition generates a scalar reward.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E Terry · 1952
Earlier work this paper cites.
Restless Bandits: Activity Allocation in A Changing World
Peter Whittle · 1988
Earlier work this paper cites.
Discrete-time controlled markov processes with average cost criterion: A survey
Aristotle Arapostathis, Vivek S Borkar, Emmanuel Fernández-Gaucherand, Mrinal K Ghosh, and Steven I Marcus · 1993
Earlier work this paper cites.
Probability Inequalities for Sums of Bounded Random Variables
Wassily Hoeffding · 1994
Earlier work this paper cites.
The Complexity of Optimal Queueing Network Control
Christos H Papadimitriou and John N Tsitsiklis · 1994
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L Puterman · 1994
Earlier work this paper cites.
Constrained Markov decision processes , volume 7
Eitan Altman · 1999
Earlier work this paper cites.
Restless Bandits, Linear Programming Relaxations, and A Primal-Dual Index Heuristic
Dimitris Bertsimas and José Niño-Mora · 2000
Earlier work this paper cites.
Near-Optimal Regret Bounds for Reinforcement Learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
Adaptive Learning of Uncontrolled Restless Bandits with Logarithmic Regret
Cem Tekin and Mingyan Liu · 2011
Earlier work this paper cites.
Regret Bounds for Restless Markov Bandits
Ronald Ortner, Daniil Ryabko, Peter Auer, and Rémi Munos · 2012
Earlier work this paper cites.
The k-armed dueling bandits problem
Yisong Yue, Josef Broder, Robert Kleinberg, and Thorsten Joachims · 2012
Earlier work this paper cites.
Fine-tuning language models from human preferences
Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 2012
Earlier work this paper cites.
Restless multi-armed bandits under time-varying activation constraints for dynamic spectrum access
Kobi Cohen, Qing Zhao, and Anna Scaglione · 2014
Earlier work this paper cites.
Index Policies for A Multi-Class Queue with Convex Holding Cost and Abandonments
Maialen Larrañaga, Urtzi Ayesta, and Ina Maria Verloop · 2014
Earlier work this paper cites.
The restless multi-armed bandit formulation of the cognitive compressive sensing problem
Saeed Bagheri and Anna Scaglione · 2015
Earlier work this paper cites.
Contextual dueling bandits
Miroslav Dudík, Katja Hofmann, Robert E Schapire, Aleksandrs Slivkins, and Masrour Zoghi · 2015
Earlier work this paper cites.
Regret lower bound and optimal algorithm in dueling bandit problem
Junpei Komiyama, Junya Honda, Hisashi Kashima, and Hiroshi Nakagawa · 2015
Earlier work this paper cites.
Optimal recommendation to users that react: Online learning for a class of pomdps
Rahul Meshram, Aditya Gopalan, and D Manjunath · 2016
Earlier work this paper cites.
Asymptotically Optimal Priority Policies for Indexable and Nonindexable Restless Bandits
Ina Maria Verloop · 2016
Earlier work this paper cites.
Double thompson sampling for dueling bandits
Huasen Wu and Xin Liu · 2016
Cited alongside, same era.
An index policy for dynamic pricing in cloud computing under price commitments
Vivek S Borkar, K Ravikumar, and Krishnakant Saboo · 2017
Cited alongside, same era.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Cited alongside, same era.
A survey of preference-based reinforcement learning methods
Christian Wirth, Riad Akrour, Gerhard Neumann, and Johannes Fürnkranz · 2017
Cited alongside, same era.
Reward learning from human preferences and demonstrations in atari
Borja Ibarz, Jan Leike, Tobias Pohlen, Geoffrey Irving, Shane Legg, and Dario Amodei · 2018
Cited alongside, same era.
Deadline Scheduling as Restless Bandits
Zhe Yu, Yunjian Xu, and Lang Tong · 2018
Cited alongside, same era.
Towards soft fairness in restless multi-armed bandits
Dexun Li and Pradeep Varakantham · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Later among the works it cites.
Versatile dueling bandits: Best-of-both world analyses for learning from relative preferences
Aadirupa Saha and Pierre Gaillard · 2022
Later among the works it cites.
Efficient and optimal algorithms for contextual dueling bandits under realizability
Aadirupa Saha and Akshay Krishnamurthy · 2022
Later among the works it cites.
Open problems and fundamental limitations of reinforcement learning from human feedback
Stephen Casper, Xander Davies, Claudia Shi, Thomas Krendl Gilbert, Jérémy Scheurer, Javier Rando, Rachel Freedman, Tomasz Korbak, David Lindner, Pedro Freire, et al · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Towards q-learning the whittle index for restless bandits
Jing Fu, Yoni Nazarathy, Sarat Moka, and Peter G Taylor · 2019
Cited alongside, same era.
Regret bounds for thompson sampling in episodic restless bandit problems
Young Hun Jung and Ambuj Tewari · 2019
Cited alongside, same era.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I Jordan · 2020
Cited alongside, same era.
Collapsing bandits and their application to public health intervention
Aditya Mate, Jackson Killian, Haifeng Xu, Andrew Perrault, and Milind Tambe · 2020
Cited alongside, same era.
Restless-ucb, an efficient and low-complexity algorithm for online restless bandits
Siwei Wang, Longbo Huang, and John Lui · 2020
Cited alongside, same era.
Learn to intervene: An adaptive learning policy for restless bandits in application to preventive healthcare
Arpita Biswas, Gaurav Aggarwal, Pradeep Varakantham, and Milind Tambe · 2021
Cited alongside, same era.
Later among the works it cites.
Planning to fairly allocate: Probabilistic fairness in the restless bandit setting
Christine Herlihy, Aviva Prins, Aravind Srinivasan, and John P Dickerson · 2023
Later among the works it cites.
Dueling rl: Reinforcement learning with trajectory preferences
Aadirupa Saha, Aldo Pacchiano, and Jonathan Lee · 2023
Later among the works it cites.
Optimistic whittle index policy: Online learning for restless bandits
Kai Wang, Lily Xu, Aparna Taneja, and Milind Tambe · 2023
Later among the works it cites.
Finite-time analysis of whittle index based q-learning for restless multi-armed bandits with neural network function approximation
Guojun Xiong and Jian Li · 2023
Later among the works it cites.
Slic-hf: Sequence likelihood calibration with human feedback
Yao Zhao, Rishabh Joshi, Tianqi Liu, Misha Khalman, Mohammad Saleh, and Peter J Liu · 2023
Later among the works it cites.
Principled reinforcement learning with human feedback from pairwise or k-wise comparisons
Banghua Zhu, Michael Jordan, and Jiantao Jiao · 2023
Later among the works it cites.
A general theoretical paradigm to understand learning from human preferences
Mohammad Gheshlaghi Azar, Zhaohan Daniel Guo, Bilal Piot, Remi Munos, Mark Rowland, Michal Valko, and Daniele Calandriello · 2024
Closest in time.
Gino-q: Learning an asymptotically optimal index policy for restless multi-armed bandits
Gongpu Chen, Soung Chang Liew, and Deniz Gunduz · 2024
Closest in time.
Exploration-driven policy optimization in rlhf: Theoretical insights on efficient data utilization
Yihan Du, Anna Winnicki, Gal Dalal, Shie Mannor, and R Srikant · 2024
Closest in time.
Kto: Model alignment as prospect theoretic optimization
Kawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky, and Douwe Kiela · 2024
Closest in time.
Linear program-based policies for restless bandits: Necessary and sufficient conditions for (exponentially fast) asymptotic optimality
Nicolas Gast, Bruno Gaujal, and Chen Yan · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2024
Closest in time.
Generalized preference optimization: A unified approach to offline alignment
Yunhao Tang, Zhaohan Daniel Guo, Zeyu Zheng, Daniele Calandriello, Rémi Munos, Mark Rowland, Pierre Harvey Richemond, Michal Valko, Bernardo Ávila Pires, and Bilal Piot · 2024
Closest in time.
Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint
Wei Xiong, Hanze Dong, Chenlu Ye, Ziqi Wang, Han Zhong, Heng Ji, Nan Jiang, and Tong Zhang · 2024
Closest in time.