Fetching the paper…
Reading the bibliography…
Reinforcement Learning algorithms that learn from human feedback (RLHF) need to be efficient in terms of statistical complexity, computational complexity, and query complexity.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E Terry · 1952
Earlier work this paper cites.
Empirical processes: theory and applications
David Pollard · 1990
Earlier work this paper cites.
Empirical Processes in M-estimation , volume 6
Sara A Geer · 2000
Earlier work this paper cites.
Empirical Processes in M-estimation , volume 6
Sara Van de Geer · 2000
Earlier work this paper cites.
Minimizing regret with label efficient prediction
Nicolo Cesa-Bianchi, Gábor Lugosi, and Gilles Stoltz · 2005
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Earlier work this paper cites.
Beat the mean bandit
Yisong Yue and Thorsten Joachims · 2011
Earlier work this paper cites.
Selective sampling and active learning from single and multiple experts
Ofer Dekel, Claudio Gentile, and Karthik Sridharan · 2012
Earlier work this paper cites.
The k-armed dueling bandits problem
Yisong Yue, Josef Broder, Robert Kleinberg, and Thorsten Joachims · 2012
Earlier work this paper cites.
Selective sampling algorithms for cost-sensitive multiclass prediction
Alekh Agarwal · 2013
Earlier work this paper cites.
Learning trajectory preferences for manipulators via iterative improvement
Ashesh Jain, Brian Wojcik, Thorsten Joachims, and Ashutosh Saxena · 2013
Earlier work this paper cites.
(more) efficient reinforcement learning via posterior sampling
Ian Osband, Daniel Russo, and Benjamin Van Roy · 2013
Earlier work this paper cites.
Learning monocular reactive uav control in cluttered natural environments
Stéphane Ross, Narek Melik-Barkhudarov, Kumar Shaurya Shankar, Andreas Wendel, Debadeepta Dey, J Andrew Bagnell, and Martial Hebert · 2013
Earlier work this paper cites.
Eluder dimension and the sample complexity of optimistic exploration
Daniel Russo and Benjamin Van Roy · 2013
Earlier work this paper cites.
Efficient exploration and value function generalization in deterministic systems
Zheng Wen and Benjamin Van Roy · 2013
Earlier work this paper cites.
Learning to optimize via posterior sampling
Daniel Russo and Benjamin Van Roy · 2014
Earlier work this paper cites.
Contextual dueling bandits
Miroslav Dudík, Katja Hofmann, Robert E Schapire, Aleksandrs Slivkins, and Masrour Zoghi · 2015
Earlier work this paper cites.
Thompson sampling for learning parameterized markov decision processes
Aditya Gopalan and Shie Mannor · 2015
Earlier work this paper cites.
Minimax analysis of active learning
Steve Hanneke and Liu Yang · 2015
Earlier work this paper cites.
Learning preferences for manipulation tasks from online coactive feedback
Ashesh Jain, Shikhar Sharma, Thorsten Joachims, and Ashutosh Saxena · 2015
Cited alongside, same era.
Shiv: Reducing supervisor burden in dagger using support vectors for efficient learning from demonstrations in high dimensional state spaces
Michael Laskey, Sam Staszak, Wesley Yu-Shu Hsieh, Jeffrey Mahler, Florian T Pokorny, Anca D Dragan, and Ken Goldberg · 2016
Cited alongside, same era.
Linear thompson sampling revisited
Marc Abeille and Alessandro Lazaric · 2017
Cited alongside, same era.
Optimistic posterior sampling for reinforcement learning: worst-case regret bounds
Shipra Agrawal and Randy Jia · 2017
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Cited alongside, same era.
Deep reinforcement learning from human preferences
Instance-dependent complexity of contextual bandits and reinforcement learning: A disagreement-based perspective
Dylan Foster, Alexander Rakhlin, David Simchi-Levi, and Yunzong Xu · 2021
Later among the works it cites.
Toward a general theory of online selective sampling: Trading off mistakes and queries
Steve Hanneke and Liu Yang · 2021
Later among the works it cites.
Randomized exploration in reinforcement learning with general value function approximation
Haque Ishfaq, Qiwen Cui, Viet Nguyen, Alex Ayoub, Zhuoran Yang, Zhaoran Wang, Doina Precup, and Lin Yang · 2021
Later among the works it cites.
Bellman eluder dimension: New rich classes of rl problems, and sample-efficient algorithms
Chi Jin, Qinghua Liu, and Sobhan Miryoosefi · 2021
Later among the works it cites.
Exponential savings in agnostic active learning through abstention
Nikita Puchkin and Nikita Zhivotovskiy · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Flambe: Structural complexity and representation learning of low rank mdps
Alekh Agarwal, Sham Kakade, Akshay Krishnamurthy, and Wen Sun · 2020
Cited alongside, same era.
Model-based reinforcement learning with value-targeted regression
Alex Ayoub, Zeyu Jia, Csaba Szepesvari, Mengdi Wang, and Lin Yang · 2020
Cited alongside, same era.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I Jordan · 2020
Cited alongside, same era.
Dueling posterior sampling for preference-based reinforcement learning
Ellen Novoseller, Yibing Wei, Yanan Sui, Yisong Yue, and Joel Burdick · 2020
Cited alongside, same era.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano · 2020
Cited alongside, same era.
Pessimistic model-based offline reinforcement learning under partial coverage
Masatoshi Uehara and Wen Sun · 2021
Later among the works it cites.
Model-based rl with optimistic posterior sampling: Structural conditions and sample complexity
Alekh Agarwal and Tong Zhang · 2022
Later among the works it cites.
Stochastic contextual dueling bandits under linear stochastic transitivity models
Viktor Bengs, Aadirupa Saha, and Eyke Hüllermeier · 2022
Later among the works it cites.
Human-in-the-loop: Provably efficient preference-based reinforcement learning with general function approximation
Xiaoyu Chen, Han Zhong, Zhuoran Yang, Zhaoran Wang, and Liwei Wang · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Later among the works it cites.
The role of coverage in online reinforcement learning
Tengyang Xie, Dylan J Foster, Yu Bai, Nan Jiang, and Sham M Kakade · 2022
Later among the works it cites.
Gec: A unified framework for interactive decision making in mdp, pomdp, and beyond
Han Zhong, Wei Xiong, Sirui Zheng, Liwei Wang, Zhaoran Wang, Zhuoran Yang, and Tong Zhang · 2022
Later among the works it cites.
Efficient active learning with abstention
Yinglun Zhu and Robert Nowak · 2022
Later among the works it cites.
Let’s verify step by step
Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe · 2023
Closest in time.
Optimistic mle: A generic model-based algorithm for partially observable sequential decision making
Qinghua Liu, Praneeth Netrapalli, Csaba Szepesvari, and Chi Jin · 2023
Closest in time.
Approximate thompson sampling via epistemic neural networks
Ian Osband, Zheng Wen, Seyed Mohammad Asghari, Vikranth Dwaracherla, Morteza Ibrahimi, Xiuyuan Lu, and Benjamin Van Roy · 2023
Closest in time.
Dueling rl: Reinforcement learning with trajectory preferences
Aadirupa Saha, Aldo Pacchiano, and Jonathan Lee · 2023
Closest in time.
Is rlhf more difficult than standard rl?
Yuanhao Wang, Qinghua Liu, and Chi Jin · 2023
Closest in time.
Principled reinforcement learning with human feedback from pairwise or k-wise comparisons
Banghua Zhu, Michael Jordan, and Jiantao Jiao · 2023
Closest in time.