Fetching the paper…
Reading the bibliography…
In this paper, we study offline Reinforcement Learning with Human Feedback (RLHF) where we aim to learn the human's underlying reward and the MDP's optimal policy from a set of trajectories induced by human choices.
Optimal replacement of gmc bus engines: An empirical model of harold zurcher
John Rust · 1987
Earlier work this paper cites.
Conditional choice probabilities and the estimation of dynamic models
V Joseph Hotz and Robert A Miller · 1993
Earlier work this paper cites.
A simulation estimator for dynamic models of discrete choice
V Joseph Hotz, Robert A Miller, Seth Sanders, and Jeffrey Smith · 1994
Earlier work this paper cites.
Swapping the nested fixed point algorithm: A class of estimators for discrete markov decision models
Victor Aguirregabiria and Pedro Mira · 2002
Earlier work this paper cites.
Support vector machines
Ingo Steinwart and Andreas Christmann · 2008
Earlier work this paper cites.
Gaussian process optimization in the bandit setting: No regret and experimental design
Niranjan Srinivas, Andreas Krause, Sham M Kakade, and Matthias Seeger · 2009
Earlier work this paper cites.
Dynamic discrete choice structural models: A survey
Victor Aguirregabiria and Pedro Mira · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Earlier work this paper cites.
Practical methods for estimation of dynamic discrete choice models
Peter Arcidiacono and Paul B Ellickson · 2011
Earlier work this paper cites.
Learning trajectory preferences for manipulators via iterative improvement
Ashesh Jain, Brian Wojcik, Thorsten Joachims, and Ashutosh Saxena · 2013
Earlier work this paper cites.
Identification and efficient semiparametric estimation of a dynamic discrete game
Patrick Bajari, Victor Chernozhukov, Han Hong, and Denis Nekipelov · 2015
Earlier work this paper cites.
Estimation from pairwise comparisons: Sharp minimax bounds with topology dependence
Nihar Shah, Sivaraman Balakrishnan, Joseph Bradley, Abhay Parekh, Kannan Ramchandran, and Martin Wainwright · 2015
Earlier work this paper cites.
On kernelized multi-armed bandits, 2017
Sayak Ray Chowdhury and Aditya Gopalan · 2017
Earlier work this paper cites.
Imitation learning: A survey of learning methods
Ahmed Hussein, Mohamed Medhat Gaber, Eyad Elyan, and Chrisina Jayne · 2017
Earlier work this paper cites.
Learning dynamic robot-to-human object handover from human feedback
Andras Kupcsik, David Hsu, and Wee Sun Lee · 2018
Earlier work this paper cites.
A gentle introduction to empirical process theory and applications
Bodhisattva Sen · 2018
Earlier work this paper cites.
Exponentially weighted imitation learning for batched historical data
Qing Wang, Jiechao Xiong, Lei Han, Han Liu, Tong Zhang, et al · 2018
Earlier work this paper cites.
Temporal-difference estimation of dynamic discrete choice models
Karun Adusumilli and Dita Eckardt · 2019
Earlier work this paper cites.
Sample-optimal parametric q-learning using linearly additive features, 2019
Lin F. Yang and Mengdi Wang · 2019
Cited alongside, same era.
Flambe: Structural complexity and representation learning of low rank mdps, 2020
Alekh Agarwal, Sham Kakade, Akshay Krishnamurthy, and Wen Sun · 2020
Cited alongside, same era.
Dynamic assortment optimization with changing contextual information
Xi Chen, Yining Wang, and Yuan Zhou · 2020
Cited alongside, same era.
Minimax-optimal off-policy evaluation with linear function approximation
Yaqi Duan, Zeyu Jia, and Mengdi Wang · 2020
Cited alongside, same era.
Risk-sensitive reinforcement learning: Near-optimal risk-sample tradeoff in regret
Yingjie Fei, Zhuoran Yang, Yudong Chen, Zhaoran Wang, and Qiaomin Xie · 2020
Cited alongside, same era.
Human-centric dialog training via offline reinforcement learning
Representation learning for online and offline rl in low-rank mdps
Masatoshi Uehara, Xuezhou Zhang, and Wen Sun · 2021
Later among the works it cites.
Pessimistic bootstrapping for uncertainty-driven offline reinforcement learning
Chenjia Bai, Lingxiao Wang, Zhuoran Yang, Zhihong Deng, Animesh Garg, Peng Liu, and Zhaoran Wang · 2022
Later among the works it cites.
Human-in-the-loop: Provably efficient preference-based reinforcement learning with general function approximation, 2022
Xiaoyu Chen, Han Zhong, Zhuoran Yang, Zhaoran Wang, and Liwei Wang · 2022
Later among the works it cites.
Locally robust semiparametric estimation
Victor Chernozhukov, Juan Carlos Escanciano, Hidehiko Ichimura, Whitney K Newey, and James M Robins · 2022
Later among the works it cites.
Scaling laws for reward model overoptimization, 2022
Leo Gao, John Schulman, and Jacob Hilton · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Natasha Jaques, Judy Hanwen Shen, Asma Ghandeharioun, Craig Ferguson, Agata Lapedriza, Noah Jones, Shixiang Shane Gu, and Rosalind Picard · 2020
Cited alongside, same era.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I Jordan · 2020
Cited alongside, same era.
Dueling posterior sampling for preference-based reinforcement learning
Ellen Novoseller, Yibing Wei, Yanan Sui, Yisong Yue, and Joel Burdick · 2020
Cited alongside, same era.
Provably efficient reinforcement learning with kernel and neural function approximations
Zhuoran Yang, Chi Jin, Zhaoran Wang, Mengdi Wang, and Michael Jordan · 2020
Cited alongside, same era.
Mopo: Model-based offline policy optimization
Tianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon, James Y Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma · 2020
Cited alongside, same era.
Provably efficient safe exploration via primal-dual policy optimization
Dongsheng Ding, Xiaohan Wei, Zhuoran Yang, Zhaoran Wang, and Mihailo Jovanovic · 2021
Cited alongside, same era.
Is pessimism provably efficient for offline rl?
Ying Jin, Zhuoran Yang, and Zhaoran Wang · 2021
Cited alongside, same era.
Improving alignment of dialogue agents via targeted human judgements
Amelia Glaese, Nat McAleese, Maja Trebacz, John Aslanides, Vlad Firoiu, Timo Ewalds, Maribeth Rauh, Laura Weidinger, Martin Chadwick, Phoebe Thacker, et al · 2022
Later among the works it cites.
Interactively learning preference constraints in linear bandits, 2022
David Lindner, Sebastian Tschiatschek, Katja Hofmann, and Andreas Krause · 2022
Later among the works it cites.
Teaching language models to support answers with verified quotes
Jacob Menick, Maja Trebacz, Vladimir Mikulik, John Aslanides, Francis Song, Martin Chadwick, Mia Glaese, Susannah Young, Lucy Campbell-Gillingham, Geoffrey Irving, et al · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Later among the works it cites.
Efficient and optimal algorithms for contextual dueling bandits under realizability
Aadirupa Saha and Akshay Krishnamurthy · 2022
Later among the works it cites.
Pessimistic q-learning for offline reinforcement learning: Towards optimal sample complexity
Laixi Shi, Gen Li, Yuting Wei, Yuxin Chen, and Yuejie Chi · 2022
Later among the works it cites.
Structural estimation of markov decision processes in high-dimensional state space with finite-time guarantees, 2022
Siliang Zeng, Mingyi Hong, and Alfredo Garcia · 2022
Later among the works it cites.
Pac reinforcement learning for predictive state representations, 2022
Wenhao Zhan, Masatoshi Uehara, Wen Sun, and Jason D. Lee · 2022
Later among the works it cites.
Pessimistic minimax value iteration: Provably efficient equilibrium learning from offline datasets
Han Zhong, Wei Xiong, Jiyuan Tan, Liwei Wang, Tong Zhang, Zhaoran Wang, and Zhuoran Yang · 2022
Later among the works it cites.
On the provable advantage of unsupervised pretraining, 2023
Jiawei Ge, Shange Tang, Jianqing Fan, and Chi Jin · 2023
Closest in time.
Provably feedback-efficient reinforcement learning via active reward learning, 2023
Dingwen Kong and Lin F. Yang · 2023
Closest in time.
Understanding expertise through demonstrations: A maximum likelihood framework for offline inverse reinforcement learning, 2023
Siliang Zeng, Chenliang Li, Alfredo Garcia, and Mingyi Hong · 2023
Closest in time.
Principled reinforcement learning with human feedback from pairwise or k k -wise comparisons, 2023
Banghua Zhu, Jiantao Jiao, and Michael I. Jordan · 2023
Closest in time.