Fetching the paper…
Reading the bibliography…
We study data corruption robustness for reinforcement learning with human feedback (RLHF) in an offline setting.
“Linear programming and sequential decisions”
Alan Manne · 1960
Earlier work this paper cites.
“Online convex optimization in the bandit setting: gradient descent without a gradient”
Abraham Flaxman, Adam Kalai and H McMahan · 2004
Earlier work this paper cites.
“Optimal rates for zero-order convex optimization: The power of two function evaluations”
John Duchi, Michael Jordan, Martin Wainwright and Andre Wibisono · 2015
Earlier work this paper cites.
“Deep Reinforcement Learning from Human Preferences”
Paul. Christiano and Jan Leike · 2017
Earlier work this paper cites.
“Being robust (in high dimensions) can be practical”
Ilias Diakonikolas et al · 2017
Earlier work this paper cites.
“Random gradient-free minimization of convex functions”
Yurii Nesterov and Vladimir Spokoiny · 2017
Earlier work this paper cites.
“A survey of preference-based reinforcement learning methods”
Christian Wirth, Riad Akrour, Gerhard Neumann and Johannes Fürnkranz · 2017
Earlier work this paper cites.
“Lectures on convex optimization”
Yurii Nesterov · 2018
Earlier work this paper cites.
“High-dimensional probability: An introduction with applications in data science”
Roman Vershynin · 2018
Earlier work this paper cites.
“Recent advances in algorithmic high-dimensional robust statistics”
Ilias Diakonikolas and Daniel Kane · 2019
Earlier work this paper cites.
“A short note on concentration inequalities for random vectors with subgaussian norm”
Chi Jin et al · 2019
Earlier work this paper cites.
“Fine-tuning Language Models from Human Preferences”
Daniel Ziegler · 2019
Earlier work this paper cites.
“Provably efficient reinforcement learning with linear function approximation”
Chi Jin, Zhuoran Yang, Zhaoran Wang and Michael Jordan · 2020
Earlier work this paper cites.
“Morel: Model-based offline reinforcement learning”
Rahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli and Thorsten Joachims · 2020
Earlier work this paper cites.
“Offline reinforcement learning: Tutorial, review, and perspectives on open problems”
Sergey Levine, Aviral Kumar, George Tucker and Justin Fu · 2020
Cited alongside, same era.
“Policy Teaching via Environment Poisoning: Training-time Adversarial Attacks against Reinforcement Learning”
Amin Rakhsha et al · 2020
Cited alongside, same era.
“Learning to Summarize with Human Feedback”
Nisan Stiennon and Long Ouyang · 2020
Cited alongside, same era.
“Improved Corruption Robust Algorithms for Episodic Reinforcement Learning”
Yifang Chen, Simon Du and Kevin Jamieson · 2021
Cited alongside, same era.
“On the theory of reinforcement learning with once-per-episode feedback”
Niladri Chatterji, Aldo Pacchiano, Peter Bartlett and Michael Jordan · 2021
Cited alongside, same era.
“Training Language Models to Follow Instructions with Human Feedback”
Long Ouyang and Jeffrey Wu · 2022
Later among the works it cites.
“Corruption-robust offline reinforcement learning”
Xuezhou Zhang, Yiding Chen, Xiaojin Zhu and Wen Sun · 2022
Later among the works it cites.
“Generalized resilience and robust statistics”
Banghua Zhu, Jiantao Jiao and Jacob Steinhardt · 2022
Later among the works it cites.
“Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback”
Stephen Casper et al · 2023
Later among the works it cites.
“Offline Primal-Dual Reinforcement Learning for Linear MDPs”
Germano Gabbianelli, Gergely Neu, Nneka Okolo and Matteo Papini · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ying Jin, Zhuoran Yang and Zhaoran Wang · 2021
Cited alongside, same era.
“B-pref: Benchmarking preference-based reinforcement learning”
Kimin Lee, Laura Smith, Anca Dragan and Pieter Abbeel · 2021
Cited alongside, same era.
“Corruption-robust exploration in episodic reinforcement learning”
Thodoris Lykouris, Max Simchowitz, Alex Slivkins and Wen Sun · 2021
Cited alongside, same era.
“Dueling rl: reinforcement learning with trajectory preferences”
Aldo Pacchiano, Aadirupa Saha and Jonathan Lee · 2021
Cited alongside, same era.
“Robust Policy Gradient against Strong Data Corruption”
Xuezhou Zhang, Yiding Chen, Xiaojin Zhu and Wen Sun · 2021
Cited alongside, same era.
“Trimmed Maximum Likelihood Estimation for Robust Generalized Linear Model”
Pranjal Awasthi, Abhimanyu Das, Weihao Kong and Rajat Sen · 2022
Cited alongside, same era.
“Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback”
Yuntao Bai and Andy Jones · 2022
Cited alongside, same era.
“Benchmarks and Algorithms for Offline Preference-Based Reward Learning”
Daniel Shin, Anca. Dragan and Daniel. Brown · 2023
Later among the works it cites.
“Is RLHF More Difficult than Standard RL? A Theoretical Perspective”
Yuanhao Wang, Qinghua Liu and Chi Jin · 2023
Later among the works it cites.
“Reinforcement Learning from Diverse Human Preferences”
Wanqi Xue, Bo An, Shuicheng Yan and Zhongwen Xu · 2023
Later among the works it cites.
“Corruption-robust algorithms with uncertainty weighting for nonlinear contextual bandits and markov decision processes”
Chenlu Ye, Wei Xiong, Quanquan Gu and Tong Zhang · 2023
Later among the works it cites.
“Corruption-Robust Offline Reinforcement Learning with General Function Approximation”
Chenlu Ye, Rui Yang, Quanquan Gu and Tong Zhang · 2023
Later among the works it cites.
“Provable Offline Reinforcement Learning with Human Feedback”
Wenhao Zhan et al · 2023
Later among the works it cites.
“Principled Reinforcement Learning with Human Feedback from Pairwise or K K -wise Comparisons”
Banghua Zhu, Jiantao Jiao and Michael Jordan · 2023
Later among the works it cites.
“Crowd-PrefRL: Preference-Based Reward Learning from Crowds”
David Chhan, Ellen Novoseller and Vernon Lawhern · 2024
Closest in time.