Fetching the paper…
Reading the bibliography…
Reinforcement Learning from Human Feedback (RLHF) is a pivotal technique that aligns language models closely with human-centric values.
Rank analysis of incomplete block designs: I. the method of paired comparisons
R. A. Bradley and M. E. Terry · 1952
Earlier work this paper cites.
The analysis of permutations
R. L. Plackett · 1975
Earlier work this paper cites.
Tamer: Training an agent manually via evaluative reinforcement
W. B. Knox and P. Stone · 2008
Earlier work this paper cites.
Interactively optimizing information retrieval systems as a dueling bandits problem
Y. Yue and T. Joachims · 2009
Earlier work this paper cites.
Beat the mean bandit
Y. Yue and T. Joachims · 2011
Earlier work this paper cites.
Individual choice behavior: A theoretical analysis
R. D. Luce · 2012
Earlier work this paper cites.
The k k -armed dueling bandits problem
Y. Yue, J. Broder, R. Kleinberg, and T. Joachims · 2012
Earlier work this paper cites.
Introduction to mathematical statistics
R. V. Hogg, J. W. McKean, A. T. Craig, et al · 2013
Earlier work this paper cites.
Learning trajectory preferences for manipulators via iterative improvement
A. Jain, B. Wojcik, T. Joachims, and A. Saxena · 2013
Earlier work this paper cites.
Reducing dueling bandits to cardinal bandits
N. Ailon, Z. S. Karnin, and T. Joachims · 2014
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
A. Romero, N. Ballas, S. E. Kahou, A. Chassang, C. Gatta, and Y. Bengio · 2014
Earlier work this paper cites.
A relative exponential weighing algorithm for adversarial utility-based dueling bandits
P. Gajane, T. Urvoy, and F. Clérot · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
G. Hinton, O. Vinyals, and J. Dean · 2015
Earlier work this paper cites.
Regret lower bound and optimal algorithm in dueling bandit problem
J. Komiyama, J. Honda, H. Kashima, and H. Nakagawa · 2015
Earlier work this paper cites.
Estimation from pairwise comparisons: Sharp minimax bounds with topology dependence
N. Shah, S. Balakrishnan, J. Bradley, A. Parekh, K. Ramchandran, and M. Wainwright · 2015
Earlier work this paper cites.
Like what you like: Knowledge distill via neuron selectivity transfer
Z. Huang and N. Wang · 2017
Earlier work this paper cites.
Interactive learning from policy-dependent human feedback
J. MacGlashan, M. K. Ho, R. Loftin, B. Peng, G. Wang, D. L. Roberts, M. E. Taylor, and M. L. Littman · 2017
Earlier work this paper cites.
Active preference-based learning of reward functions
D. Sadigh, A. D. Dragan, S. Sastry, and S. A. Seshia · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
A survey of preference-based reinforcement learning methods
C. Wirth, R. Akrour, G. Neumann, and J. Fürnkranz · 2017
Earlier work this paper cites.
A gift from knowledge distillation: Fast optimization, network minimization and transfer learning
J. Yim, D. Joo, J. Bae, and J. Kim · 2017
Earlier work this paper cites.
Born again neural networks
T. Furlanello, Z. Lipton, M. Tschannen, L. Itti, and A. Anandkumar · 2018
Earlier work this paper cites.
Learning dynamic robot-to-human object handover from human feedback
A. Kupcsik, D. Hsu, and W. S. Lee · 2018
Earlier work this paper cites.
Deep tamer: Interactive agent shaping in high-dimensional state spaces
G. Warnell, N. Waytowich, V. Lawhern, and P. Stone · 2018
Earlier work this paper cites.
Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations
D. Brown, W. Goo, P. Nagarajan, and S. Niekum · 2019
Cited alongside, same era.
On the efficacy of knowledge distillation
J. H. Cho and B. Hariharan · 2019
Cited alongside, same era.
Dueling posterior sampling for preference-based reinforcement learning
E. R. Novoseller, Y. Sui, Y. Yue, and J. W. Burdick · 2019
Cited alongside, same era.
Relational knowledge distillation
W. Park, D. Kim, Y. Lu, and M. Cho · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al · 2019
Cited alongside, same era.
PAC Battling Bandits in the Plackett-Luce Model
A. Saha and A. Gopalan · 2019
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
D. Ganguli, L. Lovitt, J. Kernion, A. Askell, Y. Bai, S. Kadavath, B. Mann, E. Perez, N. Schiefer, K. Ndousse, et al · 2022
Later among the works it cites.
Scaling laws for reward model overoptimization
L. Gao, J. Schulman, and J. Hilton · 2022
Later among the works it cites.
Exploiting correlation to achieve faster learning rates in low-rank preference bandits
S. Ghoshal and A. Saha · 2022
Later among the works it cites.
Improving alignment of dialogue agents via targeted human judgements
A. Glaese, N. McAleese, M. Tr e ⋅ \underset{\cdot}{e} bacz, J. Aslanides, V. Firoiu, T. Ewalds, M. Rauh, L. Weidinger, M. Chadwick, P. Thacker, et al · 2022
Later among the works it cites.
Teaching language models to support answers with verified quotes
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Contrastive representation distillation
Y. Tian, D. Krishnan, and P. Isola · 2019
Cited alongside, same era.
Similarity-preserving knowledge distillation
F. Tung and G. Mori · 2019
Cited alongside, same era.
Fine-tuning language models from human preferences
D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P. Christiano, and G. Irving · 2019
Cited alongside, same era.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Cited alongside, same era.
Explaining knowledge distillation by quantifying the knowledge
X. Cheng, Z. Rao, Y. Chen, and Q. Zhang · 2020
Cited alongside, same era.
Improved optimistic algorithms for logistic bandits
L. Faury, M. Abeille, C. Calauzènes, and O. Fercoq · 2020
Cited alongside, same era.
J. Menick, M. Trebacz, V. Mikulik, J. Aslanides, F. Song, M. Chadwick, M. Glaese, S. Young, L. Campbell-Gillingham, G. Irving, et al · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Later among the works it cites.
Red teaming language models with language models
E. Perez, S. Huang, F. Song, T. Cai, R. Ring, J. Aslanides, A. Glaese, N. McAleese, and G. Irving · 2022
Later among the works it cites.
Better teacher better student: Dynamic prior knowledge for knowledge distillation
Z. Qiu, X. Ma, K. Yang, C. Liu, J. Hou, S. Yi, and W. Ouyang · 2022
Later among the works it cites.
R. Ramamurthy, P. Ammanabrolu, K. Brantley, J. Hessel, R. Sifa, C. Bauckhage, H. Hajishirzi, and Y. Choi · 2022
Later among the works it cites.
Efficient and optimal algorithms for contextual dueling bandits under realizability
A. Saha and A. Krishnamurthy · 2022
Later among the works it cites.
Chatgpt: Optimizing language models for dialogue
J. Schulman, B. Zoph, C. Kim, J. Hilton, J. Menick, J. Weng, J. F. C. Uribe, L. Fedus, L. Metz, M. Pokorny, et al · 2022
Later among the works it cites.
Offline rl for natural language generation with implicit language q learning
C. Snell, I. Kostrikov, Y. Su, M. Yang, and S. Levine · 2022
Later among the works it cites.
Decoupled knowledge distillation
B. Zhao, Q. Cui, R. Song, Y. Qiu, and J. Liang · 2022
Later among the works it cites.
Model card and evaluations for claude models, 2023
Anthropic · 2023
Later among the works it cites.
Sparks of artificial general intelligence: Early experiments with gpt-4
S. Bubeck, V. Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y. T. Lee, Y. Li, S. Lundberg, et al · 2023
Later among the works it cites.
Alpacafarm: A simulation framework for methods that learn from human feedback
Y. Dubois, X. Li, R. Taori, T. Zhang, I. Gulrajani, J. Ba, C. Guestrin, P. Liang, and T. B. Hashimoto · 2023
Later among the works it cites.
Scaling laws for reward model overoptimization
L. Gao, J. Schulman, and J. Hilton · 2023
Later among the works it cites.
Gpt-4 technical report, 2023
OpenAI · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
R. Rafailov, A. Sharma, E. Mitchell, S. Ermon, C. D. Manning, and C. Finn · 2023
Later among the works it cites.
Benchmarks and algorithms for offline preference-based reward learning
D. Shin, A. D. Dragan, and D. S. Brown · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2023
Later among the works it cites.
Pairwise proximal policy optimization: Harnessing relative feedback for llm alignment, 2023
T. Wu, B. Zhu, R. Zhang, Z. Wen, K. Ramchandran, and J. Jiao · 2023
Later among the works it cites.
Rrhf: Rank responses to align language models with human feedback without tears
Z. Yuan, H. Yuan, C. Tan, W. Wang, S. Huang, and F. Huang · 2023
Later among the works it cites.
Towards the fundamental limits of knowledge transfer over finite domains
Q. Zhao and B. Zhu · 2023
Later among the works it cites.