Fetching the paper…
Reading the bibliography…
The reward model for Reinforcement Learning from Human Feedback (RLHF) has proven effective in fine-tuning Large Language Models (LLMs).
Fine-tuning language models from human preferences
Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. 2019 · 1909
Earlier work this paper cites.
Variational inference for the nested chinese restaurant process
Chong Wang and David Blei. 2009 · 2009
Earlier work this paper cites.
The bayesian case model: A generative approach for case-based reasoning and prototype classification
Been Kim, Cynthia Rudin, and Julie A Shah. 2014 · 2014
Earlier work this paper cites.
Advances in neural information processing systems 28
Corinna Cortes, N Lawarence, D Lee, M Sugiyama, and R Garnett. 2015 · 2015
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017 · 2017
Earlier work this paper cites.
Gaussian prototypical networks for few-shot learning on omniglot
Stanislav Fort. 2017 · 2017
Earlier work this paper cites.
Deep reinforcement learning: An overview
Yuxi Li. 2017 · 2017
Earlier work this paper cites.
A deep reinforced model for abstractive summarization
Romain Paulus, Caiming Xiong, and Richard Socher. 2017 · 2017
Earlier work this paper cites.
Prototypical networks for few-shot learning
Jake Snell, Kevin Swersky, and Richard Zemel. 2017 · 2017
Earlier work this paper cites.
Infinite mixture prototypes for few-shot learning
Kelsey Allen, Evan Shelhamer, Hanul Shin, and Joshua Tenenbaum. 2019 · 2019
Earlier work this paper cites.
Generative adversarial user model for reinforcement learning based recommendation system
Xinshi Chen, Shuang Li, Hui Li, Shaohua Jiang, Yuan Qi, and Le Song. 2019 · 2019
Earlier work this paper cites.
Unified language model pre-training for natural language understanding and generation
Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon. 2019 · 2019
Earlier work this paper cites.
Interpretable and steerable sequence learning via prototypes
Yao Ming, Panpan Xu, Huamin Qu, and Liu Ren. 2019 · 2019
Earlier work this paper cites.
Transferrable prototypical networks for unsupervised domain adaptation
Yingwei Pan, Ting Yao, Yehao Li, Yu Wang, Chong-Wah Ngo, and Tao Mei. 2019 · 2019
Earlier work this paper cites.
Graph prototypical networks for few-shot learning on attributed networks
Kaize Ding, Jianling Wang, Jundong Li, Kai Shu, Chenghao Liu, and Huan Liu. 2020 · 2020
Earlier work this paper cites.
Improved prototypical networks for few-shot learning
Zhong Ji, Xingliang Chai, Yunlong Yu, Yanwei Pang, and Zhongfei Zhang. 2020 · 2020
Cited alongside, same era.
Prototype rectification for few-shot learning
Jinlu Liu, Liang Song, and Yongqiang Qin. 2020 · 2020
Cited alongside, same era.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano. 2020 · 2020
Cited alongside, same era.
The online pivot: Lessons learned from teaching a text and data mining course in lockdown, enhancing online teaching with pair programming and digital badges
Beatrice Alex, Clare Llewellyn, Pawel Orzechowski, and Maria Boutchkova. 2021 · 2021
Cited alongside, same era.
Webgpt: Browser-assisted question-answering with human feedback
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al. 2021 · 2021
Cited alongside, same era.
trlX: A framework for large scale reinforcement learning from human feedback
Alexander Havrilla, Maksym Zhuravinskyi, Duy Phung, Aman Tiwari, Jonathan Tow, Stella Biderman, Quentin Anthony, and Louis Castricato. 2023 · 2023
Later among the works it cites.
Prototypical fine-tuning: Towards robust performance under varying data sizes
Yiqiao Jin, Xiting Wang, Yaru Hao, Yizhou Sun, and Xing Xie. 2023 · 2023
Later among the works it cites.
The history and risks of reinforcement learning and human feedback
Nathan Lambert, Thomas Krendl Gilbert, and Tom Zick. 2023 · 2023
Later among the works it cites.
Rlaif: Scaling reinforcement learning from human feedback with ai feedback
Harrison Lee, Samrat Phatale, Hassan Mansoor, Kellie Lu, Thomas Mesnard, Colton Bishop, Victor Carbune, and Abhinav Rastogi. 2023 · 2023
Later among the works it cites.
Chat gpt & google bard ai: A review
Shashi Kant Singh, Shubham Kumar, and Pawan Singh Mehra. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
Ben Wang and Aran Komatsuzaki. 2021 · 2021
Cited alongside, same era.
Bin Ji, Shasha Li, Shaoduo Gan, Jie Yu, Jun Ma, and Huijun Liu. 2022 · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022 · 2022
Cited alongside, same era.
Understanding adamw through proximal methods and scale-freeness
Zhenxun Zhuang, Mingrui Liu, Ashok Cutkosky, and Francesco Orabona. 2022 · 2022
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023 · 2023
Cited alongside, same era.
Meysam Alizadeh, Maël Kubli, Zeynab Samei, Shirin Dehghani, Juan Diego Bermeo, Maria Korobeynikova, and Fabrizio Gilardi. 2023 · 2023
Cited alongside, same era.
Stackllama: an rl fine-tuned llama model for stack exchange question and answering
Edward Beeching, Younes Belkada, Kashif Rasul, Lewis Tunstall, Leandro von Werra, Nazneen Rajani, and Nathan Lambert. 2023 · 2023
Cited alongside, same era.
Preference ranking optimization for human alignment
Feifan Song, Bowen Yu, Minghao Li, Haiyang Yu, Fei Huang, Yongbin Li, and Houfeng Wang. 2023 · 2023
Later among the works it cites.
Simeng Sun, Dhawal Gupta, and Mohit Iyyer. 2023 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
Later among the works it cites.
Densecl: A simple framework for self-supervised dense visual pre-training
Xinlong Wang, Rufeng Zhang, Chunhua Shen, and Tao Kong. 2023 · 2023
Later among the works it cites.
Rrhf: Rank responses to align language models with human feedback without tears
Zheng Yuan, Hongyi Yuan, Chuanqi Tan, Wei Wang, Songfang Huang, and Fei Huang. 2023 · 2023
Later among the works it cites.
Secrets of rlhf in large language models part i: Ppo
Rui Zheng, Shihan Dou, Songyang Gao, Yuan Hua, Wei Shen, Binghai Wang, Yan Liu, Senjie Jin, Qin Liu, Yuhao Zhou, et al. 2023 · 2023
Later among the works it cites.
Chatbot interaction for textual analysis and assistance
OpenAI. 2024 · 2024
Closest in time.
Secrets of rlhf in large language models part ii: Reward modeling
Binghai Wang, Rui Zheng, Lu Chen, Yan Liu, Shihan Dou, Caishuang Huang, Wei Shen, Senjie Jin, Enyu Zhou, Chenyu Shi, et al. 2024 · 2024
Closest in time.
Foundation models meet visualizations: Challenges and opportunities
Weikai Yang, Mengchen Liu, Zheng Wang, and Shixia Liu. 2024 · 2024
Closest in time.
Competeai: Understanding the competition behaviors in large language model-based agents
Qinlin Zhao, Jindong Wang, Yixuan Zhang, Yiqiao Jin, Kaijie Zhu, Hao Chen, and Xing Xie. 2024 · 2024
Closest in time.