2023

DeepSpeed-Chat: Easy, Fast and Affordable RLHF Training of ChatGPT-like Models at All Scales

Yao, Zhewei, Aminabadi, Reza Yazdani, Ruwase, Olatunji et al.

Understand

ChatGPT-like models have revolutionized various applications in artificial intelligence, from summarization and coding to translation, matching or even surpassing human performance.

  • However, the current landscape lacks an accessible, efficient, and cost-effective end-to-end RLHF (Reinforcement Learning with Human Feedback) training pipeline for these powerful models, particularly when training at the scale of billions of parameters.
  • This paper introduces DeepSpeed-Chat, a novel system that democratizes RLHF training, making it accessible to the AI community.
  • DeepSpeed-Chat offers three key capabilities: an easy-to-use training and inference experience for ChatGPT-like models, a DeepSpeed-RLHF pipeline that replicates the training pipeline from InstructGPT, and a robust DeepSpeed-RLHF system that combines various optimizations for training and inference in a unified way.

Built on

  • Proximal policy optimization algorithms

    Original

    John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017

    Earlier work this paper cites.

  • Know what you don’t know: Unanswerable questions for squad

    Original

    Pranav Rajpurkar, Robin Jia, and Percy Liang · 2018

    Earlier work this paper cites.

  • Huggingface’s transformers: State-of-the-art natural language processing

    Original

    Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al · 2019

    Earlier work this paper cites.

  • Pytorch: An imperative style, high-performance deep learning library

    Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019

    Earlier work this paper cites.

  • Learning to summarize with human feedback

    Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano · 2020

    Earlier work this paper cites.

Similar

  • Zero: Memory optimizations toward training trillion parameter models

    Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He · 2020

    Cited alongside, same era.

  • Lora: Low-rank adaptation of large language models

    Original

    Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021

    Cited alongside, same era.

  • Training language models to follow instructions with human feedback

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022

    Cited alongside, same era.

  • Opt: Open pre-trained transformer language models

    Original

    Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al · 2022

    Cited alongside, same era.

  • Colossal ai

    Colossal AI Authors · 2022

    Cited alongside, same era.

Then

  • Chatllama

    ChatLLaMa Authors · 2023

    Closest in time.

  • Stanford alpaca: An instruction-following llama model

    Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023

    Closest in time.

  • Judging llm-as-a-judge with mt-bench and chatbot arena, 2023

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric. P Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica · 2023

    Closest in time.

  • Databricks-dolly

    Databricks · 2023

    Closest in time.

Beyond the bibliography

alphaXiv searches the wider corpus for related work and actual follow-ups.

Open on alphaXiv

alphaXiv is searching for related work…