Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have revolutionized the field of AI, demonstrating unprecedented capacity across various tasks.
“Language models are few-shot learners”
Tom Brown et al · 1901
Earlier work this paper cites.
“Non-autoregressive neural machine translation”
Jiatao Gu et al · 2017
Earlier work this paper cites.
“Automatic differentiation in PyTorch”, 2017
Adam Paszke et al · 2017
Earlier work this paper cites.
“Bert: Pre-training of deep bidirectional transformers for language understanding”
Jacob Devlin, Ming-Wei Chang, Kenton Lee and Kristina Toutanova · 2018
Earlier work this paper cites.
“Deterministic non-autoregressive neural sequence modeling by iterative refinement”
Jason Lee, Elman Mansimov and Kyunghyun Cho · 2018
Earlier work this paper cites.
“Mask-predict: Parallel decoding of conditional masked language models”
Marjan Ghazvininejad, Omer Levy, Yinhan Liu and Luke Zettlemoyer · 2019
Earlier work this paper cites.
“Gpipe: Efficient training of giant neural networks using pipeline parallelism”
Yanping Huang et al · 2019
Earlier work this paper cites.
“Flowseq: Non-autoregressive conditional sequence generation with generative flow”
Xuezhe Ma et al · 2019
Earlier work this paper cites.
“Megatron-lm: Training multi-billion parameter language models using model parallelism”
Mohammad Shoeybi et al · 2019
Earlier work this paper cites.
“Fast structured decoding for sequence models”
Zhiqing Sun et al · 2019
Earlier work this paper cites.
“Iterative refinement in the continuous space for non-autoregressive neural machine translation”
Jason Lee, Raphael Shu and Kyunghyun Cho · 2020
Earlier work this paper cites.
“Latent-variable non-autoregressive neural machine translation with deterministic inference using a delta posterior”
Raphael Shu, Jason Lee, Hideki Nakayama and Kyunghyun Cho · 2020
Earlier work this paper cites.
“TurboTransformers: an efficient GPU serving system for transformer models”
Jiarui Fang, Yang Yu, Chengduo Zhao and Jie Zhou · 2021
Earlier work this paper cites.
“Lora: Low-rank adaptation of large language models”
Edward Hu et al · 2021
Cited alongside, same era.
“Colossal-AI: A unified deep learning system for large-scale parallel training”
Shenggui Li et al · 2021
Cited alongside, same era.
“Deepspeed inference: Enabling efficient inference of transformer models at unprecedented scale”
Reza Aminabadi et al · 2022
Cited alongside, same era.
“Accelerating Transformer Networks through Recomposing Softmax Layers”
Jaewan Choi et al · 2022
Cited alongside, same era.
“Palm: Scaling language modeling with pathways”
Aakanksha Chowdhery et al · 2022
Cited alongside, same era.
“Orca: A Distributed Serving System for { \{ Transformer-Based } \} Generative Models”
Gyeong-In Yu et al · 2022
Later among the works it cites.
“Batch Prompting: Efficient Inference with Large Language Model APIs”
Zhoujun Cheng, Jungo Kasai and Tao Yu · 2023
Closest in time.
“Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality”, 2023
Wei-Lin Chiang et al · 2023
Closest in time.
“Full Stack Optimization of Transformer Inference: a Survey”
Sehoon Kim et al · 2023
Closest in time.
“GPT-4 vs. GPT-3.5: A concise showdown”
Anis Koubaa · 2023
Closest in time.
“New trends in machine translation using large language models: Case examples with chatgpt”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Flashattention: Fast and memory-efficient exact attention with io-awareness”
Tri Dao et al · 2022
Cited alongside, same era.
“LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale”
Tim Dettmers, Mike Lewis, Younes Belkada and Luke Zettlemoyer · 2022
Cited alongside, same era.
“GPTQ: Accurate Post-training Compression for Generative Pretrained Transformers”
Elias Frantar, Saleh Ashkboos, Torsten Hoefler and Dan Alistarh · 2022
Cited alongside, same era.
“Non-autoregressive sequence generation”
Jiatao Gu and Xu Tan · 2022
Cited alongside, same era.
“Training compute-optimal large language models”
Jordan Hoffmann et al · 2022
Cited alongside, same era.
“Training language models to follow instructions with human feedback”
Long Ouyang et al · 2022
Cited alongside, same era.
“Efficiently Scaling Transformer Inference”
Reiner Pope et al · 2022
Cited alongside, same era.
Chenyang Lyu, Jitao Xu and Longyue Wang · 2023
Closest in time.
“High-throughput generative inference of large language models with a single gpu”
Ying Sheng et al · 2023
Closest in time.
“Stanford Alpaca: An Instruction-following LLaMA model”
Rohan Taori et al · 2023
Closest in time.
“Llama: Open and efficient foundation language models”
Hugo Touvron et al · 2023
Closest in time.
Xiaoxia Wu et al · 2023
Closest in time.
“Instruction in the Wild: A User-based Instruction Dataset”
Fuzhao Xue et al · 2023
Closest in time.
“Benchmarking large language models for news summarization”
Tianyi Zhang et al · 2023
Closest in time.