Fetching the paper…
Reading the bibliography…
The prohibitive training costs of Large Language Models (LLMs) have emerged as a significant bottleneck in the development of next-generation LLMs.
Gradient-based learning applied to document recognition
Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson · 2012
Earlier work this paper cites.
Net2net: Accelerating learning via knowledge transfer
Tianqi Chen, Ian Goodfellow, and Jonathon Shlens · 2015
Earlier work this paper cites.
Sequence-level knowledge distillation
Yoon Kim and Alexander M. Rush · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2017
Earlier work this paper cites.
Sequence-Level Training for Non-Autoregressive Neural Machine Translation
Chenze Shao, Yang Feng, Jinchao Zhang, Fandong Meng, and Jie Zhou · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord · 2018
Earlier work this paper cites.
Non-autoregressive neural machine translation
Jiatao Gu, James Bradbury, Caiming Xiong, Victor O.K. Li, and Richard Socher · 2018
Earlier work this paper cites.
Fast decoding in sequence models using discrete latent variables
Lukasz Kaiser, Samy Bengio, Aurko Roy, Ashish Vaswani, Niki Parmar, Jakob Uszkoreit, and Noam Shazeer · 2018
Earlier work this paper cites.
Deterministic non-autoregressive neural sequence modeling by iterative refinement
Jason Lee, Elman Mansimov, and Kyunghyun Cho · 2018
Earlier work this paper cites.
End-to-end non-autoregressive neural machine translation with connectionist temporal classification
Jindřich Libovický and Jindřich Helcl · 2018
Earlier work this paper cites.
Blockwise parallel decoding for deep autoregressive models
Mitchell Stern, Noam Shazeer, and Jakob Uszkoreit · 2018
Earlier work this paper cites.
mixup: Beyond empirical risk minimization
Hongyi Zhang, Moustapha Cisse, Yann N. Dauphin, and David Lopez-Paz · 2018
Earlier work this paper cites.
Mask-predict: Parallel decoding of conditional masked language models
Marjan Ghazvininejad, Omer Levy, Yinhan Liu, and Luke Zettlemoyer · 2019
Earlier work this paper cites.
Efficient training of BERT by progressively stacking
Linyuan Gong, Di He, Zhuohan Li, Tao Qin, Liwei Wang, and Tieyan Liu · 2019
Earlier work this paper cites.
Levenshtein transformer
Jiatao Gu, Changhan Wang, and Junbo Zhao · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Earlier work this paper cites.
FlowSeq: Non-autoregressive conditional sequence generation with generative flow
Xuezhe Ma, Chunting Zhou, Xian Li, Graham Neubig, and Eduard Hovy · 2019
Earlier work this paper cites.
Retrieving sequential information for non-autoregressive neural machine translation
Chenze Shao, Yang Feng, Jinchao Zhang, Fandong Meng, Xilin Chen, and Jie Zhou · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 2019
Earlier work this paper cites.
Root mean square layer normalization
Biao Zhang and Rico Sennrich · 2019
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli · 2020
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Yonatan Bisk, Rowan Zellers, Ronan Le bras, Jianfeng Gao, and Yejin Choi · 2020
Cited alongside, same era.
The pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al · 2020
Cited alongside, same era.
Aligned cross entropy for non-autoregressive machine translation
Marjan Ghazvininejad, Vladimir Karpukhin, Luke Zettlemoyer, and Omer Levy · 2020
Cited alongside, same era.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Cited alongside, same era.
Winogrande: An adversarial winograd schema challenge at scale
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi · 2020
Flexivit: One model for all patch sizes
Lucas Beyer, Pavel Izmailov, Alexander Kolesnikov, Mathilde Caron, Simon Kornblith, Xiaohua Zhai, Matthias Minderer, Michael Tschannen, Ibrahim Alabdulmohsin, and Filip Pavetic · 2023
Later among the works it cites.
Accelerating large language model decoding with speculative sampling
Charlie Chen, Sebastian Borgeaud, Geoffrey Irving, Jean-Baptiste Lespiau, Laurent Sifre, and John Jumper · 2023
Later among the works it cites.
Non-autoregressive machine translation with probabilistic context-free grammar
Shangtong Gui, Chenze Shao, Zhengrui Ma, xishan zhang, Yunji Chen, and Yang Feng · 2023
Later among the works it cites.
Fast inference from transformers via speculative decoding
Yaniv Leviathan, Matan Kalman, and Yossi Matias · 2023
Later among the works it cites.
Fuzzy alignments in directed acyclic graph for non-autoregressive machine translation
Zhengrui Ma, Chenze Shao, Shangtong Gui, Min Zhang, and Yang Feng · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Minimizing the bag-of-ngrams difference for non-autoregressive neural machine translation
Chenze Shao, Jinchao Zhang, Yang Feng, Fandong Meng, and Jie Zhou · 2020
Cited alongside, same era.
Glu variants improve transformer
Noam Shazeer · 2020
Cited alongside, same era.
Latent-variable non-autoregressive neural machine translation with deterministic inference using a delta posterior
Raphael Shu, Jason Lee, Hideki Nakayama, and Kyunghyun Cho · 2020
Cited alongside, same era.
Progressively stacking 2.0: A multi-stage layerwise training method for bert training speedup
Cheng Yang, Shengnan Wang, Chao Yang, Yuechuan Li, Ru He, and Jingqiao Zhang · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Cited alongside, same era.
Order-agnostic cross entropy for non-autoregressive machine translation
Cunxiao Du, Zhaopeng Tu, and Jing Jiang · 2021
Cited alongside, same era.
On the transformer growth for progressive BERT training
Xiaotao Gu, Liyuan Liu, Hongkun Yu, Jing Li, Chen Chen, and Jiawei Han · 2021
Cited alongside, same era.
Scaling data-constrained language models
Niklas Muennighoff, Alexander M Rush, Boaz Barak, Teven Le Scao, Nouamane Tazi, Aleksandra Piktus, Sampo Pyysalo, Thomas Wolf, and Colin Raffel · 2023
Later among the works it cites.
Hierarchical attention encoder decoder
Asier Mujika · 2023
Later among the works it cites.
Reusing pretrained models by multi-linear operators for efficient training
Yu Pan, Ye Yuan, Yichun Yin, Zenglin Xu, Lifeng Shang, Xin Jiang, and Qun Liu · 2023
Later among the works it cites.
Beyond MLE: Convex learning for text generation
Chenze Shao, Zhengrui Ma, Min Zhang, and Yang Feng · 2023
Later among the works it cites.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al · 2023
Later among the works it cites.
Efficient large language models: A survey
Zhongwei Wan, Xin Wang, Che Liu, Samiul Alam, Yu Zheng, et al · 2023
Later among the works it cites.
Learning to grow pretrained models for efficient transformer training
Peihao Wang, Rameswar Panda, Lucas Torroba Hennigen, Philip Greengard, Leonid Karlinsky, Rogerio Feris, David Daniel Cox, Zhangyang Wang, and Yoon Kim · 2023
Later among the works it cites.
Navigating scaling laws: Compute optimality in adaptive model training
Sotiris Anagnostidis, Gregor Bachmann, Imanol Schlag, and Thomas Hofmann · 2024
Closest in time.
Medusa: Simple llm inference acceleration framework with multiple decoding heads
Tianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng, Jason D Lee, Deming Chen, and Tri Dao · 2024
Closest in time.
Flashattention-2: Faster attention with better parallelism and work partitioning
Tri Dao · 2024
Closest in time.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Closest in time.
Break the sequential dependency of LLM inference using lookahead decoding
Yichao Fu, Peter Bailis, Ion Stoica, and Hao Zhang · 2024
Closest in time.
Block transformer: Global-to-local language modeling for fast inference
Namgyu Ho, Sangmin Bae, Taehyeon Kim, Hyunjik Jo, Yireun Kim, Tal Schuster, Adam Fisch, James Thorne, and Se-Young Yun · 2024
Closest in time.
Zeping Li, Xinlong Yang, Ziheng Gao, Ji Liu, Zhuang Liu, Dong Li, Jinzhang Peng, Lu Tian, and Emad Barsoum · 2024
Closest in time.
Bita: Bi-directional tuning for lossless acceleration in large language models
Feng Lin, Hanling Yi, Hongbin Li, Yifan Yang, Xiaotian Yu, Guangming Lu, and Rong Xiao · 2024
Closest in time.
Masked structural growth for 2x faster language model pre-training
Yiqun Yao, Zheng Zhang, Jing Li, and Yequan Wang · 2024
Closest in time.
Megabyte: Predicting million-byte sequences with multiscale transformers
Lili Yu, Dániel Simig, Colin Flaherty, Armen Aghajanyan, Luke Zettlemoyer, and Mike Lewis · 2024
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2024
Closest in time.