Fetching the paper…
Reading the bibliography…
Getting large language models (LLMs) to perform well on the downstream tasks requires pre-training over trillions of tokens.
T. Brown, B. Mann et al. , “Language models are few-shot learners,” in Advances in Neural Information Processing Systems , H. Larochelle, M. Ranzato et al. , Eds., vol. 33. Curran Associates, Inc., 2020, pp. 1877–1901
1901
Earlier work this paper cites.
M. Abadi, A. Agarwal et al. , “TensorFlow: Large-scale machine learning on heterogeneous systems,” 2015, software available from tensorflow.org
2015
Earlier work this paper cites.
S. Gupta, A. Agrawal et al. , “Deep learning with limited numerical precision,” in Proceedings of the 32nd International Conference on International Conference on Machine Learning - Volume 37 , ser. ICML’15. JMLR.org, 2015, p. 1737–1746
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer et al. , “Attention is all you need,” in Advances in Neural Information Processing Systems , 2017, pp. 5998–6008
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
D. A. Hudson and C. D. Manning, “Gqa: A new dataset for real-world visual reasoning and compositional question answering,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 6700–6709
2019
Earlier work this paper cites.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in International Conference on Learning Representations , 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
T. Kwiatkowski, J. Palomaki et al. , “Natural questions: a benchmark for question answering research,” Transactions of the Association of Computational Linguistics , 2019
2019
Earlier work this paper cites.
C. Clark, K. Lee et al. , “Boolq: Exploring the surprising difficulty of natural yes/no questions,” in NAACL , 2019
2019
Earlier work this paper cites.
J. Devlin, M.-W. Chang et al. , “Bert: Pre-training of deep bidirectional transformers for language understanding,” in North American Chapter of the Association for Computational Linguistics , 2019
2019
Earlier work this paper cites.
A. Radford, J. Wu et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
J. Rasley, S. Rajbhandari et al. , “Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters,” in Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining , 2020, pp. 3505–3506
2020
Earlier work this paper cites.
S. Rajbhandari, J. Rasley et al. , “Zero: Memory optimizations toward training trillion parameter models,” in SC20: International Conference for High Performance Computing, Networking, Storage and Analysis . IEEE, 2020, pp. 1–16
2020
Earlier work this paper cites.
A. Arrow, “Apache arrow, a crosslanguage development platform for in-memory analytics.” https://arrow.apache.org/ , 2020
2020
Earlier work this paper cites.
N. Shazeer, “Glu variants improve transformer,” arXiv preprint arXiv:2002.05202 , 2020
2020
Cited alongside, same era.
Y. Bisk, R. Zellers et al. , “Piqa: Reasoning about physical commonsense in natural language,” in Proceedings of the AAAI conference on artificial intelligence , vol. 34, no. 05, 2020, pp. 7432–7439
2020
Cited alongside, same era.
N. Nangia, C. Vania et al. , “CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing , Nov. 2020
2020
Cited alongside, same era.
M. Lewis, Y. Liu et al. , “BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , D. Jurafsky, J. Chai et al. , Eds. Association for Computational Linguistics, Jul. 2020, pp. 7871–7880
2023
Later among the works it cites.
V. A. Korthikanti, J. Casper et al. , “Reducing activation recomputation in large transformer models,” Proceedings of Machine Learning and Systems , vol. 5, pp. 341–353, 2023
2023
Later among the works it cites.
T. Computer, “Redpajama: an open dataset for training large language models,” 2023. [Online]. Available: https://github.com/togethercomputer/RedPajama-Data
2023
Later among the works it cites.
L. Soldaini and K. Lo, “peS2o (Pretraining Efficiently on S2ORC) Dataset,” Allen Institute for AI, Tech. Rep., 2023
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
FairScale authors, “Fairscale: A general purpose modular pytorch library for high performance and large scale training,” https://github.com/facebookresearch/fairscale , 2021
2021
Cited alongside, same era.
D. Hendrycks, C. Burns et al. , “Aligning ai with shared human values,” Proceedings of the International Conference on Learning Representations (ICLR) , 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
L. Gao, J. Tow et al. , “A framework for few-shot language model evaluation,” Version v0. 0.1. Sept , 2021
2021
Cited alongside, same era.
D. Hendrycks, C. Burns et al. , “Measuring massive multitask language understanding,” Proceedings of the International Conference on Learning Representations (ICLR) , 2021
2021
Cited alongside, same era.
K. Sakaguchi, R. L. Bras et al. , “Winogrande: An adversarial winograd schema challenge at scale,” Communications of the ACM , vol. 64, no. 9, pp. 99–106, 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2023
Later among the works it cites.
X. Geng and H. Liu, “Openllama: An open reproduction of llama,” May 2023. [Online]. Available: https://github.com/openlm-research/open_llama
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
R. Li, L. B. Allal et al. , “Starcoder: may the source be with you!” 2023
2023
Later among the works it cites.
Y. Chang, X. Wang et al. , “A survey on evaluation of large language models,” ACM Transactions on Intelligent Systems and Technology , 2023
2023
Later among the works it cites.
S. Gunasekar, Y. Zhang et al. , “Textbooks are all you need,” arXiv preprint arXiv:2306.11644 , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
R. Anil, A. M. Dai et al. , “Palm 2 technical report,” arXiv preprint arXiv:2305.10403 , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
OpenAI, “Gpt-4 technical report,” ArXiv , vol. abs/2303.08774, 2023
2023
Later among the works it cites.
A. Dubey, A. Jauhri et al. , “The llama 3 herd of models,” 2024
2024
Closest in time.
A. Borzunov, M. Ryabinin et al. , “Distributed inference and fine-tuning of large language models over the internet,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
X. Fu, Z. Zhang et al. , “Distributed training of large language models on aws trainium,” in Proceedings of the 2024 ACM Symposium on Cloud Computing , 2024
2024
Closest in time.