Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are considered important approaches towards foundational machine intelligence, achieving remarkable success in Natural Language Processing and multimodal tasks, among others.
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 1901
Earlier work this paper cites.
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Shoeybi, M.; Patwary, M.; Puri, R.; LeGresley, P.; Casper, J.; and Catanzaro, B. 2019 · 1909
Earlier work this paper cites.
ZeRO: Memory Optimization Towards Training A Trillion Parameter Models
Rajbhandari, S.; Rasley, J.; Ruwase, O.; and He, Y. 2019 · 1910
Earlier work this paper cites.
A Bridging Model for Parallel Computation
Valiant, L. G. 1990 · 1990
Earlier work this paper cites.
Contemporary practice of psychological assessment by clinical psychologists
Watkins, C. E.; Campbell, V. L.; Nieberding, R.; and Hallmark, R. 1995 · 1995
Earlier work this paper cites.
Neurogenesis in the adult human hippocampus
Eriksson, P. S.; Perfilieva, E.; Björk-Eriksson, T.; Alborn, A.-M.; Nordborg, C.; Peterson, D. A.; and Gage, F. H. 1998 · 1998
Earlier work this paper cites.
Scaling Laws for Neural Language Models
Kaplan, J.; McCandlish, S.; Henighan, T.; Brown, T. B.; Chess, B.; Child, R.; Gray, S.; Radford, A.; Wu, J.; and Amodei, D. 2020 · 2001
Earlier work this paper cites.
Scaling Laws for Autoregressive Generative Modeling
Henighan, T.; Kaplan, J.; Katz, M.; Chen, M.; Hesse, C.; Jackson, J.; Jun, H.; Brown, T. B.; Dhariwal, P.; Gray, S.; Hallacy, C.; Mann, B.; Radford, A.; Ramesh, A.; Ryder, N.; Ziegler, D. M.; Schulman, J.; Amodei, D.; and McCandlish, S. 2020 · 2010
Earlier work this paper cites.
Towards ai-complete question answering: A set of prerequisite toy tasks
Weston, J.; Bordes, A.; Chopra, S.; Rush, A. M.; Van Merriënboer, B.; Joulin, A.; and Mikolov, T. 2015 · 2015
Earlier work this paper cites.
Net2Net: Accelerating Learning via Knowledge Transfer
Chen, T.; Goodfellow, I. J.; and Shlens, J. 2016 · 2016
Earlier work this paper cites.
Fixing Weight Decay Regularization in Adam
Loshchilov, I.; and Hutter, F. 2017 · 2017
Earlier work this paper cites.
Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Clark, P.; Cowhey, I.; Etzioni, O.; Khot, T.; Sabharwal, A.; Schoenick, C.; and Tafjord, O. 2018 · 2018
Earlier work this paper cites.
Past review, current progress, and challenges ahead on the cocktail party problem
Qian, Y.; Weng, C.; Chang, X.; Wang, S.; and Yu, D. 2018 · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A.; Narasimhan, K.; Salimans, T.; Sutskever, I.; et al. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.; Lee, K.; and Toutanova, K. 2019 · 2019
Earlier work this paper cites.
Efficient training of bert by progressively stacking
Gong, L.; He, D.; Li, Z.; Qin, T.; Wang, L.; and Liu, T. 2019 · 2019
Earlier work this paper cites.
Language Models are Unsupervised Multitask Learners
Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; and Sutskever, I. 2019 · 2019
Earlier work this paper cites.
SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems
Wang, A.; Pruksachatkun, Y.; Nangia, N.; Singh, A.; Michael, J.; Hill, F.; Levy, O.; and Bowman, S. R. 2019 · 2019
Earlier work this paper cites.
HellaSwag: Can a Machine Really Finish Your Sentence?
Zellers, R.; Holtzman, A.; Bisk, Y.; Farhadi, A.; and Choi, Y. 2019 · 2019
Earlier work this paper cites.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; and Liu, P. J. 2020 · 2020
Earlier work this paper cites.
Green ai
Schwartz, R.; Dodge, J.; Smith, N. A.; and Etzioni, O. 2020 · 2020
Cited alongside, same era.
CLUE: A Chinese Language Understanding Evaluation Benchmark
Xu, L.; Hu, H.; Zhang, X.; Li, L.; Cao, C.; Li, Y.; Xu, Y.; Sun, K.; Yu, D.; Yu, C.; Tian, Y.; Dong, Q.; Liu, W.; Shi, B.; Cui, Y.; Li, J.; Zeng, J.; Wang, R.; Xie, W.; Li, Y.; Patterson, Y.; Tian, Z.; Zhang, Y.; Zhou, H.; Liu, S.; Zhao, Z.; Zhao, Q.; Yue, C.; Zhang, X.; Yang, Z.; Richardson, K.; and Lan, Z. 2020 · 2020
Cited alongside, same era.
On the Transformer Growth for Progressive BERT Training
Gu, X.; Liu, L.; Yu, H.; Li, J.; Chen, C.; and Han, J. 2021 · 2021
Cited alongside, same era.
Measuring Massive Multitask Language Understanding
Hendrycks, D.; Burns, C.; Basart, S.; Zou, A.; Mazeika, M.; Song, D.; and Steinhardt, J. 2021 · 2021
Cited alongside, same era.
Efficient Large-Scale Language Model Training on GPU Clusters
Narayanan, D.; Shoeybi, M.; Casper, J.; LeGresley, P.; Patwary, M.; Korthikanti, V.; Vainbrand, D.; Kashinkunti, P.; Bernauer, J.; Catanzaro, B.; Phanishayee, A.; and Zaharia, M. 2021 · 2021
Cited alongside, same era.
CofeNet: Context and Former-Label Enhanced Net for Complicated Quotation Extraction
Wang, Y.; Li, X.; Sun, A.; Meng, X.; Liao, H.; and Guo, J. 2022a · 2022
Later among the works it cites.
CORT: A New Baseline for Comparative Opinion Classification by Dual Prompts
Wang, Y.; Zhang, H.; Sun, A.; and Meng, X. 2022b · 2022
Later among the works it cites.
Anil, R.; Dai, A. M.; Firat, O.; Johnson, M.; Lepikhin, D.; Passos, A.; Shakeri, S.; Taropa, E.; Bailey, P.; Chen, Z.; Chu, E.; Clark, J. H.; Shafey, L. E.; Huang, Y.; Meier-Hellstern, K.; Mishra, G.; Moreira, E.; Omernick, M.; Robinson, K.; Ruder, S.; Tay, Y.; Xiao, K.; Xu, Y.; Zhang, Y.; Ábrego, G. H.; Ahn, J.; Austin, J.; Barham, P.; Botha, J. A.; Bradbury, J.; Brahma, S.; Brooks, K.; Catasta, M.; Cheng, Y.; Cherry, C.; Choquette-Choo, C. A.; Chowdhery, A.; Crepy, C.; Dave, S.; Dehghani, M.; Dev, S.; Devlin, J.; Díaz, M.; Du, N.; Dyer, E.; Feinberg, V.; Feng, F.; Fienber, V.; Freitag, M.; Garcia, X.; Gehrmann, S.; Gonzalez, L.; and et al. 2023 · 2023
Closest in time.
C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models
Huang, Y.; Bai, Y.; Zhu, Z.; Zhang, J.; Zhang, J.; Su, T.; Liu, J.; Lv, C.; Zhang, Y.; Lei, J.; Fu, Y.; Sun, M.; and He, J. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Carbon emissions and large neural network training
Patterson, D.; Gonzalez, J.; Le, Q.; Liang, C.; Munguia, L.-M.; Rothchild, D.; So, D.; Texier, M.; and Dean, J. 2021 · 2021
Cited alongside, same era.
Scaling Language Models: Methods, Analysis & Insights from Training Gopher
Rae, J. W.; Borgeaud, S.; Cai, T.; Millican, K.; Hoffmann, J.; Song, H. F.; Aslanides, J.; Henderson, S.; Ring, R.; Young, S.; Rutherford, E.; Hennigan, T.; Menick, J.; Cassirer, A.; Powell, R.; van den Driessche, G.; Hendricks, L. A.; Rauh, M.; Huang, P.; Glaese, A.; Welbl, J.; Dathathri, S.; Huang, S.; Uesato, J.; Mellor, J.; Higgins, I.; Creswell, A.; McAleese, N.; Wu, A.; Elsen, E.; Jayakumar, S. M.; Buchatskaya, E.; Budden, D.; Sutherland, E.; Simonyan, K.; Paganini, M.; Sifre, L.; Martens, L.; Li, X. L.; Kuncoro, A.; Nematzadeh, A.; Gribovskaya, E.; Donato, D.; Lazaridou, A.; Mensch, A.; Lespiau, J.; Tsimpoukelli, M.; Grigorev, N.; Fritz, D.; Sottiaux, T.; Pajarskas, M.; Pohlen, T.; Gong, Z.; Toyama, D.; de Masson d’Autume, C.; Li, Y.; Terzi, T.; Mikulik, V.; Babuschkin, I.; Clark, A.; de Las Casas, D.; Guy, A.; Jones, C.; Bradbury, J.; Johnson, M. J.; Hechtman, B. A.; Weidinger, L.; Gabriel, I.; Isaac, W.; Lockhart, E.; Osindero, S.; Rimell, L.; Dyer, C.; Vinyals, O.; Ayoub, K.; Stanway, J.; Bennett, L.; Hassabis, D.; Kavukcuoglu, K.; and Irving, G. 2021 · 2021
Cited alongside, same era.
RoFormer: Enhanced Transformer with Rotary Position Embedding
Su, J.; Lu, Y.; Pan, S.; Wen, B.; and Liu, Y. 2021 · 2021
Cited alongside, same era.
Tensor Programs IV: Feature Learning in Infinite-Width Neural Networks
Yang, G.; and Hu, E. J. 2021 · 2021
Cited alongside, same era.
Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer
Yang, G.; Hu, E. J.; Babuschkin, I.; Sidor, S.; Liu, X.; Farhi, D.; Ryder, N.; Pachocki, J.; Chen, W.; and Gao, J. 2021 · 2021
Cited alongside, same era.
bert2BERT: Towards Reusable Pretrained Language Models
Chen, C.; Yin, Y.; Shang, L.; Jiang, X.; Qin, Y.; Wang, F.; Wang, Z.; Chen, X.; Liu, Z.; and Liu, Q. 2022 · 2022
Cited alongside, same era.
Interactive Information Extraction by Semantic Information Graph
Fan, S.; Wang, Y.; Li, J.; Zhang, Z.; Shang, S.; and Han, P. 2022 · 2022
Cited alongside, same era.
Closest in time.
FreeLM: Fine-Tuning-Free Language Model
Li, X.; Jiang, X.; Meng, X.; Sun, A.; and Wang, Y. 2023 · 2023
Closest in time.
Adaptive Optimization in the ∞ \infty -Width Limit
Littwin, E.; and Yang, G. 2023 · 2023
Closest in time.
NetGPT: Generative Pretrained Transformer for Network Traffic
Meng, X.; Lin, C.; Wang, Y.; and Zhang, Y. 2023 · 2023
Closest in time.
OpenAI. 2023 · 2023
Closest in time.
Penedo, G.; Malartic, Q.; Hesslow, D.; Cojocaru, R.; Cappelli, A.; Alobeidli, H.; Pannier, B.; Almazrouei, E.; and Launay, J. 2023 · 2023
Closest in time.
Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
Srivastava, A.; Rastogi, A.; Rao, A.; Shoeb, A. A. M.; Abid, A.; Fisch, A.; Brown, A. R.; Santoro, A.; Gupta, A.; Garriga-Alonso, A.; et al. 2023 · 2023
Closest in time.
A Length-Extrapolatable Transformer
Sun, Y.; Dong, L.; Patra, B.; Ma, S.; Huang, S.; Benhaim, A.; Chaudhary, V.; Song, X.; and Wei, F. 2023 · 2023
Closest in time.
Symbol tuning improves in-context learning in language models
Wei, J. W.; Hou, L.; Lampinen, A. K.; Chen, X.; Huang, D.; Tay, Y.; Chen, X.; Lu, Y.; Zhou, D.; Ma, T.; and Le, Q. V. 2023 · 2023
Closest in time.
Yao, Y.; and Wang, Y. 2023 · 2023
Closest in time.
GLM-130B: An Open Bilingual Pre-trained Model
Zeng, A.; Liu, X.; Du, Z.; Wang, Z.; Lai, H.; Ding, M.; Yang, Z.; Xu, Y.; Zheng, W.; Xia, X.; Tam, W. L.; Ma, Z.; Xue, Y.; Zhai, J.; Chen, W.; Liu, Z.; Zhang, P.; Dong, Y.; and Tang, J. 2023 · 2023
Closest in time.
A Survey of Large Language Models
Zhao, W. X.; Zhou, K.; Li, J.; Tang, T.; Wang, X.; Hou, Y.; Min, Y.; Zhang, B.; Zhang, J.; Dong, Z.; Du, Y.; Yang, C.; Chen, Y.; Chen, Z.; Jiang, J.; Ren, R.; Li, Y.; Tang, X.; Liu, Z.; Liu, P.; Nie, J.; and Wen, J. 2023 · 2023
Closest in time.
Deepseek llm: Scaling open-source language models with longtermism
Bi, X.; Chen, D.; Chen, G.; Chen, S.; Dai, D.; Deng, C.; Ding, H.; Dong, K.; Du, Q.; Fu, Z.; et al. 2024 · 2024
Closest in time.
Introducing Meta Llama 3: The most capable openly available LLM to date
Meta. 2024 · 2024
Closest in time.
Mistral 8x22B
Mistral. 2024 · 2024
Closest in time.
Masked Structural Growth for 2x Faster Language Model Pre-training
Yao, Y.; Zhang, Z.; Li, J.; and Wang, Y. 2024 · 2024
Closest in time.