Fetching the paper…
Reading the bibliography…
A primary challenge in large language model (LLM) development is their onerous pre-training cost.
Fast Transformer Decoding: One Write-Head is All You Need
Noam Shazeer · 1911
Earlier work this paper cites.
Stochastic Processes
S. M. Ross · 1983
Earlier work this paper cites.
Statistical Analysis of Some Multi-Category Large Margin Classification Methods
Tong Zhang · 2004
Earlier work this paper cites.
Convexity, classification, and risk bounds
Peter L Bartlett, Michael I Jordan, and Jon D McAuliffe · 2006
Earlier work this paper cites.
Model Compression
Cristian Bucilǎ, Rich Caruana, and Alexandru Niculescu-Mizil · 2006
Earlier work this paper cites.
The PASCAL Recognising Textual Entailment Challenge
Ido Dagan, Oren Glickman, and Bernardo Magnini · 2006
Earlier work this paper cites.
How to Compare Different Loss Functions and Their Risks
Ingo Steinwart · 2007
Earlier work this paper cites.
Empirical Bernstein Bounds and Sample Variance Penalization
Andreas Maurer and Massimiliano Pontil · 2009
Earlier work this paper cites.
Hoeffding’s inequality for supermartingales
Xiequan Fan, Ion Grama, and Quansheng Liu · 2012
Earlier work this paper cites.
SemEval-2012 task 7: Choice of Plausible Alternatives: An Evaluation of Commonsense Causal Reasoning
Andrew Gordon, Zornitsa Kozareva, and Melissa Roemmele · 2012
Earlier work this paper cites.
The Winograd Schema Challenge
Hector J. Levesque, Ernest Davis, and Leora Morgenstern · 2012
Earlier work this paper cites.
Semantic parsing on freebase from question-answer pairs
Jonathan Berant, Andrew Chou, Roy Frostig, and Percy Liang · 2013
Earlier work this paper cites.
Distilling the knowledge in a neural network, 2015
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Sequence-Level Knowledge Distillation
Yoon Kim and Alexander M Rush · 2016
Earlier work this paper cites.
A Corpus and Cloze Evaluation for Deeper Understanding of Commonsense Stories
Nasrin Mostafazadeh, Nathanael Chambers, Xiaodong He, Devi Parikh, Dhruv Batra, Lucy Vanderwende, Pushmeet Kohli, and James Allen · 2016
Earlier work this paper cites.
Abstractive Text Summarization using Sequence-to-sequence RNNs and Beyond
Ramesh Nallapati, Bowen Zhou, Cicero dos Santos, Çağlar Gulçehre, and Bing Xiang · 2016
Earlier work this paper cites.
The LAMBADA dataset: Word prediction requiring a broad discourse context
Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Ngoc Quan Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández · 2016
Earlier work this paper cites.
Multiclass Classification Calibration Functions
Bernardo Ávila Pires and Csaba Szepesvári · 2016
Earlier work this paper cites.
TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
Mandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer · 2017
Earlier work this paper cites.
RACE: Large-scale ReAding Comprehension Dataset From Examinations
Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy · 2017
Earlier work this paper cites.
SGDR: Stochastic Gradient Descent with Warm Restarts
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Attention Is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang · 2018
Earlier work this paper cites.
Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord · 2018
Earlier work this paper cites.
Born-Again Neural Networks
Tommaso Furlanello, Zachary Lipton, Michael Tschannen, Laurent Itti, and Anima Anandkumar · 2018
Earlier work this paper cites.
Not all samples are created equal: Deep learning with importance sampling
Angelos Katharopoulos and Francois Fleuret · 2018
Earlier work this paper cites.
Looking Beyond the Surface: A Challenge Set for Reading Comprehension over Multiple Sentences
Daniel Khashabi, Snigdha Chaturvedi, Michael Roth, Shyam Upadhyay, and Dan Roth · 2018
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing
Taku Kudo and John Richardson · 2018
Earlier work this paper cites.
Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal · 2018
Earlier work this paper cites.
Don’t Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization
Shashi Narayan, Shay B. Cohen, and Mirella Lapata · 2018
Earlier work this paper cites.
WiC: 10,000 Example Pairs for Evaluating Context-Sensitive Representations
Mohammad Taher Pilehvar and José Camacho-Collados · 2018
Earlier work this paper cites.
Adafactor: Adaptive Learning Rates with Sublinear Memory Cost
Noam Shazeer and Mitchell Stern · 2018
Earlier work this paper cites.
ReCoRD: Bridging the Gap between Human and Machine Commonsense Reading Comprehension
Sheng Zhang, Xiaodong Liu, Jingjing Liu, Jianfeng Gao, Kevin Duh, and Benjamin Van Durme · 2018
Earlier work this paper cites.
BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova · 2019
Earlier work this paper cites.
The CommitmentBank: Investigating projection in naturally occurring discourse
Marie-Catherine de Marneffe, Mandy Simons, and Judith Tonhauser · 2019
Earlier work this paper cites.
Efficient Training of BERT by Progressively Stacking
Linyuan Gong, Di He, Zhuohan Li, Tao Qin, Liwei Wang, and Tieyan Liu · 2019
Earlier work this paper cites.
TinyBERT: Distilling BERT for Natural Language Understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu · 2019
Earlier work this paper cites.
Sgd on neural networks learns functions of increasing complexity
Dimitris Kalimeris, Gal Kaplun, Preetum Nakkiran, Benjamin Edelman, Tristan Yang, Boaz Barak, and Haofeng Zhang · 2019
Earlier work this paper cites.
Latent Retrieval for Weakly Supervised Open Domain Question Answering
Kenton Lee, Ming-Wei Chang, and Kristina Toutanova · 2019
Cited alongside, same era.
Towards Understanding Knowledge Distillation
Mary Phuong and Christoph Lampert · 2019
Cited alongside, same era.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf · 2019
Cited alongside, same era.
Patient Knowledge Distillation for BERT Model Compression
Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu · 2019
Cited alongside, same era.
Well-Read Students Learn Better: On the Importance of Pre-training Compact Models
Iulia Turc, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al · 2023
Later among the works it cites.
Weak-to-strong generalization: Eliciting strong capabilities with weak supervision
Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker, Leo Gao, Leopold Aschenbrenner, Yining Chen, Adrien Ecoffet, Manas Joglekar, Jan Leike, et al · 2023
Later among the works it cites.
FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance
Lingjiao Chen, Matei Zaharia, and James Zou · 2023
Later among the works it cites.
RedPajama: an Open Dataset for Training Large Language Models
Together Computer · 2023
Later among the works it cites.
Specializing Smaller Language Models towards Multi-Step Reasoning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
SuperGLUE: A Stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman · 2019
Cited alongside, same era.
HellaSwag: Can a Machine Really Finish Your Sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 2019
Cited alongside, same era.
PIQA: Reasoning about Physical Commonsense in Natural Language
Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi · 2020
Cited alongside, same era.
TyDi QA: A Benchmark for Information-Seeking Question Answering in Typologically Diverse Languages
Jonathan H. Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev, and Jennimaria Palomaki · 2020
Cited alongside, same era.
The Pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy · 2020
Cited alongside, same era.
On the Transformer Growth for Progressive BERT Training
Xiaotao Gu, Liyuan Liu, Hongkun Yu, Jing Li, Chen Chen, and Jiawei Han · 2020
Cited alongside, same era.
WikiLingua: A New Benchmark Dataset for Cross-Lingual Abstractive Summarization
Faisal Ladhak, Esin Durmus, Claire Cardie, and Kathleen Mckeown · 2020
Cited alongside, same era.
Yao Fu, Hao Peng, Litu Ou, Ashish Sabharwal, and Tushar Khot · 2023
Later among the works it cites.
Gemini: A Family of Highly Capable Multimodal Models
Gemini-Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al · 2023
Later among the works it cites.
Supervision complexity and its role in knowledge distillation
Hrayr Harutyunyan, Ankit Singh Rawat, Aditya Krishna Menon, Seungyeon Kim, and Sanjiv Kumar · 2023
Later among the works it cites.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al · 2023
Later among the works it cites.
On student-teacher deviations in distillation: does it pay to disobey?
Vaishnavh Nagarajan, Aditya K Menon, Srinadh Bhojanapalli, Hossein Mobahi, and Sanjiv Kumar · 2023
Later among the works it cites.
OpenAI · 2023
Later among the works it cites.
Baolin Peng, Chunyuan Li, Pengcheng He, Michel Galley, and Jianfeng Gao · 2023
Later among the works it cites.
Efficient Training of Language Models using Few-Shot Learning
Sashank J. Reddi, Sobhan Miryoosefi, Stefani Karp, Shankar Krishnan, Satyen Kale, Seungyeon Kim, and Sanjiv Kumar · 2023
Later among the works it cites.
Neural networks trained with SGD learn distributions of increasing complexity
Maria Refinetti, Alessandro Ingrosso, and Sebastian Goldt · 2023
Later among the works it cites.
Andre Niyongabo Rubungo, Craig Arnold, Barry P Rand, and Adji Bousso Dieng · 2023
Later among the works it cites.
Knowledge Distillation Performs Partial Variance Reduction
Mher Safaryan, Alexandra Peste, and Dan Alistarh · 2023
Later among the works it cites.
Stanford Alpaca: An Instruction-following LLaMA model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Later among the works it cites.
UL2: Unifying Language Learning Paradigms
Yi Tay, Mostafa Dehghani, Vinh Q. Tran, Xavier Garcia, Jason Wei, Xuezhi Wang, Hyung Won Chung, Siamak Shakeri, Dara Bahri, Tal Schuster, Huaixiu Steven Zheng, Denny Zhou, Neil Houlsby, and Donald Metzler · 2023
Later among the works it cites.
LLaMA: Open and Efficient Foundation Language Models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Later among the works it cites.
Mimetic Initialization of Self-Attention Layers
Asher Trockman and J Zico Kolter · 2023
Later among the works it cites.
LEMON: Lossless model expansion
Yite Wang, Jiahao Su, Hanlin Lu, Cong Xie, Tianyi Liu, Jianbo Yuan, Haibin Lin, Ruoyu Sun, and Hongxia Yang · 2023
Later among the works it cites.
f-Divergence Minimization for Sequence-Level Knowledge Distillation
Yuqiao Wen, Zichao Li, Wenyu Du, and Lili Mou · 2023
Later among the works it cites.
On-policy distillation of language models: Learning from self-generated mistakes
Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos Garea, Matthieu Geist, and Olivier Bachem · 2024
Closest in time.
Llama 3 Model Card
AI@Meta · 2024
Closest in time.
A Survey on Data Selection for Language Models
Alon Albalak, Yanai Elazar, Sang Michael Xie, Shayne Longpre, Nathan Lambert, Xinyi Wang, Niklas Muennighoff, Bairu Hou, Liangming Pan, Haewon Jeong, Colin Raffel, Shiyu Chang, Tatsunori Hashimoto, and William Yang Wang · 2024
Closest in time.
Perplexed by perplexity: Perplexity-based data pruning with small reference models, 2024
Zachary Ankner, Cody Blakeney, Kartik Sreenivasan, Max Marion, Matthew L. Leavitt, and Mansheej Paul · 2024
Closest in time.
The Claude 3 Model Family: Opus, Sonnet, Haiku
AI Anthropic · 2024
Closest in time.
Language models scale reliably with over-training and on downstream tasks, 2024
Samir Yitzhak Gadre, Georgios Smyrnis, Vaishaal Shankar, Suchin Gururangan, Mitchell Wortsman, Rulin Shao, Jean Mercat, Alex Fang, Jeffrey Li, Sedrick Keh, Rui Xin, Marianna Nezhurina, Igor Vasiljevic, Jenia Jitsev, Luca Soldaini, Alexandros G. Dimakis, Gabriel Ilharco, Pang Wei Koh, Shuran Song, Thomas Kollar, Yair Carmon, Achal Dave, Reinhard Heckel, Niklas Muennighoff, and Ludwig Schmidt · 2024
Closest in time.
MiniLLM: Knowledge Distillation of Large Language Models
Yuxian Gu, Li Dong, Furu Wei, and Minlie Huang · 2024
Closest in time.
Language Model Cascades: Token-Level Uncertainty And Beyond
Neha Gupta, Harikrishna Narasimhan, Wittawat Jitkrittum, Ankit Singh Rawat, Aditya Krishna Menon, and Sanjiv Kumar · 2024
Closest in time.
Rho-1: Not All Tokens Are What You Need
Zhenghao Lin, Zhibin Gou, Yeyun Gong, Xiao Liu, Yelong Shen, Ruochen Xu, Chen Lin, Yujiu Yang, Jian Jiao, Nan Duan, and Weizhu Chen · 2024
Closest in time.
An Emulator for Fine-tuning Large Language Models using Small Language Models
Eric Mitchell, Rafael Rafailov, Archit Sharma, Chelsea Finn, and Christopher D Manning · 2024
Closest in time.
Scaling Smart: Accelerating Large Language Model Pre-training with Small Model Initialization, 2024
Mohammad Samragh, Iman Mirzadeh, Keivan Alizadeh Vahid, Fartash Faghri, Minsik Cho, Moin Nabi, Devang Naik, and Mehrdad Farajtabar · 2024
Closest in time.
D4: Improving LLM pretraining via document de-duplication and diversification
Kushal Tirumala, Daniel Simig, Armen Aghajanyan, and Ari Morcos · 2024
Closest in time.
A Survey on Knowledge Distillation of Large Language Models
Xiaohan Xu, Ming Li, Chongyang Tao, Tao Shen, Reynold Cheng, Jinyang Li, Can Xu, Dacheng Tao, and Tianyi Zhou · 2024
Closest in time.
Yu Yang, Siddhartha Mishra, Jeffrey N Chiang, and Baharan Mirzasoleiman · 2024
Closest in time.
Large Language Model Cascades with Mixture of Thought Representations for Cost-Efficient Reasoning
Murong Yue, Jie Zhao, Min Zhang, Liang Du, and Ziyu Yao · 2024
Closest in time.
Large language models as markov chains, 2024
Oussama Zekri, Ambroise Odonnat, Abdelhakim Benechehab, Linus Bleistein, Nicolas Boullé, and Ievgen Redko · 2024
Closest in time.
Knowledge distillation based on transformed teacher matching
Kaixiang Zheng and En-Hui Yang · 2024
Closest in time.