Fetching the paper…
Reading the bibliography…
Deploying large language models (LLMs) of several billion parameters can be impractical in most industrial use cases due to constraints such as cost, latency limitations, and hardware accessibility.
Binary coors capable or ‘correcting deletions, insertions, and reversals
VI Lcvenshtcin · 1966
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Model compression
Cristian Buciluundefined, Rich Caruana, and Alexandru Niculescu-Mizil · 2006
Earlier work this paper cites.
Beyond accuracy, f-score and roc: A family of discriminant measures for performance evaluation
Marina Sokolova, Nathalie Japkowicz, and Stan Szpakowicz · 2006
Earlier work this paper cites.
Optimal transport: old and new , volume 338
Cédric Villani et al · 2009
Earlier work this paper cites.
Sinkhorn distances: Lightspeed computation of optimal transport
Marco Cuturi · 2013
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Sequence-level knowledge distillation
Yoon Kim and Alexander M. Rush · 2016
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text, 2016
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Enriching word vectors with subword information, 2017
Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov · 2017
Earlier work this paper cites.
Fast discrete distribution clustering using wasserstein barycenter with sparse support, 2017
Jianbo Ye, Panruo Wu, James Z. Wang, and Jia Li · 2017
Earlier work this paper cites.
Stochastic wasserstein autoencoder for probabilistic sentence generation
Hareesh Bahuleyan, Lili Mou, Hao Zhou, and Olga Vechtomova · 2018
Earlier work this paper cites.
Distilled wasserstein learning for word embedding and topic modeling, 2018
Hongteng Xu, Wenlin Wang, Wei Liu, and Lawrence Carin · 2018
Earlier work this paper cites.
Pubmedqa: A dataset for biomedical research question answering, 2019
Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William W. Cohen, and Xinghua Lu · 2019
Earlier work this paper cites.
A study of bfloat16 for deep learning training, 2019
Dhiraj Kalamkar, Dheevatsa Mudigere, Naveen Mellempudi, Dipankar Das, Kunal Banerjee, Sasikanth Avancha, Dharma Teja Vooturi, Nataraj Jammalamadaka, Jianyu Huang, Hector Yuen, Jiyan Yang, Jongsoo Park, Alexander Heinecke, Evangelos Georganas, Sudarshan Srinivasan, Abhisek Kundu, Misha Smelyanskiy, Bharat Kaul, and Pradeep Dubey · 2019
Earlier work this paper cites.
Computational optimal transport: With applications to data science
Gabriel Peyré, Marco Cuturi, et al · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Tinybert: Distilling bert for natural language understanding, 2020
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu · 2020
Earlier work this paper cites.
Qed: A framework and dataset for explanations in question answering, 2020
Matthew Lamm, Jennimaria Palomaki, Chris Alberti, Daniel Andor, Eunsol Choi, Livio Baldini Soares, and Michael Collins · 2020
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter, 2020
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf · 2020
Earlier work this paper cites.
Mobilebert: a compact task-agnostic bert for resource-limited devices, 2020
Zhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu, Yiming Yang, and Denny Zhou · 2020
Cited alongside, same era.
Bertscore: Evaluating text generation with bert
Tianyi Zhang*, Varsha Kishore*, Felix Wu*, Kilian Q. Weinberger, and Yoav Artzi · 2020
Cited alongside, same era.
Dialogsum: A real-life scenario dialogue summarization dataset, 2021
Yulong Chen, Yang Liu, Liang Chen, and Yue Zhang · 2021
Cited alongside, same era.
Automatic text evaluation through the lens of Wasserstein barycenters
Pierre Colombo, Guillaume Staerman, Chloé Clavel, and Pablo Piantanida · 2021
Cited alongside, same era.
Pot: Python optimal transport
Rémi Flamary, Nicolas Courty, Alexandre Gramfort, Mokhtar Z. Alaya, Aurélie Boisbunon, Stanislas Chambon, Laetitia Chapel, Adrien Corenflos, Kilian Fatras, Nemo Fournier, Léo Gautheron, Nathalie T.H. Gayraud, Hicham Janati, Alain Rakotomamonjy, Ievgen Redko, Antoine Rolet, Antony Schutz, Vivien Seguy, Danica J. Sutherland, Romain Tavenard, Alexander Tong, and Titouan Vayer · 2021
Cited alongside, same era.
The falcon series of open language models
Ebtesam Almazrouei, Hamza Alobeidli, Abdulaziz Alshamsi, Alessandro Cappelli, Ruxandra Cojocaru, Mérouane Debbah, Étienne Goffinet, Daniel Hesslow, Julien Launay, Quentin Malartic, Daniele Mazzotta, Badreddine Noune, Baptiste Pannier, and Guilherme Penedo · 2023
Later among the works it cites.
Pythia: A suite for analyzing large language models across training and scaling
Stella Biderman, Hailey Schoelkopf, Quentin Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar van der Wal · 2023
Later among the works it cites.
Sparks of artificial general intelligence: Early experiments with gpt-4, 2023
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, Harsha Nori, Hamid Palangi, Marco Tulio Ribeiro, and Yi Zhang · 2023
Later among the works it cites.
Cost-effective distillation of large language models
Sayantan Dasgupta, Trevor Cohn, and Timothy Baldwin · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
dis, 2021
Tianda Li, Yassir El Mesbahi, Ivan Kobyzev, Ahmad Rashid, Atif Mahmud, Nithin Anchuri, Habib Hajimolahoseini, Yang Liu, and Mehdi Rezagholizadeh · 2021
Cited alongside, same era.
Retrieval augmentation reduces hallucination in conversation, 2021
Kurt Shuster, Spencer Poff, Moya Chen, Douwe Kiela, and Jason Weston · 2021
Cited alongside, same era.
Minilmv2: Multi-head self-attention relation distillation for compressing pretrained transformers, 2021
Wenhui Wang, Hangbo Bao, Shaohan Huang, Li Dong, and Furu Wei · 2021
Cited alongside, same era.
Knot: Knowledge distillation using optimal transport for solving nlp tasks, 2022
Rishabh Bhardwaj, Tushar Vaidya, and Soujanya Poria · 2022
Cited alongside, same era.
Gpt-neox-20b: An open-source autoregressive language model, 2022
Sid Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, Michael Pieler, USVSN Sai Prashanth, Shivanshu Purohit, Laria Reynolds, Jonathan Tow, Ben Wang, and Samuel Weinbach · 2022
Cited alongside, same era.
Optimal transport for unsupervised hallucination detection in neural machine translation
Nuno M Guerreiro, Pierre Colombo, Pablo Piantanida, and André FT Martins · 2022
Cited alongside, same era.
Generate, Annotate, and Learn: NLP with Synthetic Text
Xuanli He, Islam Nassar, Jamie Kiros, Gholamreza Haffari, and Mohammad Norouzi · 2022
Cited alongside, same era.
Qlora: Efficient finetuning of quantized llms, 2023
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer · 2023
Later among the works it cites.
Training on synthetic data beats real data in multimodal relation extraction, 2023
Zilin Du, Haoxin Li, Xu Guo, and Boyang Li · 2023
Later among the works it cites.
Teacherlm: Teaching to fish rather than giving the fish, language modeling likewise
Nan He, Hanyu Lai, Chenyang Zhao, Zirui Cheng, Junting Pan, Ruoyu Qin, Ruofan Lu, Rui Lu, Yunchen Zhang, Gangming Zhao, Zhaohui Hou, Zhiyuan Huang, Shaoqing Lu, Ding Liang, and Mingjie Zhan · 2023
Later among the works it cites.
Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes, 2023
Cheng-Yu Hsieh, Chun-Liang Li, Chih-Kuan Yeh, Hootan Nakhost, Yasuhisa Fujii, Alexander Ratner, Ranjay Krishna, Chen-Yu Lee, and Tomas Pfister · 2023
Later among the works it cites.
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed · 2023
Later among the works it cites.
Fast inference from transformers via speculative decoding
Yaniv Leviathan, Matan Kalman, and Yossi Matias · 2023
Later among the works it cites.
Llm-pruner: On the structural pruning of large language models, 2023
Xinyin Ma, Gongfan Fang, and Xinchao Wang · 2023
Later among the works it cites.
Crosslingual generalization through multitask finetuning
Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M. Saiful Bari, Sheng Shen, Zheng-Xin Yong, Hailey Schoelkopf, Xiangru Tang, Dragomir Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid Alyafeai, Albert Webson, Edward Raff, and Colin Raffel · 2023
Later among the works it cites.
For distillation, tokens are not all you need
Mrigank Raman, Pranav Mani, Davis Liang, and Zachary Lipton · 2023
Later among the works it cites.
A survey of hallucination in large foundation models
Vipula Rawte, Amit Sheth, and Amitava Das · 2023
Later among the works it cites.
Baby llama: knowledge distillation from an ensemble of teachers trained on a small dataset with no performance penalty, 2023
Inar Timiryasov and Jean-Loup Tastet · 2023
Later among the works it cites.
An empirical comparison of lm-based question and answer generation methods
Asahi Ushio, Fernando Alva-Manchego, and Jose Camacho-Collados · 2023
Later among the works it cites.
Lamini-lm: A diverse herd of distilled models from large-scale instructions, 2023
Minghao Wu, Abdul Waheed, Chiyu Zhang, Muhammad Abdul-Mageed, and Alham Fikri Aji · 2023
Later among the works it cites.
FiD-ICL: A fusion-in-decoder approach for efficient in-context learning
Qinyuan Ye, Iz Beltagy, Matthew Peters, Xiang Ren, and Hannaneh Hajishirzi · 2023
Later among the works it cites.
Synthetic data generation method for data-free knowledge distillation in regression neural networks
Tianxun Zhou and Keng-Hwee Chiam · 2023
Later among the works it cites.
Mixtral of experts, 2024
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Théophile Gervet, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed · 2024
Closest in time.