Fetching the paper…
Reading the bibliography…
The quality of training data impacts the performance of pre-trained large language models (LMs).
Abilities and learning sets in knowledge acquisition
Robert M Gagne and Noel E Paradise · 1961
Earlier work this paper cites.
The acquisition of knowledge
Robert M Gagne · 1962
Earlier work this paper cites.
Research into learning hierarchies
Richard T White · 1973
Earlier work this paper cites.
Past and future research on learning hierarchies
Richard T. White and Robert M. Gagné · 1974
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
Arkadij Semenovič Nemirovskij and David Borisovich Yudin · 1983
Earlier work this paper cites.
A sequential algorithm for training text classifiers: Corrigendum and additional data
David D Lewis · 1995
Earlier work this paper cites.
Curriculum learning
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston · 2009
Earlier work this paper cites.
Intelligent selection of language model training data
Robert C. Moore and William Lewis · 2010
Earlier work this paper cites.
Coresets and sketches, 2016
Jeff M. Phillips · 2016
Earlier work this paper cites.
Latent skill embedding for personalized lesson sequence recommendation, 2016
Siddharth Reddy, Igor Labutov, and Thorsten Joachims · 2016
Earlier work this paper cites.
Understanding black-box predictions via influence functions, 2017
Pang Wei Koh and Percy Liang · 2017
Earlier work this paper cites.
Data selection strategies for multi-domain sentiment analysis, 2017
Sebastian Ruder, Parsa Ghaffari, and John G. Breslin · 2017
Earlier work this paper cites.
Learning to select data for transfer learning with bayesian optimization
Sebastian Ruder and Barbara Plank · 2017
Earlier work this paper cites.
Constituency parsing with a self-attentive encoder
Nikita Kitaev and Dan Klein · 2018
Earlier work this paper cites.
An empirical study of example forgetting during deep neural network learning, 2018
Mariya Toneva, Alessandro Sordoni, Remi Tachet des Combes, Adam Trischler, Yoshua Bengio, and Geoffrey J. Gordon · 2018
Earlier work this paper cites.
Taskonomy: Disentangling task transfer learning
Amir R. Zamir, Alexander Sax, William Shen, Leonidas Guibas, Jitendra Malik, and Silvio Savarese · 2018
Earlier work this paper cites.
Selection via proxy: Efficient data selection for deep learning, 2019
Cody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman, Peter Bailis, Percy Liang, Jure Leskovec, and Matei Zaharia · 2019
Earlier work this paper cites.
Data shapley: Equitable valuation of data for machine learning
Amirata Ghorbani and James Zou · 2019
Earlier work this paper cites.
Multilingual constituency parsing with self-attention and pre-training
Nikita Kitaev, Steven Cao, and Dan Klein · 2019
Earlier work this paper cites.
Coresets for data-efficient training of machine learning models, 2019
Baharan Mirzasoleiman, Jeff Bilmes, and Jure Leskovec · 2019
Earlier work this paper cites.
Data parameters: A new family of parameters for learning a differentiable curriculum
Shreyas Saxena, Oncel Tuzel, and Dennis DeCoste · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al · 2020
Earlier work this paper cites.
Advanced algorithms: Notes for cmu 15-850 (fall 2020), 2020
Anupam Gupta · 2020
Earlier work this paper cites.
Don’t stop pretraining: Adapt language models to domains and tasks
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith · 2020
Cited alongside, same era.
When do curricula work?, 2020
Xiaoxia Wu, Ethan Dyer, and Behnam Neyshabur · 2020
Cited alongside, same era.
GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow, March 2021
Sid Black, Gao Leo, Phil Wang, Connor Leahy, and Stella Biderman · 2021
Cited alongside, same era.
On the opportunities and risks of foundation models, 2021
Rishi Bommasani, Percy Liang, et al · 2021
Cited alongside, same era.
Evaluating large language models trained on code, 2021
Mark Chen, Jerry Tworek, et al · 2021
Cited alongside, same era.
A framework for few-shot language model evaluation, September 2021
Leo Gao, Jonathan Tow, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Kyle McDonell, Niklas Muennighoff, Jason Phang, Laria Reynolds, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou · 2021
Self-instruct: Aligning language model with self generated instructions, 2022
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi · 2022
Later among the works it cites.
Super-naturalinstructions: Generalization via declarative instructions on 1600+ nlp tasks, 2022
Yizhong Wang, Swaroop Mishra, et al · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models, 2022
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer · 2022
Later among the works it cites.
Unveiling transformers with lego: a synthetic reasoning task, 2022
Yi Zhang, Arturs Backurs, Sébastien Bubeck, Ronen Eldan, Suriya Gunasekar, and Tal Wagner · 2022
Later among the works it cites.
Semdedup: Data-efficient learning at web-scale through semantic deduplication, 2023
Amro Abbas, Kushal Tirumala, Dániel Simig, Surya Ganguli, and Ari S. Morcos · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Bias out-of-the-box: An empirical analysis of intersectional occupational biases in popular generative language models
Hannah Rose Kirk, Yennie Jun, Filippo Volpin, Haider Iqbal, Elias Benussi, Frederic Dreyer, Aleksandar Shtedritski, and Yuki Asano · 2021
Cited alongside, same era.
Efficient multi-task auxiliary learning: Selecting auxiliary data by feature similarity
Po-Nien Kung, Sheng-Siang Yin, Yi-Cheng Chen, Tse-Hsuan Yang, and Yun-Nung Chen · 2021
Cited alongside, same era.
Towards understanding and mitigating social biases in language models
Paul Pu Liang, Chiyu Wu, Louis-Philippe Morency, and Ruslan Salakhutdinov · 2021
Cited alongside, same era.
Prioritized training on points that are learnable, worth learning, and not yet learned (workshop version), 2021
Sören Mindermann, Muhammed Razzak, Winnie Xu, Andreas Kirsch, Mrinank Sharma, Adrien Morisot, Aidan N. Gomez, Sebastian Farquhar, Jan Brauner, and Yarin Gal · 2021
Cited alongside, same era.
Stereoset: Measuring stereotypical bias in pretrained language models
Moin Nadeem, Anna Bethke, and Siva Reddy · 2021
Cited alongside, same era.
Deep learning on a data diet: Finding important examples early in training, 2021
Mansheej Paul, Surya Ganguli, and Gintare Karolina Dziugaite · 2021
Cited alongside, same era.
Pythia: A suite for analyzing large language models across training and scaling, 2023
Stella Biderman, Hailey Schoelkopf, Quentin Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar van der Wal · 2023
Closest in time.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing · 2023
Closest in time.
Koala: A dialogue model for academic research
Xinyang Geng, Arnav Gudibande, Hao Liu, Eric Wallace, Pieter Abbeel, Sergey Levine, and Dawn Song · 2023
Closest in time.
Palm2 technical report
Google · 2023
Closest in time.
Simfluence: Modeling the influence of individual training examples by simulating training runs, 2023
Kelvin Guu, Albert Webson, Ellie Pavlick, Lucas Dixon, Ian Tenney, and Tolga Bolukbasi · 2023
Closest in time.
Taskweb: Selecting better source tasks for multi-task nlp, 2023
Joongwon Kim, Akari Asai, Gabriel Ilharco, and Hannaneh Hajishirzi · 2023
Closest in time.
The bigscience roots corpus: A 1.6tb composite multilingual dataset, 2023
Hugo Laurençon, Lucile Saulnier, et al · 2023
Closest in time.
How do transformers learn topic structure: Towards a mechanistic understanding, 2023
Yuchen Li, Yuanzhi Li, and Andrej Risteski · 2023
Closest in time.
The quantization model of neural scaling, 2023
Eric J. Michaud, Ziming Liu, Uzay Girit, and Max Tegmark · 2023
Closest in time.
Introducing mpt-7b: A new standard for open-source, commercially usable llms, 2023
MosaicML · 2023
Closest in time.
Progress measures for grokking via mechanistic interpretability
Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt · 2023
Closest in time.
Why think step-by-step? reasoning emerges from the locality of experience, 2023
Ben Prystawski and Noah D. Goodman · 2023
Closest in time.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Closest in time.
Redpajama-data: An open source recipe to reproduce llama training dataset, 2023
Together · 2023
Closest in time.
Llama: Open and efficient foundation language models, 2023
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample · 2023
Closest in time.
Farewell to aimless large-scale pretraining: Influential subset selection for language model, 2023
Xiao Wang, Weikang Zhou, Qi Zhang, Jie Zhou, Songyang Gao, Junzhe Wang, Menghan Zhang, Xiang Gao, Yunwen Chen, and Tao Gui · 2023
Closest in time.
Doremi: Optimizing data mixtures speeds up language model pretraining, 2023
Sang Michael Xie, Hieu Pham, Xuanyi Dong, Nan Du, Hanxiao Liu, Yifeng Lu, Percy Liang, Quoc V. Le, Tengyu Ma, and Adams Wei Yu · 2023
Closest in time.
Data selection for language models via importance resampling, 2023
Sang Michael Xie, Shibani Santurkar, Tengyu Ma, and Percy Liang · 2023
Closest in time.