Fetching the paper…
Reading the bibliography…
Recently published work on rephrasing natural text data for pre-training LLMs has shown promising results when combining the original dataset with the synthetically rephrased data.
Slurm: Simple linux utility for resource management
Andy B Yoo, Morris A Jette, and Mark Grondona · 2003
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality, 2013
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg Corrado, and Jeffrey Dean · 2013
Earlier work this paper cites.
Character-level convolutional networks for text classification, 2016
Xiang Zhang, Junbo Zhao, and Yann LeCun · 2016
Earlier work this paper cites.
The lambada dataset: Word prediction requiring a broad discourse context, 2016
Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Quan Ngoc Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández · 2016
Earlier work this paper cites.
Crowdsourcing multiple choice science questions, 2017
Johannes Welbl, Nelson F. Liu, and Matt Gardner · 2017
Earlier work this paper cites.
Text data augmentation made simple by leveraging nlp cloud apis, 2018
Claude Coulombe · 2018
Earlier work this paper cites.
Sequence-to-sequence data augmentation for dialogue language understanding, 2018
Yutai Hou, Yijia Liu, Wanxiang Che, and Ting Liu · 2018
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge, 2018
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord · 2018
Earlier work this paper cites.
Ccnet: Extracting high quality monolingual datasets from web crawl data, 2019
Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzmán, Armand Joulin, and Edouard Grave · 2019
Earlier work this paper cites.
EDA: Easy data augmentation techniques for boosting performance on text classification tasks
Jason Wei and Kai Zou · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?, 2019
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 2019
Earlier work this paper cites.
Winogrande: An adversarial winograd schema challenge at scale, 2019
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi · 2019
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language, Apr. 2020
Yonatan Bisk, Rowan Zellers, Ronan Le bras, Jianfeng Gao, and Yejin Choi · 2020
Earlier work this paper cites.
Data augmentation approaches in natural language processing: A survey
Bohan Li, Yutai Hou, and Wanxiang Che · 2022
Earlier work this paper cites.
Whose language counts as high quality? measuring language ideologies in text data selection
Suchin Gururangan, Dallas Card, Sarah Dreier, Emily Gade, Leroy Wang, Zeyu Wang, Luke Zettlemoyer, and Noah A. Smith · 2022
Earlier work this paper cites.
Deduplicating training data makes language models better
Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini · 2022
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer, 2023
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2023
Earlier work this paper cites.
Culturax: A cleaned, enormous, and multilingual dataset for large language models in 167 languages, 2023
Thuat Nguyen, Chien Van Nguyen, Viet Dac Lai, Hieu Man, Nghia Trung Ngo, Franck Dernoncourt, Ryan A. Rossi, and Thien Huu Nguyen · 2023
Earlier work this paper cites.
Mistral 7b, 2023
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed · 2023
Earlier work this paper cites.
A pretrainer’s guide to training data: Measuring the effects of data age, domain coverage, quality, & toxicity, 2023
Shayne Longpre, Gregory Yauney, Emily Reif, Katherine Lee, Adam Roberts, Barret Zoph, Denny Zhou, Jason Wei, Kevin Robinson, David Mimno, and Daphne Ippolito · 2023
Cited alongside, same era.
The refinedweb dataset for falcon llm: Outperforming curated corpora with web data, and web data only, 2023
Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay · 2023
Cited alongside, same era.
Semdedup: Data-efficient learning at web-scale through semantic deduplication, 2023
Amro Abbas, Kushal Tirumala, Dániel Simig, Surya Ganguli, and Ari S. Morcos · 2023
Cited alongside, same era.
Doremi: Optimizing data mixtures speeds up language model pretraining, 2023
Sang Michael Xie, Hieu Pham, Xuanyi Dong, Nan Du, Hanxiao Liu, Yifeng Lu, Percy Liang, Quoc V. Le, Tengyu Ma, and Adams Wei Yu · 2023
Cited alongside, same era.
Tinystories: How small can language models be and still speak coherent english?, 2023
Cosmopedia: how to create large-scale synthetic data for pre-training
Daniel van Strien Loubna Ben Allal, Anton Lozhkov · 2024
Closest in time.
Self-alignment with instruction backtranslation, 2024
Xian Li, Ping Yu, Chunting Zhou, Timo Schick, Omer Levy, Luke Zettlemoyer, Jason Weston, and Mike Lewis · 2024
Closest in time.
Lab: Large-scale alignment for chatbots, 2024
Shivchander Sudalairaj, Abhishek Bhandwaldar, Aldo Pareja, Kai Xu, David D. Cox, and Akash Srivastava · 2024
Closest in time.
Magpie: Alignment data synthesis from scratch by prompting aligned llms with nothing, 2024
Zhangchen Xu, Fengqing Jiang, Luyao Niu, Yuntian Deng, Radha Poovendran, Yejin Choi, and Bill Yuchen Lin · 2024
Closest in time.
Synthetic data (almost) from scratch: Generalized instruction tuning for language models, 2024
Haoran Li, Qingxiu Dong, Zhengyang Tang, Chaojun Wang, Xingxing Zhang, Haoyang Huang, Shaohan Huang, Xiaolong Huang, Zeqiang Huang, Dongdong Zhang, Yuxian Gu, Xin Cheng, Xun Wang, Si-Qing Chen, Li Dong, Wei Lu, Zhifang Sui, Benyou Wang, Wai Lam, and Furu Wei · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ronen Eldan and Yuanzhi Li · 2023
Cited alongside, same era.
Textbooks are all you need, 2023
Suriya Gunasekar, Yi Zhang, Jyoti Aneja, Caio César Teodoro Mendes, Allie Del Giorno, Sivakanth Gopi, Mojan Javaheripi, Piero Kauffmann, Gustavo de Rosa, Olli Saarikivi, Adil Salim, Shital Shah, Harkirat Singh Behl, Xin Wang, Sébastien Bubeck, Ronen Eldan, Adam Tauman Kalai, Yin Tat Lee, and Yuanzhi Li · 2023
Cited alongside, same era.
Wizardlm: Empowering large language models to follow complex instructions, 2023
Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, and Daxin Jiang · 2023
Cited alongside, same era.
Open llm leaderboard (2023-2024)
Edward Beeching, Clémentine Fourrier, Nathan Habib, Sheon Han, Nathan Lambert, Nazneen Rajani, Omar Sanseviero, Lewis Tunstall, and Thomas Wolf · 2023
Cited alongside, same era.
Self-consuming generative models go mad, 2023
Sina Alemohammad, Josue Casco-Rodriguez, Lorenzo Luzi, Ahmed Imtiaz Humayun, Hossein Babaei, Daniel LeJeune, Ali Siahkoohi, and Richard G. Baraniuk · 2023
Cited alongside, same era.
Enhancing chat language models by scaling high-quality instructional conversations, 2023
Ning Ding, Yulin Chen, Bokai Xu, Yujia Qin, Zhi Zheng, Shengding Hu, Zhiyuan Liu, Maosong Sun, and Bowen Zhou · 2023
Cited alongside, same era.
Efficient memory management for large language model serving with pagedattention, 2023
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica · 2023
Cited alongside, same era.
Will we run out of data? limits of llm scaling based on human-generated data, 2024
Pablo Villalobos, Anson Ho, Jaime Sevilla, Tamay Besiroglu, Lennart Heim, and Marius Hobbhahn · 2024
Cited alongside, same era.
Codeclm: Aligning language models with tailored synthetic data, 2024
Zifeng Wang, Chun-Liang Li, Vincent Perot, Long T. Le, Jin Miao, Zizhao Zhang, Chen-Yu Lee, and Tomas Pfister · 2024
Closest in time.
Scaling synthetic data creation with 1,000,000,000 personas, 2024
Tao Ge, Xin Chan, Xiaoyang Wang, Dian Yu, Haitao Mi, and Dong Yu · 2024
Closest in time.
Improving black-box robustness with in-context rewriting, 2024
Kyle O’Brien, Nathan Ng, Isha Puri, Jorge Mendez, Hamid Palangi, Yoon Kim, Marzyeh Ghassemi, and Thomas Hartvigsen · 2024
Closest in time.
Best practices and lessons learned on synthetic data for language models, 2024
Ruibo Liu, Jerry Wei, Fangyu Liu, Chenglei Si, Yanzhe Zhang, Jinmeng Rao, Steven Zheng, Daiyi Peng, Diyi Yang, Denny Zhou, and Andrew M. Dai · 2024
Closest in time.
Generative ai for synthetic data generation: Methods, challenges and the future, 2024
Xu Guo and Yiqiang Chen · 2024
Closest in time.
On llms-driven synthetic data generation, curation, and evaluation: A survey, 2024
Lin Long, Rui Wang, Ruixuan Xiao, Junbo Zhao, Xiao Ding, Gang Chen, and Haobo Wang · 2024
Closest in time.
Frontiers in synthetic data
Nathan Lambert · 2024
Closest in time.
Stable lm 2 1.6b technical report, 2024
Marco Bellagente, Jonathan Tow, Dakota Mahan, Duy Phung, Maksym Zhuravinskyi, Reshinth Adithyan, James Baicoianu, Ben Brooks, Nathan Cooper, Ashish Datta, Meng Lee, Emad Mostaque, Michael Pieler, Nikhil Pinnaparju, Paulo Rocha, Harry Saini, Hannah Teufel, Niccolo Zanichelli, and Carlos Riquelme · 2024
Closest in time.
How to train data-efficient llms, 2024
Noveen Sachdeva, Benjamin Coleman, Wang-Cheng Kang, Jianmo Ni, Lichan Hong, Ed H. Chi, James Caverlee, Julian McAuley, and Derek Zhiyuan Cheng · 2024
Closest in time.
A framework for few-shot language model evaluation, 07 2024
Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac’h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang Sutawika, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou · 2024
Closest in time.
Open llm leaderboard v2
Clémentine Fourrier, Nathan Habib, Alina Lozovskaya, Konrad Szafer, and Thomas Wolf · 2024
Closest in time.
Is model collapse inevitable? breaking the curse of recursion by accumulating real and synthetic data, 2024
Matthias Gerstgrasser, Rylan Schaeffer, Apratim Dey, Rafael Rafailov, Henry Sleight, John Hughes, Tomasz Korbak, Rajashree Agrawal, Dhruv Pai, Andrey Gromov, Daniel A. Roberts, Diyi Yang, David L. Donoho, and Sanmi Koyejo · 2024
Closest in time.
Training on the test task confounds evaluation and emergence, 2024
Ricardo Dominguez-Olmedo, Florian E. Dorner, and Moritz Hardt · 2024
Closest in time.