Fetching the paper…
Reading the bibliography…
A major factor in the recent success of large language models is the use of enormous and ever-growing text datasets for unsupervised pre-training.
Space/time trade-offs in hash coding with allowable errors
Burton H. Bloom · 1970
Earlier work this paper cites.
D-optimality for regression designs: A review
R. C. St. John and N. R. Draper · 1975
Earlier work this paper cites.
Suffix arrays: A new method for on-line string searches
Udi Manber and Gene Myers · 1993
Earlier work this paper cites.
On the resemblance and containment of documents
A.Z. Broder · 1997
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2001
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Peter Auer, Nicolò Cesa-Bianchi, Yoav Freund, and Robert E. Schapire · 2002
Earlier work this paper cites.
Similarity estimation techniques from rounding algorithms
Moses S. Charikar · 2002
Earlier work this paper cites.
On-line construction of compact directed acyclic word graphs
Shunsuke Inenaga, Hiromasa Hoshino, Ayumi Shinohara, Masayuki Takeda, Setsuo Arikawa, Giancarlo Mauri, and Giulio Pavesi · 2004
Earlier work this paper cites.
Training data selection for support vector machines
Jigang Wang, Predrag Neskovic, and Leon N. Cooper · 2005
Earlier work this paper cites.
Text data acquisition for domain-specific language models
Abhinav Sethy, Panayiotis G. Georgiou, and Shrikanth Narayanan · 2006
Earlier work this paper cites.
Url normalization for de-duplication of web pages
Amit Agarwal, Hema Swetha Koppula, Krishna P. Leela, Krishna Prasad Chitrapura, Sachin Garg, Pavan Kumar GM, Chittaranjan Haty, Anirban Roy, and Amit Sasturkar · 2009
Earlier work this paper cites.
KenLM: Faster and smaller language model queries
Kenneth Heafield · 2011
Earlier work this paper cites.
Super-samples from kernel herding
Yutian Chen, Max Welling, and Alex Smola · 2012
Earlier work this paper cites.
Data preprocessing techniques for classification without discrimination
Faisal Kamiran and Toon Calders · 2012
Earlier work this paper cites.
Machine Learning: A Probabilistic Perspective
Kevin P. Murphy · 2012
Earlier work this paper cites.
More Effective Boilerplate Removal - the GoldMiner Algorithm
István Endrédy and Attila Novák · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts · 2013
Earlier work this paper cites.
Data curation: A study of researcher practices and needs
Merinda McLure, Allison V Level, Catherine L Cranston, Beth Oehlerts, and Mike Culbertson · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning · 2015
Earlier work this paper cites.
Detecting data errors: Where are we and what needs to be done?
Ziawasch Abedjan, Xu Chu, Dong Deng, Raul Castro Fernandez, Ihab F. Ilyas, Mourad Ouzzani, Paolo Papotti, Michael Stonebraker, and Nan Tang · 2016
Earlier work this paper cites.
Demographic dialectal variation in social media: A case study of African-American English
Su Lin Blodgett, Lisa Green, and Brendan O’Connor · 2016
Earlier work this paper cites.
Big data curation
André Freitas and Edward Curry · 2016
Earlier work this paper cites.
Fasttext. zip: Compressing text classification models
Armand Joulin, Edouard Grave, Piotr Bojanowski, Matthijs Douze, Hérve Jégou, and Tomas Mikolov · 2016
Earlier work this paper cites.
Cynical selection of language model training data
Amittai Axelrod · 2017
Earlier work this paper cites.
Active bias: Training more accurate neural networks by emphasizing high variance samples
Haw-Shiuan Chang, Erik G. Learned-Miller, and Andrew McCallum · 2017
Earlier work this paper cites.
Automated curriculum learning for neural networks
Alex Graves, Marc G. Bellemare, Jacob Menick, Rémi Munos, and Koray Kavukcuoglu · 2017
Earlier work this paper cites.
The secret sharer: Evaluating and testing unintended memorization in neural networks
Nicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos, and Dawn Xiaodong Song · 2018
Earlier work this paper cites.
Learning word vectors for 157 languages
Edouard Grave, Piotr Bojanowski, Prakhar Gupta, Armand Joulin, and Tomas Mikolov · 2018
Earlier work this paper cites.
Learning adversarially fair and transferable representations
David Madras, Elliot Creager, Toniann Pitassi, and Richard Zemel · 2018
Earlier work this paper cites.
The natural language decathlon: Multitask learning as question answering
Bryan McCann, Nitish Shirish Keskar, Caiming Xiong, and Richard Socher · 2018
Earlier work this paper cites.
Deep contextualized word representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer · 2018
Earlier work this paper cites.
Sentence encoders on stilts: Supplementary training on intermediate labeled-data tasks
Jason Phang, Thibault Févry, and Samuel R Bowman · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training, 2018
Alec Radford and Karthik Narasimhan · 2018
Earlier work this paper cites.
Active learning for convolutional neural networks: A core-set approach
Ozan Sener and Silvio Savarese · 2018
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman · 2018
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman · 2018
Earlier work this paper cites.
One-network adversarial fairness
Tameem Adel, Isabel Valera, Zoubin Ghahramani, and Adrian Weller · 2019
Earlier work this paper cites.
Selection via proxy: Efficient data selection for deep learning
Cody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman, Peter Bailis, Percy Liang, Jure Leskovec, and Matei Zaharia · 2019
Earlier work this paper cites.
Cross-lingual language model pretraining
Alexis Conneau and Guillaume Lample · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Data shapley: Equitable valuation of data for machine learning
Amirata Ghorbani and James Zou · 2019
Earlier work this paper cites.
Unifying question answering, text classification, and regression via span extraction
Nitish Shirish Keskar, Bryan McCann, Caiming Xiong, and Richard Socher · 2019
Earlier work this paper cites.
Multi-task deep neural networks for natural language understanding
Xiaodong Liu, Pengcheng He, Weizhu Chen, and Jianfeng Gao · 2019
Earlier work this paper cites.
Distributionally robust language modeling
Yonatan Oren, Shiori Sagawa, Tatsunori B. Hashimoto, and Percy Liang · 2019
Earlier work this paper cites.
Asynchronous pipelines for processing huge corpora on medium to low resource infrastructures
Pedro Javier Ortiz Suárez, Benoît Sagot, and Laurent Romary · 2019
Earlier work this paper cites.
Sentence-BERT: Sentence embeddings using Siamese BERT-networks
Nils Reimers and Iryna Gurevych · 2019
Earlier work this paper cites.
Fairprep: Promoting data to a first-class citizen in studies on fairness-enhancing interventions
Sebastian Schelter, Yuxuan He, Jatin Khilnani, and Julia Stoyanovich · 2019
Earlier work this paper cites.
Balanced datasets are not enough: Estimating and mitigating gender bias in deep image representations
Tianlu Wang, Jieyu Zhao, Mark Yatskar, Kai-Wei Chang, and Vicente Ordonez · 2019
Earlier work this paper cites.
Neural network acceptability judgments
Alex Warstadt, Amanpreet Singh, and Samuel R Bowman · 2019
Earlier work this paper cites.
Coresets via bilevel optimization for continual learning and streaming
Zalán Borsos, Mojmir Mutny, and Andreas Krause · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D. Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
Extracting training data from large language models
Nicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom B. Brown, Dawn Xiaodong Song, Úlfar Erlingsson, Alina Oprea, and Colin Raffel · 2020
Earlier work this paper cites.
Language id in the wild: Unexpected challenges on the path to a thousand-language web text corpus
Isaac Caswell, Theresa Breiner, Daan van Esch, and Ankur Bapna · 2020
Earlier work this paper cites.
Selection via proxy: Efficient data selection for deep learning
Cody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman, Peter Bailis, Percy Liang, Jure Leskovec, and Matei Zaharia · 2020
Earlier work this paper cites.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov · 2020
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al · 2020
Earlier work this paper cites.
Identifying and correcting label bias in machine learning
Heinrich Jiang and Ofir Nachum · 2020
Earlier work this paper cites.
UnifiedQA: Crossing format boundaries with a single QA system
Daniel Khashabi, Sewon Min, Tushar Khot, Ashish Sabharwal, Oyvind Tafjord, Peter Clark, and Hannaneh Hajishirzi · 2020
Earlier work this paper cites.
Dataset meta-learning from kernel ridge-regression
Timothy Nguyen, Zhourong Chen, and Jaehoon Lee · 2020
Earlier work this paper cites.
Identifying mislabeled data using the area under the margin ranking
Geoff Pleiss, Tianyi Zhang, Ethan Elenberg, and Kilian Q Weinberger · 2020
Earlier work this paper cites.
Intermediate-task transfer learning with pretrained language models: When and why does it work?
Yada Pruksachatkun, Jason Phang, Haokun Liu, Phu Mon Htut, Xiaoyi Zhang, Richard Yuanzhe Pang, Clara Vania, Katharina Kann, and Samuel R. Bowman · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2020
Earlier work this paper cites.
Distributionally robust neural networks
Shiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, and Percy Liang · 2020
Earlier work this paper cites.
Apricot: Submodular selection for data summarization in python
Jacob Schreiber, Jeffrey Bilmes, and William Stafford Noble · 2020
Earlier work this paper cites.
Carpe diem, seize the samples uncertain "at the moment" for adaptive batch selection
Hwanjun Song, Minseok Kim, Sundong Kim, and Jae-Gil Lee · 2020
Earlier work this paper cites.
Dataset cartography: Mapping and diagnosing datasets with training dynamics
Swabha Swayamdipta, Roy Schwartz, Nicholas Lourie, Yizhong Wang, Hannaneh Hajishirzi, Noah A. Smith, and Yejin Choi · 2020
Earlier work this paper cites.
Data curation with deep learning
Saravanan Thirumuruganathan, Nan Tang, Mourad Ouzzani, and AnHai Doan · 2020
Earlier work this paper cites.
CCNet: Extracting high quality monolingual datasets from web crawl data
Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzmán, Armand Joulin, and Edouard Grave · 2020
Earlier work this paper cites.
Curriculum learning for natural language understanding
Benfeng Xu, Licheng Zhang, Zhendong Mao, Quan Wang, Hongtao Xie, and Yongdong Zhang · 2020
Earlier work this paper cites.
Fair Class Balancing: Enhancing Model Fairness without Observing Sensitive Attributes
Shen Yan, Hsien-te Kao, and Emilio Ferrara · 2020
Earlier work this paper cites.
Towards fairer datasets: filtering and balancing the distribution of the people subtree in the ImageNet hierarchy
Kaiyu Yang, Klint Qinami, Li Fei-Fei, Jia Deng, and Olga Russakovsky · 2020
Earlier work this paper cites.
Dataset condensation with gradient matching
Bo Zhao, Konda Reddy Mopuri, and Hakan Bilen · 2020
Earlier work this paper cites.
Muppet: Massive multi-task representations with pre-finetuning
Armen Aghajanyan, Anchit Gupta, Akshat Shrivastava, Xilun Chen, Luke Zettlemoyer, and Sonal Gupta · 2021
Earlier work this paper cites.
Ext5: Towards extreme multi-task scaling for transfer learning
Vamsi Aribandi, Yi Tay, Tal Schuster, Jinfeng Rao, Huaixiu Steven Zheng, Sanket Vaibhav Mehta, Honglei Zhuang, Vinh Q Tran, Dara Bahri, Jianmo Ni, et al · 2021
Earlier work this paper cites.
Program synthesis with large language models
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al · 2021
Earlier work this paper cites.
Fair preprocessing: towards understanding compositional fairness of data transformers in machine learning pipeline
Sumon Biswas and Hridesh Rajan · 2021
Earlier work this paper cites.
When is memorization of irrelevant training data necessary for high-accuracy learning?
Gavin Brown, Mark Bun, Vitaly Feldman, Adam Smith, and Kunal Talwar · 2021
Earlier work this paper cites.
Bias in machine learning software: why? how? what to do?
Joymallya Chakraborty, Suvodeep Majumder, and Tim Menzies · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al · 2021
Earlier work this paper cites.
Nl-augmenter: A framework for task-sensitive natural language augmentation
Kaustubh D Dhole, Varun Gangal, Sebastian Gehrmann, Aadesh Gupta, Zhenhao Li, Saad Mahamood, Abinaya Mahendiran, Simon Mille, Ashish Shrivastava, Samson Tan, et al · 2021
Earlier work this paper cites.
Documenting large webtext corpora: A case study on the colossal clean crawled corpus
Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner · 2021
Earlier work this paper cites.
A survey of data augmentation approaches for NLP
Steven Y. Feng, Varun Gangal, Jason Wei, Sarath Chandar, Soroush Vosoughi, Teruko Mitamura, and Eduard Hovy · 2021
Earlier work this paper cites.
Lightweight inspection of data preprocessing in native machine learning pipelines
Stefan Grafberger, Julia Stoyanovich, and Sebastian Schelter · 2021
Earlier work this paper cites.
Measuring coding challenge competence with apps
Dan Hendrycks, Steven Basart, Saurav Kadavath, Mantas Mazeika, Akul Arora, Ethan Guo, Collin Burns, Samir Puranik, Horace He, Dawn Song, and Jacob Steinhardt · 2021
Earlier work this paper cites.
On the complementarity of data selection and fine tuning for domain adaptation
Dan Iter and David Grangier · 2021
Earlier work this paper cites.
Cleanml: A study for evaluating the impact of data cleaning on ml classification tasks
Peng Li, Xi Rao, Jennifer Blase, Yue Zhang, Xu Chu, and Ce Zhang · 2021
Earlier work this paper cites.
What’s in the box? an analysis of undesirable content in the Common Crawl corpus
Alexandra Luccioni and Joseph Viviano · 2021
Earlier work this paper cites.
Cross-task generalization via natural language crowdsourcing instructions
Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi · 2021
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al · 2021
Earlier work this paper cites.
Dataset distillation with infinitely wide convolutional networks
Timothy Nguyen, Roman Novak, Lechao Xiao, and Jaehoon Lee · 2021
Cited alongside, same era.
Deep learning on a data diet: Finding important examples early in training
Mansheej Paul, Surya Ganguli, and Gintare Karolina Dziugaite · 2021
Cited alongside, same era.
True few-shot learning with language models
Ethan Perez, Douwe Kiela, and Kyunghyun Cho · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
Scaling language models: Methods, analysis & insights from training gopher
Jack W Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, et al · 2021
Cited alongside, same era.
Kazuma Hashimoto, Karthik Raman, and Michael Bendersky · 2023
Later among the works it cites.
In-context demonstration selection with cross entropy difference
Dan Iter, Reid Pryzant, Ruochen Xu, Shuohang Wang, Yang Liu, Yichong Xu, and Chenguang Zhu · 2023
Later among the works it cites.
Data-efficient finetuning using cross-task nearest neighbors
Hamish Ivison, Noah A. Smith, Hannaneh Hajishirzi, and Pradeep Dasigi · 2023
Later among the works it cites.
Stop uploading test data in plain text: Practical strategies for mitigating data contamination by evaluation benchmarks
Alon Jacovi, Avi Caciularu, Omer Goldman, and Yoav Goldberg · 2023
Later among the works it cites.
A note on interpreting canary exposure
Matthew Jagielski · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A survey of deep active learning
Pengzhen Ren, Yun Xiao, Xiaojun Chang, Po-Yao Huang, Zhihui Li, Brij B. Gupta, Xiaojiang Chen, and Xin Wang · 2021
Cited alongside, same era.
FAWOS: Fairness-Aware Oversampling Algorithm Based on Distributions of Sensitive Attributes
Teresa Salazar, Miriam Seoane Santos, Helder Araújo, and Pedro Henriques Abreu · 2021
Cited alongside, same era.
A reproduction of apple’s bi-directional LSTM models for language identification in short strings
Mads Toftrup, Søren Asger Sørensen, Manuel R. Ciosici, and Ira Assent · 2021
Cited alongside, same era.
Finetuned language models are zero-shot learners
Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le · 2021
Cited alongside, same era.
Challenges in detoxifying language models
Johannes Welbl, Amelia Glaese, Jonathan Uesato, Sumanth Dathathri, John Mellor, Lisa Anne Hendricks, Kirsty Anderson, Pushmeet Kohli, Ben Coppin, and Po-Sen Huang · 2021
Cited alongside, same era.
Detoxifying language models risks marginalizing minority voices
Albert Xu, Eshaan Pathak, Eric Wallace, Suchin Gururangan, Maarten Sap, and Dan Klein · 2021
Cited alongside, same era.
mT5: A massively multilingual pre-trained text-to-text transformer
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel · 2021
Cited alongside, same era.
Later among the works it cites.
The minipile challenge for data-efficient language models
Jean Kaddour · 2023
Later among the works it cites.
No train no gain: Revisiting efficient training algorithms for transformer-based language models
Jean Kaddour, Oscar Key, Piotr Nawrot, Pasquale Minervini, and Matt Kusner · 2023
Later among the works it cites.
Prometheus: Inducing fine-grained evaluation capability in language models
Seungone Kim, Jamin Shin, Yejin Cho, Joel Jang, Shayne Longpre, Hwaran Lee, Sangdoo Yun, Seongjin Shin, Sungdong Kim, James Thorne, et al · 2023
Later among the works it cites.
Openassistant conversations–democratizing large language model alignment
Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotiris Anagnostidis, Zhi-Rui Tam, Keith Stevens, Abdullah Barhoum, Nguyen Minh Duc, Oliver Stanley, Richárd Nagyfi, et al · 2023
Later among the works it cites.
Madlad-400: A multilingual and document-level large audited dataset
Sneha Kudugunta, Isaac Caswell, Biao Zhang, Xavier Garcia, Christopher A Choquette-Choo, Katherine Lee, Derrick Xin, Aditya Kusupati, Romi Stella, Ankur Bapna, et al · 2023
Later among the works it cites.
Active instruction tuning: Improving cross-task generalization by training on prompt sensitive tasks
Po-Nien Kung, Fan Yin, Di Wu, Kai-Wei Chang, and Nanyun Peng · 2023
Later among the works it cites.
Ds-1000: a natural and reliable benchmark for data science code generation
Yuhang Lai, Chengxi Li, Yiming Wang, Tianyi Zhang, Ruiqi Zhong, Luke Zettlemoyer, Wen-tau Yih, Daniel Fried, Sida Wang, and Tao Yu · 2023
Later among the works it cites.
Making large language models better data creators
Dong-Ho Lee, Jay Pujara, Mohit Sewak, Ryen White, and Sujay Jauhar · 2023
Later among the works it cites.
Diverse demonstrations improve in-context compositional generalization
Itay Levy, Ben Bogin, and Jonathan Berant · 2023
Later among the works it cites.
Finding supporting examples for in-context learning
Xiaonan Li and Xipeng Qiu · 2023
Later among the works it cites.
Unified demonstration retriever for in-context learning
Xiaonan Li, Kai Lv, Hang Yan, Tianyang Lin, Wei Zhu, Yuan Ni, Guotong Xie, Xiaoling Wang, and Xipeng Qiu · 2023
Later among the works it cites.
Slimorca: An open dataset of gpt-4 augmented flan reasoning traces, with verification, 2023
Wing Lian, Guan Wang, Bleys Goodson, Eugene Pentland, Austin Cook, Chanvichet Vong, and "Teknium" · 2023
Later among the works it cites.
Fingpt: Large generative models for a small language
Risto Luukkonen, Ville Komulainen, Jouni Luoma, Anni Eskelinen, Jenna Kanerva, Hanna-Mari Kupari, Filip Ginter, Veronika Laippala, Niklas Muennighoff, Aleksandra Piktus, et al · 2023
Later among the works it cites.
Paloma: A benchmark for evaluating language model fit
Ian Magnusson, Akshita Bhagia, Valentin Hofmann, Luca Soldaini, Ananya Harsh Jha, Oyvind Tafjord, Dustin Schwenk, Evan Pete Walsh, Yanai Elazar, Kyle Lo, et al · 2023
Later among the works it cites.
Which examples to annotate for in-context learning? towards effective and efficient selection
Costas Mavromatis, Balasubramaniam Srinivasan, Zhengyuan Shen, Jiani Zhang, Huzefa Rangwala, Christos Faloutsos, and George Karypis · 2023
Later among the works it cites.
Dataperf: Benchmarks for data-centric AI development
Mark Mazumder, Colby Banbury, Xiaozhe Yao, Bojan Karlaš, William A Gaviria Rojas, Sudnya Diamos, Greg Diamos, Lynn He, Alicia Parrish, Hannah Rose Kirk, Jessica Quaye, Charvi Rastogi, Douwe Kiela, David Jurado, David Kanter, Rafael Mosquera, Will Cukierski, Juan Ciro, Lora Aroyo, Bilge Acun, Lingjiao Chen, Mehul Smriti Raje, Max Bartolo, Sabri Eyuboglu, Amirata Ghorbani, Emmett Daniel Goodman, Addison Howard, Oana Inel, Tariq Kane, Christine Kirkpatrick, D. Sculley, Tzu-Sheng Kuo, Jonas Mueller, Tristan Thrush, Joaquin Vanschoren, Margaret Warren, Adina Williams, Serena Yeung, Newsha Ardalani, Praveen Paritosh, Ce Zhang, James Y. Zou, Carole-Jean Wu, Cody Coleman, Andrew Ng, Peter Mattson, and Vijay Janapa Reddi · 2023
Later among the works it cites.
Large-scale near-deduplication behind bigcode, May 2023
Chenghao Mou · 2023
Later among the works it cites.
In-context example selection with influences
Tai Nguyen and Eric Wong · 2023
Later among the works it cites.
Logic-LM: Empowering large language models with symbolic solvers for faithful logical reasoning
Liangming Pan, Alon Albalak, Xinyi Wang, and William Wang · 2023
Later among the works it cites.
Trak: Attributing model behavior at scale
Sung Min Park, Kristian Georgiev, Andrew Ilyas, Guillaume Leclerc, and Aleksander Madry · 2023
Later among the works it cites.
The refinedweb dataset for falcon llm: Outperforming curated corpora with web data only
Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Hamza Alobeidli, Alessandro Cappelli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay · 2023
Later among the works it cites.
Dynamic pretraining of vision-language models
AJ Piergiovanni, Weicheng Kuo, Wei Li, and Anelia Angelova · 2023
Later among the works it cites.
GAIA search: Hugging face and pyserini interoperability for NLP training data exploration
Aleksandra Piktus, Odunayo Ogundepo, Christopher Akiki, Akintunde Oladipo, Xinyu Zhang, Hailey Schoelkopf, Stella Biderman, Martin Potthast, and Jimmy Lin · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D Manning, and Chelsea Finn · 2023
Later among the works it cites.
No robots
Nazneen Rajani, Lewis Tunstall, Edward Beeching, Nathan Lambert, Alexander M. Rush, and Thomas Wolf · 2023
Later among the works it cites.
Paul tremblay, mona awad vs. openai, inc., et al., 2023
Joseph R. Saveri, Cadio Zirpoli, Christopher K.L. Young, and Kathleen J. McMahon · 2023
Later among the works it cites.
Scaling multilingual language models under constrained data
Teven Le Scao · 2023
Later among the works it cites.
Toolformer: Language models can teach themselves to use tools
Timo Schick, Jane Dwivedi-Yu, Roberto Dessi, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom · 2023
Later among the works it cites.
Cross-lingual supervision improves large language models pre-training
Andrea Schioppa, Xavier Garcia, and Orhan Firat · 2023
Later among the works it cites.
Effective robustness against natural distribution shifts for models with different training data
Zhouxing Shi, Nicholas Carlini, Ananth Balashankar, Ludwig Schmidt, Cho-Jui Hsieh, Alex Beutel, and Yao Qin · 2023
Later among the works it cites.
Fair oversampling technique using heterogeneous clusters
Ryosuke Sonoda · 2023
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R. Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, Agnieszka Kluska, Aitor Lewkowycz, Akshat Agarwal, Alethea Power, Alex Ray, Alex Warstadt, Alexander W. Kocurek, Ali Safaya, Ali Tazarv, Alice Xiang, Alicia Parrish, Allen Nie, Aman Hussain, Amanda Askell, Amanda Dsouza, Ambrose Slone, Ameet Rahane, Anantharaman S. Iyer, Anders Johan Andreassen, Andrea Madotto, Andrea Santilli, Andreas Stuhlmüller, Andrew M. Dai, Andrew La, Andrew Lampinen, Andy Zou, Angela Jiang, Angelica Chen, Anh Vuong, Animesh Gupta, Anna Gottardi, Antonio Norelli, Anu Venkatesh, Arash Gholamidavoodi, Arfa Tabassum, Arul Menezes, Arun Kirubarajan, Asher Mullokandov, Ashish Sabharwal, Austin Herrick, Avia Efrat, Aykut Erdem, Ayla Karakaş, B. Ryan Roberts, Bao Sheng Loe, Barret Zoph, Bartłomiej Bojanowski, Batuhan Özyurt, Behnam Hedayatnia, Behnam Neyshabur, Benjamin Inden, Benno Stein, Berk Ekmekci, Bill Yuchen Lin, Blake Howald, Bryan Orinion, Cameron Diao, Cameron Dour, Catherine Stinson, Cedrick Argueta, Cesar Ferri, Chandan Singh, Charles Rathkopf, Chenlin Meng, Chitta Baral, Chiyu Wu, Chris Callison-Burch, Christopher Waites, Christian Voigt, Christopher D Manning, Christopher Potts, Cindy Ramirez, Clara E. Rivera, Clemencia Siro, Colin Raffel, Courtney Ashcraft, Cristina Garbacea, Damien Sileo, Dan Garrette, Dan Hendrycks, Dan Kilman, Dan Roth, C. Daniel Freeman, Daniel Khashabi, Daniel Levy, Daniel Moseguí González, Danielle Perszyk, Danny Hernandez, Danqi Chen, Daphne Ippolito, Dar Gilboa, David Dohan, David Drakard, David Jurgens, Debajyoti Datta, Deep Ganguli, Denis Emelin, Denis Kleyko, Deniz Yuret, Derek Chen, Derek Tam, Dieuwke Hupkes, Diganta Misra, Dilyar Buzan, Dimitri Coelho Mollo, Diyi Yang, Dong-Ho Lee, Dylan Schrader, Ekaterina Shutova, Ekin Dogus Cubuk, Elad Segal, Eleanor Hagerman, Elizabeth Barnes, Elizabeth Donoway, Ellie Pavlick, Emanuele Rodolà, Emma Lam, Eric Chu, Eric Tang, Erkut Erdem, Ernie Chang, Ethan A Chi, Ethan Dyer, Ethan Jerzak, Ethan Kim, Eunice Engefu Manyasi, Evgenii Zheltonozhskii, Fanyue Xia, Fatemeh Siar, Fernando Martínez-Plumed, Francesca Happé, Francois Chollet, Frieda Rong, Gaurav Mishra, Genta Indra Winata, Gerard de Melo, Germán Kruszewski, Giambattista Parascandolo, Giorgio Mariani, Gloria Xinyue Wang, Gonzalo Jaimovitch-Lopez, Gregor Betz, Guy Gur-Ari, Hana Galijasevic, Hannah Kim, Hannah Rashkin, Hannaneh Hajishirzi, Harsh Mehta, Hayden Bogar, Henry Francis Anthony Shevlin, Hinrich Schuetze, Hiromu Yakura, Hongming Zhang, Hugh Mee Wong, Ian Ng, Isaac Noble, Jaap Jumelet, Jack Geissinger, Jackson Kernion, Jacob Hilton, Jaehoon Lee, Jaime Fernández Fisac, James B Simon, James Koppel, James Zheng, James Zou, Jan Kocon, Jana Thompson, Janelle Wingfield, Jared Kaplan, Jarema Radom, Jascha Sohl-Dickstein, Jason Phang, Jason Wei, Jason Yosinski, Jekaterina Novikova, Jelle Bosscher, Jennifer Marsh, Jeremy Kim, Jeroen Taal, Jesse Engel, Jesujoba Alabi, Jiacheng Xu, Jiaming Song, Jillian Tang, Joan Waweru, John Burden, John Miller, John U. Balis, Jonathan Batchelder, Jonathan Berant, Jörg Frohberg, Jos Rozen, Jose Hernandez-Orallo, Joseph Boudeman, Joseph Guerr, Joseph Jones, Joshua B. Tenenbaum, Joshua S. Rule, Joyce Chua, Kamil Kanclerz, Karen Livescu, Karl Krauth, Karthik Gopalakrishnan, Katerina Ignatyeva, Katja Markert, Kaustubh Dhole, Kevin Gimpel, Kevin Omondi, Kory Wallace Mathewson, Kristen Chiafullo, Ksenia Shkaruta, Kumar Shridhar, Kyle McDonell, Kyle Richardson, Laria Reynolds, Leo Gao, Li Zhang, Liam Dugan, Lianhui Qin, Lidia Contreras-Ochando, Louis-Philippe Morency, Luca Moschella, Lucas Lam, Lucy Noble, Ludwig Schmidt, Luheng He, Luis Oliveros-Colón, Luke Metz, Lütfi Kerem Senel, Maarten Bosma, Maarten Sap, Maartje Ter Hoeve, Maheen Farooqi, Manaal Faruqui, Mantas Mazeika, Marco Baturan, Marco Marelli, Marco Maru, Maria Jose Ramirez-Quintana, Marie Tolkiehn, Mario Giulianelli, Martha Lewis, Martin Potthast, Matthew L Leavitt, Matthias Hagen, Mátyás Schubert, Medina Orduna Baitemirova, Melody Arnaud, Melvin McElrath, Michael Andrew Yee, Michael Cohen, Michael Gu, Michael Ivanitskiy, Michael Starritt, Michael Strube, Michał Swędrowski, Michele Bevilacqua, Michihiro Yasunaga, Mihir Kale, Mike Cain, Mimee Xu, Mirac Suzgun, Mitch Walker, Mo Tiwari, Mohit Bansal, Moin Aminnaseri, Mor Geva, Mozhdeh Gheini, Mukund Varma T, Nanyun Peng, Nathan Andrew Chi, Nayeon Lee, Neta Gur-Ari Krakover, Nicholas Cameron, Nicholas Roberts, Nick Doiron, Nicole Martinez, Nikita Nangia, Niklas Deckers, Niklas Muennighoff, Nitish Shirish Keskar, Niveditha S. Iyer, Noah Constant, Noah Fiedel, Nuan Wen, Oliver Zhang, Omar Agha, Omar Elbaghdadi, Omer Levy, Owain Evans, Pablo Antonio Moreno Casares, Parth Doshi, Pascale Fung, Paul Pu Liang, Paul Vicol, Pegah Alipoormolabashi, Peiyuan Liao, Percy Liang, Peter W Chang, Peter Eckersley, Phu Mon Htut, Pinyu Hwang, Piotr Miłkowski, Piyush Patil, Pouya Pezeshkpour, Priti Oli, Qiaozhu Mei, Qing Lyu, Qinlang Chen, Rabin Banjade, Rachel Etta Rudolph, Raefer Gabriel, Rahel Habacker, Ramon Risco, Raphaël Millière, Rhythm Garg, Richard Barnes, Rif A. Saurous, Riku Arakawa, Robbe Raymaekers, Robert Frank, Rohan Sikand, Roman Novak, Roman Sitelew, Ronan Le Bras, Rosanne Liu, Rowan Jacobs, Rui Zhang, Russ Salakhutdinov, Ryan Andrew Chi, Seungjae Ryan Lee, Ryan Stovall, Ryan Teehan, Rylan Yang, Sahib Singh, Saif M. Mohammad, Sajant Anand, Sam Dillavou, Sam Shleifer, Sam Wiseman, Samuel Gruetter, Samuel R. Bowman, Samuel Stern Schoenholz, Sanghyun Han, Sanjeev Kwatra, Sarah A. Rous, Sarik Ghazarian, Sayan Ghosh, Sean Casey, Sebastian Bischoff, Sebastian Gehrmann, Sebastian Schuster, Sepideh Sadeghi, Shadi Hamdan, Sharon Zhou, Shashank Srivastava, Sherry Shi, Shikhar Singh, Shima Asaadi, Shixiang Shane Gu, Shubh Pachchigar, Shubham Toshniwal, Shyam Upadhyay, Shyamolima Shammie Debnath, Siamak Shakeri, Simon Thormeyer, Simone Melzi, Siva Reddy, Sneha Priscilla Makini, Soo-Hwan Lee, Spencer Torene, Sriharsha Hatwar, Stanislas Dehaene, Stefan Divic, Stefano Ermon, Stella Biderman, Stephanie Lin, Stephen Prasad, Steven Piantadosi, Stuart Shieber, Summer Misherghi, Svetlana Kiritchenko, Swaroop Mishra, Tal Linzen, Tal Schuster, Tao Li, Tao Yu, Tariq Ali, Tatsunori Hashimoto, Te-Lin Wu, Théo Desbordes, Theodore Rothschild, Thomas Phan, Tianle Wang, Tiberius Nkinyili, Timo Schick, Timofei Kornev, Titus Tunduny, Tobias Gerstenberg, Trenton Chang, Trishala Neeraj, Tushar Khot, Tyler Shultz, Uri Shaham, Vedant Misra, Vera Demberg, Victoria Nyamai, Vikas Raunak, Vinay Venkatesh Ramasesh, vinay uday prabhu, Vishakh Padmakumar, Vivek Srikumar, William Fedus, William Saunders, William Zhang, Wout Vossen, Xiang Ren, Xiaoyu Tong, Xinran Zhao, Xinyi Wu, Xudong Shen, Yadollah Yaghoobzadeh, Yair Lakretz, Yangqiu Song, Yasaman Bahri, Yejin Choi, Yichi Yang, Yiding Hao, Yifu Chen, Yonatan Belinkov, Yu Hou, Yufang Hou, Yuntao Bai, Zachary Seid, Zhuoye Zhao, Zijian Wang, Zijie J. Wang, Zirui Wang, and Ziyi Wu · 2023
Later among the works it cites.
Selective annotation makes language models better few-shot learners
Hongjin SU, Jungo Kasai, Chen Henry Wu, Weijia Shi, Tianlu Wang, Jiayi Xin, Rui Zhang, Mari Ostendorf, Luke Zettlemoyer, Noah A. Smith, and Tao Yu · 2023
Later among the works it cites.
Detecting personal information in training corpora: an analysis
Nishant Subramani, Sasha Luccioni, Jesse Dodge, and Margaret Mitchell · 2023
Later among the works it cites.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Later among the works it cites.
A survey on active learning: State-of-the-art, practical challenges and research directions
Alaa Tharwat and Wolfram Schenck · 2023
Later among the works it cites.
Zephyr: Direct distillation of lm alignment
Lewis Tunstall, Edward Beeching, Nathan Lambert, Nazneen Rajani, Kashif Rasul, Younes Belkada, Shengyi Huang, Leandro von Werra, Clémentine Fourrier, Nathan Habib, et al · 2023
Later among the works it cites.
Self-instruct: Aligning language models with self-generated instructions
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi · 2023
Later among the works it cites.
Findings of the BabyLM challenge: Sample-efficient pretraining on developmentally plausible corpora
Alex Warstadt, Aaron Mueller, Leshem Choshen, Ethan Wilcox, Chengxu Zhuang, Juan Ciro, Rafael Mosquera, Bhargavi Paranjabe, Adina Williams, Tal Linzen, and Ryan Cotterell · 2023
Later among the works it cites.
Scattershot: Interactive in-context example curation for text transformation
Sherry Wu, Hua Shen, Daniel S Weld, Jeffrey Heer, and Marco Tulio Ribeiro · 2023
Later among the works it cites.
Moderate coreset: A universal method of data selection for real-world data-efficient deep learning
Xiaobo Xia, Jiale Liu, Jun Yu, Xu Shen, Bo Han, and Tongliang Liu · 2023
Later among the works it cites.
C-pack: Packaged resources to advance general chinese embedding
Shitao Xiao, Zheng Liu, Peitian Zhang, and Niklas Muennighoff · 2023
Later among the works it cites.
Wizardlm: Empowering large language models to follow complex instructions
Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, and Daxin Jiang · 2023
Later among the works it cites.
Compositional exemplars for in-context learning
Jiacheng Ye, Zhiyong Wu, Jiangtao Feng, Tao Yu, and Lingpeng Kong · 2023
Later among the works it cites.
Scaling relationship on learning mathematical reasoning with large language models
Zheng Yuan, Hongyi Yuan, Chengpeng Li, Guanting Dong, Chuanqi Tan, and Chang Zhou · 2023
Later among the works it cites.
Dataset condensation with distribution matching
Bo Zhao and Hakan Bilen · 2023
Later among the works it cites.
LIMA: Less is more for alignment
Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, LILI YU, Susan Zhang, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, and Omer Levy · 2023
Later among the works it cites.
Starling-7b: Improving llm helpfulness & harmlessness with rlaif, November 2023
Banghua Zhu, Evan Frick, Tianhao Wu, Hanlin Zhu, and Jiantao Jiao · 2023
Later among the works it cites.
Data augmentation approaches for source code models: A survey
Terry Yue Zhuo, Zhou Yang, Zhensu Sun, Yufei Wang, Li Li, Xiaoning Du, Zhenchang Xing, and David Lo · 2023
Later among the works it cites.
Fine-tuning language models from human preferences
Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 2023
Later among the works it cites.
Theoretical guarantees on the best-of-n alignment policy
Ahmad Beirami, Alekh Agarwal, Jonathan Berant, Alexander D’Amour, Jacob Eisenstein, Chirag Nagpal, and Ananda Theertha Suresh · 2024
Closest in time.
Genie: Generative interactive environments
Jake Bruce, Michael Dennis, Ashley Edwards, Jack Parker-Holder, Yuge Shi, Edward Hughes, Matthew Lai, Aditi Mavalankar, Richie Steigerwald, Chris Apps, et al · 2024
Closest in time.
Suppressing pink elephants with direct principle feedback
Louis Castricato, Nathan Lile, Suraj Anand, Hailey Schoelkopf, Siddharth Verma, and Stella Biderman · 2024
Closest in time.
Self-play fine-tuning converts weak language models to strong language models
Zixiang Chen, Yihe Deng, Huizhuo Yuan, Kaixuan Ji, and Quanquan Gu · 2024
Closest in time.
What’s in my big data?
Yanai Elazar, Akshita Bhagia, Ian Helgi Magnusson, Abhilasha Ravichander, Dustin Schwenk, Alane Suhr, Evan Pete Walsh, Dirk Groeneveld, Luca Soldaini, Sameer Singh, Hanna Hajishirzi, Noah A. Smith, and Jesse Dodge · 2024
Closest in time.
Dsdm: Model-aware dataset selection with datamodels
Logan Engstrom, Axel Feldmann, and Aleksander Madry · 2024
Closest in time.
Kto: Model alignment as prospect theoretic optimization
Kawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky, and Douwe Kiela · 2024
Closest in time.
Data filtering networks
Alex Fang, Albin Madappally Jose, Amit Jain, Ludwig Schmidt, Alexander T Toshev, and Vaishaal Shankar · 2024
Closest in time.
Olmo: Accelerating the science of language models
Dirk Groeneveld, Iz Beltagy, Pete Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, et al · 2024
Closest in time.
Towards lossless dataset distillation via difficulty-aligned trajectory matching
Ziyao Guo, Kai Wang, George Cazenavette, Hui Li, Kaipeng Zhang, and Yang You · 2024
Closest in time.
The responsible foundation model development cheatsheet: A review of tools & resources
Shayne Longpre, Stella Biderman, Alon Albalak, Hailey Schoelkopf, Daniel McDuff, Sayash Kapoor, Kevin Klyman, Kyle Lo, Gabriel Ilharco, Nay San, et al · 2024
Closest in time.
$\mathbb{D}^2$ pruning: Message passing for balancing diversity & difficulty in data pruning
Adyasha Maharana, Prateek Yadav, and Mohit Bansal · 2024
Closest in time.
Data portraits: Recording foundation model training data
Marc Marone and Benjamin Van Durme · 2024
Closest in time.
Evaluating n n -gram novelty of language models using rusty-dawg
William Merrill, Noah A. Smith, and Yanai Elazar · 2024
Closest in time.
Large language models: A survey
Shervin Minaee, Tomas Mikolov, Narjes Nikzad, Meysam Chenaghlu, Richard Socher, Xavier Amatriain, and Jianfeng Gao · 2024
Closest in time.
Generative representational instruction tuning
Niklas Muennighoff, Hongjin Su, Liang Wang, Nan Yang, Furu Wei, Tao Yu, Amanpreet Singh, and Douwe Kiela · 2024
Closest in time.
Proving test set contamination for black-box language models
Yonatan Oren, Nicole Meister, Niladri S. Chatterji, Faisal Ladhak, and Tatsunori Hashimoto · 2024
Closest in time.
West-of-n: Synthetic preference generation for improved reward modeling
Alizée Pace, Jonathan Mallinson, Eric Malmi, Sebastian Krause, and Aliaksei Severyn · 2024
Closest in time.
Datatrove: large scale data processing, 2024
Guilherme Penedo, Alessandro Cappelli, Thomas Wolf, and Mario Sasko · 2024
Closest in time.
tinybenchmarks: evaluating llms with fewer examples
Felipe Maia Polo, Lucas Weber, Leshem Choshen, Yuekai Sun, Gongjun Xu, and Mikhail Yurochkin · 2024
Closest in time.
Rainbow teaming: Open-ended generation of diverse adversarial prompts
Mikayel Samvelyan, Sharath Chandra Raparthy, Andrei Lupu, Eric Hambro, Aram H Markosyan, Manish Bhatt, Yuning Mao, Minqi Jiang, Jack Parker-Holder, Jakob Foerster, et al · 2024
Closest in time.
Slimpajama-dc: Understanding data combinations for llm training, 2024
Zhiqiang Shen, Tianhua Tao, Liqun Ma, Willie Neiswanger, Zhengzhong Liu, Hongyi Wang, Bowen Tan, Joel Hestness, Natalia Vassilieva, Daria Soboleva, and Eric Xing · 2024
Closest in time.
Detecting pretraining data from large language models
Weijia Shi, Anirudh Ajith, Mengzhou Xia, Yangsibo Huang, Daogao Liu, Terra Blevins, Danqi Chen, and Luke Zettlemoyer · 2024
Closest in time.
Aya dataset: An open-access collection for multilingual instruction tuning
Shivalika Singh, Freddie Vargus, Daniel Dsouza, Börje F Karlsson, Abinaya Mahendiran, Wei-Yin Ko, Herumb Shandilya, Jay Patel, Deividas Mataciunas, Laura OMahony, et al · 2024
Closest in time.
Dolma: An open corpus of three trillion tokens for language model pretraining research
Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, et al · 2024
Closest in time.
D4: Improving llm pretraining via document de-duplication and diversification
Kushal Tirumala, Daniel Simig, Armen Aghajanyan, and Ari Morcos · 2024
Closest in time.
Aya model: An instruction finetuned open-access multilingual language model
Ahmet Üstün, Viraat Aryabumi, Zheng-Xin Yong, Wei-Yin Ko, Daniel D’souza, Gbemileke Onilude, Neel Bhandari, Shivalika Singh, Hui-Lee Ooi, Amr Kayid, et al · 2024
Closest in time.
How far can camels go? exploring the state of instruction tuning on open resources
Yizhong Wang, Hamish Ivison, Pradeep Dasigi, Jack Hessel, Tushar Khot, Khyathi Chandu, David Wadden, Kelsey MacMillan, Noah A Smith, Iz Beltagy, et al · 2024
Closest in time.
Qurating: Selecting high-quality data for training language models
Alexander Wettig, Aatmik Gupta, Saumya Malik, and Danqi Chen · 2024
Closest in time.
Misconfidence-based demonstration selection for llm in-context learning
Shangqing Xu and Chao Zhang · 2024
Closest in time.
Data mixing laws: Optimizing data mixtures by predicting language modeling performance, 2024
Jiasheng Ye, Peiju Liu, Tianxiang Sun, Yunhua Zhou, Jun Zhan, and Xipeng Qiu · 2024
Closest in time.
Multimodal c4: An open, billion-scale corpus of images interleaved with text
Wanrong Zhu, Jack Hessel, Anas Awadalla, Samir Yitzhak Gadre, Jesse Dodge, Alex Fang, Youngjae Yu, Ludwig Schmidt, William Yang Wang, and Yejin Choi · 2024
Closest in time.
Astraios: Parameter-efficient instruction tuning code large language models
Terry Yue Zhuo, Armel Zebaze, Nitchakarn Suppattarachai, Leandro von Werra, Harm de Vries, Qian Liu, and Niklas Muennighoff · 2024
Closest in time.
Intelligent selection of language model training data
Robert C. Moore and William Lewis · 2041
Closest in time.
A sequential algorithm for training text classifiers
David D. Lewis and William A. Gale · 2099
Closest in time.