Fetching the paper…
Reading the bibliography…
Instruction tuning plays a critical role in aligning large language models (LLMs) with human preference.
A new readability yardstick
Rudolph Flesch · 1948
Earlier work this paper cites.
A mathematical theory of communication
Claude Elwood Shannon · 1948
Earlier work this paper cites.
The concept of readability
Edgar Dale and Jeanne S Chall · 1949
Earlier work this paper cites.
Measurement of diversity
Edward H Simpson · 1949
Earlier work this paper cites.
The principle of minimized iterations in the solution of the matrix eigenvalue problem
Walter E. Arnoldi · 1951
Earlier work this paper cites.
The technique of clear writing
Robert Gunning · 1952
Earlier work this paper cites.
Certain language skills in children; their development and interrelationships
Mildred C Templin · 1957
Earlier work this paper cites.
On measures of entropy and information
Alfréd Rényi · 1961
Earlier work this paper cites.
The measurement of readability
George R Klare et al · 1963
Earlier work this paper cites.
A statistical interpretation of term specificity and its application in retrieval
Karen Sparck Jones · 1972
Earlier work this paper cites.
Assessing readability
George R Klare · 1974
Earlier work this paper cites.
The measurement of species diversity
Robert K Peet · 1974
Earlier work this paper cites.
D-optimality for regression designs: a review
RC St John and Norman R Draper · 1975
Earlier work this paper cites.
Perplexity—a measure of the difficulty of speech recognition tasks
Fred Jelinek, Robert L Mercer, Lalit R Bahl, and James K Baker · 1977
Earlier work this paper cites.
Interpolated estimation of markov source parameters from sparse data
Frederick Jelinek · 1980
Earlier work this paper cites.
The uncapicitated facility location problem
Gérard Cornuéjols, George Nemhauser, and Laurence Wolsey · 1983
Earlier work this paper cites.
Measuring the inference load of a text
Susan Kemper · 1983
Earlier work this paper cites.
The rules of spelling errors
Emmanuel J Yannakoudakis and David Fawthrop · 1983
Earlier work this paper cites.
Readability
George R Klare et al · 1984
Earlier work this paper cites.
Readability in esl
Patricia L Carrell · 1987
Earlier work this paper cites.
Type/token ratios: What do they really tell us?
Brian Richards · 1987
Earlier work this paper cites.
Principal component analysis
Svante Wold, Kim Esbensen, and Paul Geladi · 1987
Earlier work this paper cites.
Least squares methods: Handbook of numerical analysis
A Bjork · 1988
Earlier work this paper cites.
Readability: Its Past, Present, and Future
Beverly L Zakaluk and S Jay Samuels · 1988
Earlier work this paper cites.
Combinatorial optimization
William J Cook, William H Cunningham, William R Pulleyblank, and Alexander Schrijver · 1994
Earlier work this paper cites.
Fast exact multiplication by the hessian
Barak A. Pearlmutter · 1994
Earlier work this paper cites.
An introduction to the conjugate gradient method without the agonizing pain
Jonathan Richard Shewchuk et al · 1994
Earlier work this paper cites.
Readability revisited: The new dale-chall readability formula
Jeanne Sternlicht Chall and Edgar Dale · 1995
Earlier work this paper cites.
Active learning literature survey
Burr Settles · 1995
Earlier work this paper cites.
The farthest point strategy for progressive image sampling
Yuval Eldar, Michael Lindenbaum, Moshe Porat, and Yehoshua Y Zeevi · 1997
Earlier work this paper cites.
A new measure of lexical diversity
David D Malvern and Brian J Richards · 1997
Earlier work this paper cites.
Last rites for readability formulas in technical communication
Bradford R Connatser · 1999
Earlier work this paper cites.
A comparison of species diversity estimators
David Mouillot and Alain Lepretre · 1999
Earlier work this paper cites.
Feature selection as a preprocessing step for hierarchical clustering
Luis Talavera · 1999
Earlier work this paper cites.
The analysis of a simple k-means clustering algorithm
Tapas Kanungo, David M Mount, Nathan S Netanyahu, Christine Piatko, Ruth Silverman, and Angela Y Wu · 2000
Earlier work this paper cites.
Algorithms for non-negative matrix factorization
Daniel Lee and H Sebastian Seung · 2000
Earlier work this paper cites.
A mathematical theory of communication
Claude Elwood Shannon · 2001
Earlier work this paper cites.
A statistical model for scientific readability
Luo Si and Jamie Callan · 2001
Earlier work this paper cites.
Support vector machine active learning with applications to text classification
Simon Tong and Daphne Koller · 2001
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Measuring lexical diversity in children who stutter: Application of vocd
Stacy Silverman and Nan Bernstein Ratner · 2002
Earlier work this paper cites.
Learning spectral clustering
Francis Bach and Michael Jordan · 2003
Earlier work this paper cites.
Latent dirichlet allocation
David M Blei, Andrew Y Ng, and Michael I Jordan · 2003
Earlier work this paper cites.
The principles of readability. impact information
William H Dubay · 2004
Earlier work this paper cites.
Semi-supervised learning by entropy minimization
Yves Grandvalet and Yoshua Bengio · 2004
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Lexical diversity and language development
David Malvern, Brian Richards, Ngoni Chipere, and Pilar Durán · 2004
Earlier work this paper cites.
Predicting reading difficulty with statistical language models
Kevyn Collins-Thompson and Jamie Callan · 2005
Earlier work this paper cites.
Hubbell’s fundamental biodiversity parameter and the simpson diversity index
Fangliang He and Xin-Sheng Hu · 2005
Earlier work this paper cites.
An assessment of the range and usefulness of lexical diversity measures and the potential of the measure of textual, lexical diversity (MTLD)
Philip M McCarthy · 2005
Earlier work this paper cites.
Dqi: Measuring data quality in nlp
Swaroop Mishra, Anjana Arunkumar, Bhavdeep Sachdeva, Chris Bryan, and Chitta Baral · 2005
Earlier work this paper cites.
Reading level assessment using support vector machines and statistical language models
Sarah E Schwarm and Mari Ostendorf · 2005
Earlier work this paper cites.
Agnostic active learning
Maria-Florina Balcan, Alina Beygelzimer, and John Langford · 2006
Earlier work this paper cites.
Dataset condensation with gradient matching
Bo Zhao, Konda Reddy Mopuri, and Hakan Bilen · 2006
Earlier work this paper cites.
Dataset condensation with gradient matching, 2021
Bo Zhao, Konda Reddy Mopuri, and Hakan Bilen · 2006
Earlier work this paper cites.
An overview of bilevel optimization
Benoît Colson, Patrice Marcotte, and Gilles Savard · 2007
Earlier work this paper cites.
A tutorial on spectral clustering
Ulrike Von Luxburg · 2007
Earlier work this paper cites.
The moving-average type-token ratio
M Covington and Joe D McFall · 2008
Earlier work this paper cites.
Generalized simpson-diversity
Hans-Rolf Gregorius and Elizabeth M Gillet · 2008
Earlier work this paper cites.
Similarity measures for text document clustering
Anna Huang et al · 2008
Earlier work this paper cites.
Dqi: A guide to benchmark evaluation
Swaroop Mishra, Anjana Arunkumar, Bhavdeep Sachdeva, Chris Bryan, and Chitta Baral · 2008
Earlier work this paper cites.
Methodologies for data quality assessment and improvement
Carlo Batini, Cinzia Cappiello, Chiara Francalanci, and Andrea Maurino · 2009
Earlier work this paper cites.
Facility location: concepts, models, algorithms and case studies
Reza Zanjirani Farahani and Masoud Hekmatfar · 2009
Earlier work this paper cites.
Non negative matrix factorization clustering capabilities; application on multivariate image segmentation
Cosmin Lazar and Andrei Doncescu · 2009
Earlier work this paper cites.
K-nearest neighbor
Leif E Peterson · 2009
Earlier work this paper cites.
Herding dynamical weights to learn
Max Welling · 2009
Earlier work this paper cites.
Cutting the gordian knot: The moving-average type–token ratio (mattr)
Michael A Covington and Joe D McFall · 2010
Earlier work this paper cites.
Semi-supervised learning via generalized maximum entropy
Ayse Erkan and Yasemin Altun · 2010
Earlier work this paper cites.
A comparison of features for automatic readability assessment
Lijun Feng, Martin Jansche, Matt Huenerfauth, and Noémie Elhadad · 2010
Earlier work this paper cites.
La lisibilité computationnelle: un renouveau pour la lisibilité du français langue première et seconde?
Thomas François · 2010
Earlier work this paper cites.
Mtld, vocd-d, and hd-d: A validation study of sophisticated approaches to lexical diversity assessment
Philip M McCarthy and Scott Jarvis · 2010
Earlier work this paper cites.
Intelligent selection of language model training data
Robert C Moore and William Lewis · 2010
Earlier work this paper cites.
Non-negative matrix factorization clustering on multiple manifolds
Bin Shen and Luo Si · 2010
Earlier work this paper cites.
Efficient k-nearest neighbor graph construction for generic similarity measures
Wei Dong, Charikar Moses, and Kai Li · 2011
Earlier work this paper cites.
A unified framework for approximating and clustering data
Dan Feldman and Michael Langberg · 2011
Earlier work this paper cites.
Les apports du traitement automatique du langage à la lisibilité du français langue étrangère
Thomas François · 2011
Earlier work this paper cites.
Fast and efficient saliency detection using sparse sampling and kernel density estimation
Hamed Rezazadegan Tavakoli, Esa Rahtu, and Janne Heikkilä · 2011
Earlier work this paper cites.
From theories to queries: Active learning in practice
Burr Settles · 2011
Earlier work this paper cites.
Super-samples from kernel herding
Yutian Chen, Max Welling, and Alex Smola · 2012
Earlier work this paper cites.
Linguistic features for quality estimation
Mariano Felice and Lucia Specia · 2012
Earlier work this paper cites.
Machine learning: the art and science of algorithms that make sense of data
Peter Flach · 2012
Earlier work this paper cites.
An “ai readability” formula for french as a foreign language
Thomas François and Cédrick Fairon · 2012
Earlier work this paper cites.
Do nlp and machine learning improve traditional readability formulas?
Thomas François and Eleni Miltsakaki · 2012
Earlier work this paper cites.
Optimally-weighted herding is bayesian quadrature
Ferenc Huszár and David Duvenaud · 2012
Earlier work this paper cites.
Machine learning: a probabilistic perspective
Kevin P Murphy · 2012
Earlier work this paper cites.
Legal documents clustering using latent dirichlet allocation
K Raghuveer et al · 2012
Earlier work this paper cites.
Data quality: A survey of data quality dimensions
Fatimah Sidi, Payam Hassany Shariat Panahy, Lilly Suriani Affendey, Marzanah A Jabar, Hamidah Ibrahim, and Aida Mustapha · 2012
Earlier work this paper cites.
Nonnegative matrix factorization: A comprehensive review
Yu-Xiong Wang and Yu-Jin Zhang · 2012
Earlier work this paper cites.
Density-based clustering based on hierarchical density estimates
Ricardo JGB Campello, Davoud Moulavi, and Jörg Sander · 2013
Earlier work this paper cites.
How good are my data and what is the resolution?
Philip R Evans and Garib N Murshudov · 2013
Earlier work this paper cites.
Capturing the diversity in lexical diversity
Scott Jarvis · 2013
Earlier work this paper cites.
Defining and measuring lexical diversity
Scott Jarvis and M Daller · 2013
Earlier work this paper cites.
Learning with noisy labels
Nagarajan Natarajan, Inderjit S Dhillon, Pradeep K Ravikumar, and Ambuj Tewari · 2013
Earlier work this paper cites.
A novel approach for feature selection method tf-idf in document clustering
Leena H Patil and Mohammed Atique · 2013
Earlier work this paper cites.
Can readability formulas be used to successfully gauge difficulty of reading materials?
John C Begeny and Diana J Greene · 2014
Earlier work this paper cites.
Evaluating the comparability of two measures of lexical diversity
Fredrik deBoer · 2014
Earlier work this paper cites.
Influence function learning in information diffusion networks
Nan Du, Yingyu Liang, Maria Balcan, and Le Song · 2014
Earlier work this paper cites.
Near-optimal herding
Nick Harvey and Samira Samadi · 2014
Earlier work this paper cites.
The latest research progress on spectral clustering
Hongjie Jia, Shifei Ding, Xinzheng Xu, and Ru Nie · 2014
Earlier work this paper cites.
Can type-token ratio be used to show morphological complexity of languages?
Kimmo Kettunen · 2014
Earlier work this paper cites.
Dbscan: Past, present and future
Kamran Khan, Saif Ur Rehman, Kamran Aziz, Simon Fong, and Sababady Sarasvady · 2014
Earlier work this paper cites.
Reading comprehension and readability in educational practice and psychological theory
Walter Kintsch and Douglas Vipond · 2014
Earlier work this paper cites.
Active learning with support vector machines
Jan Kremer, Kim Steenstrup Pedersen, and Christian Igel · 2014
Earlier work this paper cites.
Fast approximation of rotations and hessians matrices
Michael Mathieu and Yann LeCun · 2014
Earlier work this paper cites.
Web document clustering and ranking using tf-idf based apriori approach
Rajendra Kumar Roul, Omanwar Rohit Devanand, and Sanjay Kumar Sahay · 2014
Earlier work this paper cites.
On variants of k-means clustering
Sayan Bandyapadhyay and Kasturi Varadarajan · 2015
Earlier work this paper cites.
A diversity-promoting objective function for neural conversation models
Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and Bill Dolan · 2015
Earlier work this paper cites.
A survey and comparative study of data deduplication techniques
Jyoti Malhotra and Jagdish Bakal · 2015
Earlier work this paper cites.
Improving diversity in image search via supervised relevance scoring
Eleftherios Spyromitros-Xioufis, Symeon Papadopoulos, Alexandru Lucian Ginsca, Adrian Popescu, Yiannis Kompatsiaris, and Ioannis Vlahavas · 2015
Earlier work this paper cites.
Second-order stochastic optimization in linear time
Naman Agarwal, Brian Bullins, and Elad Hazan · 2016
Earlier work this paper cites.
Document clustering: Tf-idf approach
Prafulla Bafna, Dhanya Pramod, and Anagha Vaidya · 2016
Earlier work this paper cites.
Herded gibbs sampling
Yutian Chen, Luke Bornn, Nando De Freitas, Mareija Eskelin, Jing Fang, and Max Welling · 2016
Earlier work this paper cites.
Hierarchical data topology based selection for large scale learning
Hmida Hmida, Sana Ben Hamida, Amel Borgi, and Marta Rukoz · 2016
Earlier work this paper cites.
How well do computers solve math word problems? large-scale dataset construction and evaluation
Danqing Huang, Shuming Shi, Chin-Yew Lin, Jian Yin, and Wei-Ying Ma · 2016
Earlier work this paper cites.
Mawps: A math word problem repository
Rik Koncel-Kedziorski, Subhro Roy, Aida Amini, Nate Kushman, and Hannaneh Hajishirzi · 2016
Earlier work this paper cites.
Combining latent dirichlet allocation and k-means for documents clustering: effect of probabilistic based distance measures
Quang Vu Bui, Karim Sayadi, Soufian Ben Amor, and Marc Bui · 2017
Earlier work this paper cites.
Semantics derived automatically from language corpora contain human-like biases
Aylin Caliskan, Joanna J Bryson, and Arvind Narayanan · 2017
Earlier work this paper cites.
Latent variable dialogue models and their diversity
Kris Cao and Stephen Clark · 2017
Earlier work this paper cites.
Understanding black-box predictions via influence functions
Pang Wei Koh and Percy Liang · 2017
Earlier work this paper cites.
On the state of the art of evaluation in neural language models
Gábor Melis, Chris Dyer, and Phil Blunsom · 2017
Earlier work this paper cites.
Active learning for convolutional neural networks: A core-set approach
Ozan Sener and Silvio Savarese · 2017
Earlier work this paper cites.
A review on bilevel optimization: From classical to evolutionary approaches and applications
Ankur Sinha, Pekka Malo, and Kalyanmoy Deb · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Datasets column: diversity and credibility for social images and image retrieval
Bogdan Ionescu, Mihai Lupu, Maia Rohm, Alexandru Lucian Gînsca, and Henning Müller · 2018
Earlier work this paper cites.
Regularization and the small-ball method i: sparse recovery
Guillaume Lecué and Shahar Mendelson · 2018
Cited alongside, same era.
Influence maximization on social graphs: A survey
Yuchen Li, Ju Fan, Yanhao Wang, and Kian-Lee Tan · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Cited alongside, same era.
Aditya Siddhant and Zachary C Lipton · 2018
Cited alongside, same era.
An empirical study of example forgetting during deep neural network learning
Mariya Toneva, Alessandro Sordoni, Remi Tachet des Combes, Adam Trischler, Yoshua Bengio, and Geoffrey J Gordon · 2018
Cited alongside, same era.
Quantifying uncertainty in answers from any language model and enhancing their trustworthiness
Jiuhai Chen and Jonas Mueller · 2023
Later among the works it cites.
Free dolly: Introducing the world’s first truly open instruction-tuned llm
Mike Conover, Matt Hayes, Ankit Mathur, Jianwei Xie, Jun Wan, Sam Shah, Ali Ghodsi, Patrick Wendell, Matei Zaharia, and Reynold Xin · 2023
Later among the works it cites.
The vendi score: A diversity evaluation metric for machine learning
Dan Dan Friedman and Adji Bousso Dieng · 2023
Later among the works it cites.
Investigating data contamination in modern benchmarks for large language models
Chunyuan Deng, Yilun Zhao, Xiangru Tang, Mark Gerstein, and Arman Cohan · 2023
Later among the works it cites.
Enhancing chat language models by scaling high-quality instructional conversations
Ning Ding, Yulin Chen, Bokai Xu, Yujia Qin, Zhi Zheng, Shengding Hu, Zhiyuan Liu, Maosong Sun, and Bowen Zhou · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Texygen: A benchmarking platform for text generation models
Yaoming Zhu, Sidi Lu, Lei Zheng, Jiaxian Guo, Weinan Zhang, Jun Wang, and Yong Yu · 2018
Cited alongside, same era.
Task2vec: Task embedding for meta-learning
Alessandro Achille, Michael Lam, Rahul Tewari, Avinash Ravichandran, Subhransu Maji, Charless C Fowlkes, Stefano Soatto, and Pietro Perona · 2019
Cited alongside, same era.
Analysis methods in neural language processing: A survey
Yonatan Belinkov and James Glass · 2019
Cited alongside, same era.
Identifying and reducing gender bias in word-level language models
Shikha Bordia and Samuel R Bowman · 2019
Cited alongside, same era.
The secret sharer: Evaluating and testing unintended memorization in neural networks
Nicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos, and Dawn Song · 2019
Cited alongside, same era.
Dbscan algorithm for document clustering
Radu G Creţulescu, Daniel I Morariu, Macarie Breazu, and Daniel Volovici · 2019
Cited alongside, same era.
Unified language model pre-training for natural language understanding and generation
Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon · 2019
Cited alongside, same era.
Later among the works it cites.
Measuring the robustness of ml models against data quality issues in industrial time series data
Marcel Dix, Gianluca Manca, Kenneth Chigozie Okafor, Reuben Borrison, Konstantin Kirchheim, Divyasheel Sharma, Kr Chandrika, Deepti Maduskar, and Frank Ortmeier · 2023
Later among the works it cites.
Raft: Reward ranked finetuning for generative foundation model alignment
Hanze Dong, Wei Xiong, Deepanshu Goyal, Yihan Zhang, Winnie Chow, Rui Pan, Shizhe Diao, Jipeng Zhang, Kashun Shum, and Tong Zhang · 2023
Later among the works it cites.
Mods: Model-oriented data selection for instruction tuning
Qianlong Du, Chengqing Zong, and Jiajun Zhang · 2023
Later among the works it cites.
Gio: Gradient information optimization for training dataset selection
Dante Everaert and Christopher Potts · 2023
Later among the works it cites.
Finetune like you pretrain: Improved finetuning of zero-shot vision models
Sachin Goyal, Ananya Kumar, Sankalp Garg, Zico Kolter, and Aditi Raghunathan · 2023
Later among the works it cites.
Studying large language model generalization with influence functions, 2023
Roger Grosse, Juhan Bae, Cem Anil, Nelson Elhage, Alex Tamkin, Amirhossein Tajdini, Benoit Steiner, Dustin Li, Esin Durmus, Ethan Perez, Evan Hubinger, Kamilė Lukošiūtė, Karina Nguyen, Nicholas Joseph, Sam McCandlish, Jared Kaplan, and Samuel R. Bowman · 2023
Later among the works it cites.
Suriya Gunasekar, Yi Zhang, Jyoti Aneja, Caio César Teodoro Mendes, Allie Del Giorno, Sivakanth Gopi, Mojan Javaheripi, Piero Kauffmann, Gustavo de Rosa, Olli Saarikivi, et al · 2023
Later among the works it cites.
Simfluence: Modeling the influence of individual training examples by simulating training runs
Kelvin Guu, Albert Webson, Ellie Pavlick, Lucas Dixon, Ian Tenney, and Tolga Bolukbasi · 2023
Later among the works it cites.
K-means clustering algorithms: A comprehensive review, variants analysis, and advances in the era of big data
Abiodun M Ikotun, Absalom E Ezugwu, Laith Abualigah, Belal Abuhaija, and Jia Heming · 2023
Later among the works it cites.
A data-based perspective on transfer learning
Saachi Jain, Hadi Salman, Alaa Khaddaj, Eric Wong, Sung Min Park, and Aleksander Mądry · 2023
Later among the works it cites.
Exploring the benefits of training expert language models over instruction tuning
Joel Jang, Seungone Kim, Seonghyeon Ye, Doyoung Kim, Lajanugen Logeswaran, Moontae Lee, Kyungjae Lee, and Minjoon Seo · 2023
Later among the works it cites.
Delving into effective gradient matching for dataset condensation
Zixuan Jiang, Jiaqi Gu, Mingjie Liu, and David Z Pan · 2023
Later among the works it cites.
Does" deep learning on a data diet" reproduce? overall yes, but grand at initialization does not
Andreas Kirsch · 2023
Later among the works it cites.
GPT-3: The Ultimate Guide to Building NLP Products with OpenAI API
Sandra Kublik and Shubham Saboo · 2023
Later among the works it cites.
Do models really learn to follow instructions? an empirical study of instruction tuning
Po-Nien Kung and Nanyun Peng · 2023
Later among the works it cites.
Active instruction tuning: Improving cross-task generalization by training on prompt sensitive tasks
Po-Nien Kung, Fan Yin, Di Wu, Kai-Wei Chang, and Nanyun Peng · 2023
Later among the works it cites.
A fully first-order method for stochastic bilevel optimization
Jeongyeol Kwon, Dohyun Kwon, Stephen Wright, and Robert D Nowak · 2023
Later among the works it cites.
Aisha: A custom ai library chatbot using the chatgpt api
Yrjo Lappalainen and Nikesh Narayanan · 2023
Later among the works it cites.
Alycia Lee, Brando Miranda, Sudharsan Sundar, and Sanmi Koyejo · 2023
Later among the works it cites.
The flan collection: Designing data and methods for effective instruction tuning
Shayne Longpre, Le Hou, Tu Vu, Albert Webson, Hyung Won Chung, Yi Tay, Denny Zhou, Quoc V Le, Barret Zoph, Jason Wei, et al · 2023
Later among the works it cites.
Understanding and mitigating overfitting in prompt tuning for vision-language models
Chengcheng Ma, Yang Liu, Jiankang Deng, Lingxi Xie, Weiming Dong, and Changsheng Xu · 2023
Later among the works it cites.
When less is more: Investigating data pruning for pretraining llms at scale
Max Marion, Ahmet Üstün, Luiza Pozzobon, Alex Wang, Marzieh Fadaee, and Sara Hooker · 2023
Later among the works it cites.
Trak: Attributing model behavior at scale
Sung Min Park, Kristian Georgiev, Andrew Ilyas, Guillaume Leclerc, and Aleksander Madry · 2023
Later among the works it cites.
Amey Pasarkar and Adji Bousso Dieng · 2023
Later among the works it cites.
Openwebmath: An open dataset of high-quality mathematical web text
Keiran Paster, Marco Dos Santos, Zhangir Azerbayev, and Jimmy Ba · 2023
Later among the works it cites.
Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay · 2023
Later among the works it cites.
Baolin Peng, Chunyuan Li, Pengcheng He, Michel Galley, and Jianfeng Gao · 2023
Later among the works it cites.
A survey of data quality requirements that matter in ml development pipelines
Maria Priestley, Fionntán O’donnell, and Elena Simperl · 2023
Later among the works it cites.
Comprehensive survey on hierarchical clustering algorithms and the recent developments
Xingcheng Ran, Yue Xi, Yonggang Lu, Xiangwen Wang, and Zhenyu Lu · 2023
Later among the works it cites.
Data selection for fine-tuning large language models using transferred shapley values
Stephanie Schoch, Ritwick Mishra, and Yangfeng Ji · 2023
Later among the works it cites.
Don’t stop pretraining? make prompt-based fine-tuning powerful learner
Zhengxiang Shi and Aldo Lipani · 2023
Later among the works it cites.
On the exploitability of instruction tuning
Manli Shu, Jiongxiao Wang, Chen Zhu, Jonas Geiping, Chaowei Xiao, and Tom Goldstein · 2023
Later among the works it cites.
Does fine-tuning gpt-3 with the openai api leak personally-identifiable information?
Albert Yu Sun, Eliott Zemour, Arushi Saxena, Udith Vaidyanathan, Eric Lin, Christian Lau, and Vaikkunth Mugunthan · 2023
Later among the works it cites.
Extracting representative subset from extensive text data for training pre-trained language models
Jun Suzuki, Heiga Zen, and Hideto Kazawa · 2023
Later among the works it cites.
Alpaca: A strong, replicable instruction-following model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto · 2023
Later among the works it cites.
Introducing mpt-7b: A new standard for open-source, commercially usable llms, 2023
MosaicML NLP Team · 2023
Later among the works it cites.
Self-evolved diverse data sampling for efficient instruction tuning
Shengguang Wu, Keming Lu, Benfeng Xu, Junyang Lin, Qi Su, and Chang Zhou · 2023
Later among the works it cites.
Sheared llama: Accelerating language model pre-training via structured pruning
Mengzhou Xia, Tianyu Gao, Zhiyuan Zeng, and Danqi Chen · 2023
Later among the works it cites.
Data selection for language models via importance resampling
Sang Michael Xie, Shibani Santurkar, Tengyu Ma, and Percy S Liang · 2023
Later among the works it cites.
Data similarity is not enough to explain language model performance
Gregory Yauney, Emily Reif, and David Mimno · 2023
Later among the works it cites.
Investigating openai’s chatgpt potentials in generating chatbot’s dialogue for english as a foreign language learning
Julio Christian Young and Makoto Shishido · 2023
Later among the works it cites.
Evaluating large language models at evaluating instruction following
Zhiyuan Zeng, Jiatong Yu, Tianyu Gao, Yu Meng, Tanya Goyal, and Danqi Chen · 2023
Later among the works it cites.
Data-centric artificial intelligence: A survey
Daochen Zha, Zaid Pervaiz Bhat, Kwei-Herng Lai, Fan Yang, Zhimeng Jiang, Shaochen Zhong, and Xia Hu · 2023
Later among the works it cites.
Dataset condensation with distribution matching
Bo Zhao and Hakan Bilen · 2023
Later among the works it cites.
Dataset quantization
Daquan Zhou, Kai Wang, Jianyang Gu, Xiangyu Peng, Dongze Lian, Yifan Zhang, Yang You, and Jiashi Feng · 2023
Later among the works it cites.
Judgelm: Fine-tuned large language models are scalable judges
Lianghui Zhu, Xinggang Wang, and Xinlong Wang · 2023
Later among the works it cites.
Llama 3 model card
AI@Meta · 2024
Closest in time.
A survey on data selection for language models
Alon Albalak, Yanai Elazar, Sang Michael Xie, Shayne Longpre, Nathan Lambert, Xinyi Wang, Niklas Muennighoff, Bairu Hou, Liangming Pan, Haewon Jeong, et al · 2024
Closest in time.
Alexandre Alcoforado, Thomas Palmeira Ferraz, Lucas Hideki Okamura, Israel Campos Fama, Arnold Moya Lavado, Bárbara Dias Bueno, Bruno Veloso, and Anna Helena Reali Costa · 2024
Closest in time.
Perplexed by perplexity: Perplexity-based data pruning with small reference models
Zachary Ankner, Cody Blakeney, Kartik Sreenivasan, Max Marion, Matthew L Leavitt, and Mansheej Paul · 2024
Closest in time.
Generalization vs memorization: Tracing language models’ capabilities back to pretraining data
Antonis Antoniades, Xinyi Wang, Yanai Elazar, Alfonso Amayuelas, Alon Albalak, Kexun Zhang, and William Yang Wang · 2024
Closest in time.
Proxylm: Predicting language model performance on multilingual tasks via proxy models
David Anugraha, Genta Indra Winata, Chenyue Li, Patrick Amadeus Irawan, and En-Shiun Annie Lee · 2024
Closest in time.
Data-efficient learning via clustering-based sensitivity sampling: Foundation models and beyond
Kyriakos Axiotis, Vincent Cohen-Addad, Monika Henzinger, Sammy Jerome, Vahab Mirrokni, David Saulpic, David Woodruff, and Michael Wunder · 2024
Closest in time.
Training data attribution via approximate unrolled differentation
Juhan Bae, Wu Lin, Jonathan Lorraine, and Roger Grosse · 2024
Closest in time.
An experimental design framework for label-efficient supervised finetuning of large language models
Gantavya Bhatt, Yifang Chen, Arnav M Das, Jifan Zhang, Sang T Truong, Stephen Mussmann, Yinglun Zhu, Jeffrey Bilmes, Simon S Du, Kevin Jamieson, et al · 2024
Closest in time.
Emergent and predictable memorization in large language models
Stella Biderman, Usvsn Prashanth, Lintang Sutawika, Hailey Schoelkopf, Quentin Anthony, Shivanshu Purohit, and Edward Raff · 2024
Closest in time.
Index1.9b technical report
Bilibili · 2024
Closest in time.
Data summarization via bilevel optimization
Zalán Borsos, Mojmír Mutnỳ, Marco Tagliasacchi, and Andreas Krause · 2024
Closest in time.
Concerned with data contamination? assessing countermeasures in code language model
Jialun Cao, Wuqi Zhang, and Shing-Chi Cheung · 2024
Closest in time.
A survey on evaluation of large language models
Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al · 2024
Closest in time.
Data-juicer: A one-stop data processing system for large language models
Daoyuan Chen, Yilun Huang, Zhijian Ma, Hesen Chen, Xuchen Pan, Ce Ge, Dawei Gao, Yuexiang Xie, Zhaoyang Liu, Jinyang Gao, et al · 2024
Closest in time.
Automated data curation for robust language model fine-tuning
Jiuhai Chen and Jonas Mueller · 2024
Closest in time.
Instruction pre-training: Language models are supervised multitask learners
Daixuan Cheng, Yuxian Gu, Shaohan Huang, Junyu Bi, Minlie Huang, and Furu Wei · 2024
Closest in time.
" what data benefits my classifier?" enhancing model performance and interpretability through influence-based data selection
Anshuman Chhabra, Peizhao Li, Prasant Mohapatra, and Hongfu Liu · 2024
Closest in time.
Prompting fairness: Learning prompts for debiasing large language models
Andrei-Victor Chisca, Andrei-Cristian Rad, and Camelia Lemnaru · 2024
Closest in time.
Fairness in large language models: A taxonomic survey
Zhibo Chu, Zichong Wang, and Wenbin Zhang · 2024
Closest in time.
Continual pre-training mitigates forgetting in language and vision
Andrea Cossu, Antonio Carta, Lucia Passaro, Vincenzo Lomonaco, Tinne Tuytelaars, and Davide Bacciu · 2024
Closest in time.
Scaling laws for the value of individual data points in machine learning
Ian Covert, Wenlong Ji, Tatsunori Hashimoto, and James Zou · 2024
Closest in time.
Data quality in nlp: Metrics and a comprehensive taxonomy
Vu Minh Hoang Dang and Rakesh M Verma · 2024
Closest in time.
Fairness in large language models in three hours
Thang Viet Doan, Zichong Wang, Nhat Nguyen Minh Hoang, and Wenbin Zhang · 2024
Closest in time.
Guanting Dong, Keming Lu, Chengpeng Li, Tingyu Xia, Bowen Yu, Chang Zhou, and Jingren Zhou · 2024
Closest in time.
Sequential subset matching for dataset distillation
Jiawei Du, Qin Shi, and Joey Tianyi Zhou · 2024
Closest in time.
Alpacafarm: A simulation framework for methods that learn from human feedback
Yann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy S Liang, and Tatsunori B Hashimoto · 2024
Closest in time.
Dsdm: Model-aware dataset selection with datamodels
Logan Engstrom, Axel Feldmann, and Aleksander Madry · 2024
Closest in time.
Bias and fairness in large language models: A survey
Isabel O Gallegos, Ryan A Rossi, Joe Barrow, Md Mehrab Tanjim, Sungchul Kim, Franck Dernoncourt, Tong Yu, Ruiyi Zhang, and Nesreen K Ahmed · 2024
Closest in time.
Fairness of large language models in education
Rong Gao, Qin Ni, and Bingying Hu · 2024
Closest in time.
A closer look at the limitations of instruction tuning
Sreyan Ghosh, Chandra Kiran Reddy Evuru, Sonal Kumar, Deepali Aneja, Zeyu Jin, Ramani Duraiswami, Dinesh Manocha, et al · 2024
Closest in time.
How to detect bad data in your instruction tuning dataset (for better llm fine-tuning), 2024
Jimming He, Sanjana Garg, and Jonas Mueller · 2024
Closest in time.
Hui Huang, Yingqi Qu, Jing Liu, Muyun Yang, and Tiejun Zhao · 2024
Closest in time.
Simple and scalable strategies to continually pre-train large language models
Adam Ibrahim, Benjamin Thérien, Kshitij Gupta, Mats L Richter, Quentin Anthony, Timothée Lesort, Eugene Belilovsky, and Irina Rish · 2024
Closest in time.
Instructed to bias: Instruction-tuned language models exhibit emergent cognitive bias
Itay Itzhak, Gabriel Stanovsky, Nir Rosenfeld, and Yonatan Belinkov · 2024
Closest in time.
Enhancing fairness in llm evaluations: Unveiling and mitigating biases in standard-answer-based evaluations
Tong Jiao, Jian Zhang, Kui Xu, Rui Li, Xi Du, Shangqi Wang, and Zhenbo Song · 2024
Closest in time.
Performance scaling via optimal transport: Enabling data selection from partially revealed sources
Feiyang Kang, Hoang Anh Just, Anit Kumar Sahu, and Ruoxi Jia · 2024
Closest in time.
Continual Learning with Language Models
Zixuan Ke · 2024
Closest in time.
Lmd3: Language model data density dependence
John Kirchenbauer, Garrett Honke, Gowthami Somepalli, Jonas Geiping, Daphne Ippolito, Katherine Lee, Tom Goldstein, and David Andre · 2024
Closest in time.
Openassistant conversations-democratizing large language model alignment
Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotiris Anagnostidis, Zhi Rui Tam, Keith Stevens, Abdullah Barhoum, Duc Nguyen, Oliver Stanley, Richárd Nagyfi, et al · 2024
Closest in time.
Data-efficient fine-tuning for llm-based recommendation, 2024
Xinyu Lin, Wenjie Wang, Yongqi Li, Shuo Yang, Fuli Feng, Yinwei Wei, and Tat-Seng Chua · 2024
Closest in time.
Large language model instruction following: A survey of progresses and challenges
Renze Lou, Kai Zhang, and Wenpeng Yin · 2024
Closest in time.
First-order penalty methods for bilevel optimization
Zhaosong Lu and Sanyou Mei · 2024
Closest in time.
Codeact: Code adaptive compute-efficient tuning framework for code llms
Weijie Lv, Xuan Xia, and Sheng-Jun Huang · 2024
Closest in time.
Data portraits: Recording foundation model training data
Marc Marone and Benjamin Van Durme · 2024
Closest in time.
Data quality assessment: Challenges and opportunities
Sedir Mohammed, Hazar Harmouch, Felix Naumann, and Divesh Srivastava · 2024
Closest in time.
Aidar Myrzakhan, Sondos Mahmoud Bsharat, and Zhiqiang Shen · 2024
Closest in time.
Diversity measures: Domain-independent proxies for failure in language model queries
Noel Ngu, Nathaniel Lee, and Paulo Shakarian · 2024
Closest in time.
Quality-weighted vendi scores and their application to diverse experimental design
Quan Nguyen and Adji Bousso Dieng · 2024
Closest in time.
Large-scale dataset pruning in adversarial training through data importance extrapolation
Björn Nieth, Thomas Altstidl, Leo Schwinn, and Björn Eskofier · 2024
Closest in time.
Scalebio: Scalable bilevel optimization for llm data reweighting
Rui Pan, Jipeng Zhang, Xingyuan Pan, Renjie Pi, Xiaoyu Wang, and Tong Zhang · 2024
Closest in time.
Improving data efficiency via curating llm-driven rating systems
Jinlong Pang, Jiaheng Wei, Ankit Parag Shah, Zhaowei Zhu, Yaxuan Wang, Chen Qian, Yang Liu, Yujia Bao, and Wei Wei · 2024
Closest in time.
Reuse, don’t retrain: A recipe for continued pretraining of language models
Jupinder Parmar, Sanjev Satheesh, Mostofa Patwary, Mohammad Shoeybi, and Bryan Catanzaro · 2024
Closest in time.
Influenciæ: A library for tracing the influence back to the data-points
Agustin Picard, Lucas Hervier, Thomas Fel, and David Vigouroux · 2024
Closest in time.
Capro: webly supervised learning with cross-modality aligned prototypes
Yulei Qin, Xingyu Chen, Yunhang Shen, Chaoyou Fu, Yun Gu, Ke Li, Xing Sun, and Rongrong Ji · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2024
Closest in time.
Dele: Data efficient llm evaluation
Gayathri Saranathan, Mahammad Parwez Alam, James Lim, Suparna Bhattacharya, Soon Yee Wong, Martin Foltin, and Cong Xu · 2024
Closest in time.
Balanced data sampling for language model training with clustering
Yunfan Shao, Linyang Li, Zhaoye Fei, Hang Yan, Dahua Lin, and Xipeng Qiu · 2024
Closest in time.
Instruction tuning with loss over instructions
Zhengyan Shi, Adam X Yang, Bin Wu, Laurence Aitchison, Emine Yilmaz, and Aldo Lipani · 2024
Closest in time.
D4: Improving llm pretraining via document de-duplication and diversification
Kushal Tirumala, Daniel Simig, Armen Aghajanyan, and Ari Morcos · 2024
Closest in time.
Qurating: Selecting high-quality data for training language models
Alexander Wettig, Aatmik Gupta, Saumya Malik, and Danqi Chen · 2024
Closest in time.
Result diversification in search and recommendation: A survey
Haolun Wu, Yansen Zhang, Chen Ma, Fuyuan Lyu, Bowei He, Bhaskar Mitra, and Xue Liu · 2024
Closest in time.
Magpie: Alignment data synthesis from scratch by prompting aligned llms with nothing
Zhangchen Xu, Fengqing Jiang, Luyao Niu, Yuntian Deng, Radha Poovendran, Yejin Choi, and Bill Yuchen Lin · 2024
Closest in time.
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, et al · 2024
Closest in time.
Distilled datamodel with reverse gradient matching
Jingwen Ye, Ruonan Yu, Songhua Liu, and Xinchao Wang · 2024
Closest in time.
Investigating continual pretraining in large language models: Insights and implications
Çağatay Yıldız, Nishaanth Kanna Ravichandran, Prishruit Punia, Matthias Bethge, and Beyza Ermis · 2024
Closest in time.
Mates: Model-aware data selection for efficient pretraining with data influence models
Zichun Yu, Spandan Das, and Chenyan Xiong · 2024
Closest in time.
Bilevel optimization in the deep learning era: Methods and applications
Lei Zhang · 2024
Closest in time.
Dissecting learning and forgetting in language model finetuning
Xiao Zhang and Ji Wu · 2024
Closest in time.
Autonomous data selection with language models for mathematical texts
Yifan Zhang, Yifan Luo, Yang Yuan, and Andrew C Yao · 2024
Closest in time.
Technical report: Competition solution for bettermixture
Shuaijiang Zhao and Xiaoquan Fang · 2024
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2024
Closest in time.
Xinlin Zhuang, Xin Mao, Yuan-Hao Jiang, Hongyi Wu, Shangqing Zhao, Li Cai, Shu Liu, Yang Chen, Yuxiang Song, Chenghao Jia, et al · 2024
Closest in time.