Fetching the paper…
Reading the bibliography…
Selecting high-quality pre-training data is important for creating capable language models, but existing methods rely on simple heuristics.
A law of comparative judgment
Thurstone, L. L · 1927
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
Could comparative judgements of script quality replace traditional marking and improve the validity of exam questions
Pollitt, A. and Crisp, V · 2004
Earlier work this paper cites.
Spam filtering with naive bayes-which naive bayes?
Metsis, V., Androutsopoulos, I., and Paliouras, G · 2006
Earlier work this paper cites.
Reinforcement learning by reward-weighted regression for operational space control
Peters, J. and Schaal, S · 2007
Earlier work this paper cites.
SMS Spam Collection
Almeida, T. and Hidalgo, J · 2012
Earlier work this paper cites.
The method of adaptive comparative judgement
Pollitt, A · 2012
Earlier work this paper cites.
Genome editing. the new frontier of genome engineering with crispr-cas9
Doudna, J. A. and Charpentier, E · 2014
Earlier work this paper cites.
Gumbel-max trick and weighted reservoir sampling, 2014
Vieira, T · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G. E., Vinyals, O., and Dean, J · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. and Ba, J · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Zhang, X., Zhao, J., and LeCun, Y · 2015
Earlier work this paper cites.
Exact sampling with integer linear programs and random perturbations
Kim, C., Sabharwal, A., and Ermon, S · 2016
Earlier work this paper cites.
Sequence-level knowledge distillation
Kim, Y. and Rush, A. M · 2016
Earlier work this paper cites.
The problem with bias: Allocative versus representational harms in machine learning
Barocas, S., Crawford, K., Shapiro, A., and Wallach, H · 2017
Earlier work this paper cites.
Semantics derived automatically from language corpora contain human-like biases
Caliskan, A., Bryson, J. J., and Narayanan, A · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I · 2017
Earlier work this paper cites.
Crowdsourcing multiple choice science questions
Welbl, J., Liu, N. F., and Gardner, M · 2017
Earlier work this paper cites.
Think you have solved question answering? Try ARC, the AI2 reasoning challenge
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
BoolQ: Exploring the surprising difficulty of natural yes/no questions
Clark, C., Lee, K., Chang, M.-W., Kwiatkowski, T., Collins, M., and Toutanova, K · 2019
Earlier work this paper cites.
Stochastic beams and where to find them: The Gumbel-top-k trick for sampling sequences without replacement
Kool, W., Van Hoof, H., and Welling, M · 2019
Earlier work this paper cites.
Natural questions: A benchmark for question answering research
Kwiatkowski, T., Palomaki, J., Redfield, O., Collins, M., Parikh, A., Alberti, C., Epstein, D., Polosukhin, I., Devlin, J., Lee, K., Toutanova, K., Jones, L., Kelcey, M., Chang, M.-W., Dai, A. M., Uszkoreit, J., Le, Q., and Petrov, S · 2019
Earlier work this paper cites.
Quantifying the carbon emissions of machine learning
Lacoste, A., Luccioni, A., Schmidt, V., and Dandres, T · 2019
Earlier work this paper cites.
Black is to criminal as Caucasian is to police: Detecting and removing multiclass bias in word embeddings
Manzini, T., Yao Chong, L., Black, A. W., and Tsvetkov, Y · 2019
Earlier work this paper cites.
Language models as knowledge bases?
Petroni, F., Rocktäschel, T., Riedel, S., Lewis, P., Bakhtin, A., Wu, Y., and Miller, A · 2019
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Sanh, V., Debut, L., Chaumond, J., and Wolf, T · 2019
Earlier work this paper cites.
Energy and policy considerations for deep learning in NLP
Strubell, E., Ganesh, A., and McCallum, A · 2019
Earlier work this paper cites.
Assessing social and intersectional biases in contextualized word representations
Tan, Y. C. and Celis, L. E · 2019
Earlier work this paper cites.
HellaSwag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Bisk, Y., Zellers, R., Gao, J., Choi, Y., et al · 2020
Earlier work this paper cites.
Language (technology) is power: A critical survey of “bias” in NLP
Blodgett, S. L., Barocas, S., Daumé III, H., and Wallach, H · 2020
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Cited alongside, same era.
The Pile: An 800GB dataset of diverse text for language modeling
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., et al · 2020
Cited alongside, same era.
RealToxicityPrompts: Evaluating neural toxic degeneration in language models
Gehman, S., Gururangan, S., Sap, M., Choi, Y., and Smith, N. A · 2020
Cited alongside, same era.
Logiqa: A challenge dataset for machine reading comprehension with logical reasoning
Liu, J., Cui, L., Liu, H., Huang, D., Wang, Y., and Zhang, Y · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
SemDeDup: Data-efficient learning at web-scale through semantic deduplication
Abbas, A., Tirumala, K., Simig, D., Ganguli, S., and Morcos, A. S · 2023
Later among the works it cites.
Anil, R., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., Passos, A., Shakeri, S., Taropa, E., Bailey, P., Chen, Z., et al · 2023
Later among the works it cites.
A theory for emergence of complex skills in language models
Arora, S. and Goyal, A · 2023
Later among the works it cites.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2023
Later among the works it cites.
Alpacafarm: A simulation framework for methods that learn from human feedback
Dubois, Y., Li, X., Taori, R., Zhang, T., Gulrajani, I., Ba, J., Guestrin, C., Liang, P., and Hashimoto, T · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Cited alongside, same era.
GLU variants improve transformer
Shazeer, N. M · 2020
Cited alongside, same era.
CCNet: Extracting high quality monolingual datasets from web crawl data
Wenzek, G., Lachaux, M.-A., Conneau, A., Chaudhary, V., Guzmán, F., Joulin, A., and Grave, E · 2020
Cited alongside, same era.
Persistent anti-muslim bias in large language models
Abid, A., Farooqi, M., and Zou, J · 2021
Cited alongside, same era.
On the dangers of stochastic parrots: Can language models be too big?
Bender, E. M., Gebru, T., McMillan-Major, A., and Shmitchell, S · 2021
Cited alongside, same era.
Harms of gender exclusivity and challenges in non-binary representation in language technologies
Dev, S., Monajatipoor, M., Ovalle, A., Subramonian, A., Phillips, J., and Chang, K.-W · 2021
Cited alongside, same era.
Documenting large webtext corpora: A case study on the colossal clean crawled corpus
Dodge, J., Sap, M., Marasović, A., Agnew, W., Ilharco, G., Groeneveld, D., Mitchell, M., and Gardner, M · 2021
Cited alongside, same era.
A framework for few-shot language model evaluation, September 2021
Gao, L., Tow, J., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., McDonell, K., Muennighoff, N., Phang, J., Reynolds, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A · 2021
Cited alongside, same era.
Later among the works it cites.
Aligning language models with preferences through f f -divergence minimization
Go, D., Korbak, T., Kruszewski, G., Rozen, J., Ryu, N., and Dymetman, M · 2023
Later among the works it cites.
Textbooks are all you need, 2023
Gunasekar, S., Zhang, Y., Aneja, J., Mendes, C. C. T., Giorno, A. D., Gopi, S., Javaheripi, M., Kauffmann, P., de Rosa, G., Saarikivi, O., Salim, A., Shah, S., Behl, H. S., Wang, X., Bubeck, S., Eldan, R., Kalai, A. T., Lee, Y. T., and Li, Y · 2023
Later among the works it cites.
Scaling expert language models with unsupervised domain discovery
Gururangan, S., Li, M., Lewis, M., Shi, W., Althoff, T., Smith, N. A., and Zettlemoyer, L · 2023
Later among the works it cites.
Pretraining language models with human preferences
Korbak, T., Shi, K., Chen, A., Bhalerao, R. V., Buckley, C., Phang, J., Bowman, S. R., and Perez, E · 2023
Later among the works it cites.
Textbooks are all you need ii: phi-1.5 technical report, 2023
Li, Y., Bubeck, S., Eldan, R., Giorno, A. D., Gunasekar, S., and Lee, Y. T · 2023
Later among the works it cites.
When less is more: Investigating data pruning for pretraining LLMs at scale, 2023
Marion, M., Üstün, A., Pozzobon, L., Wang, A., Fadaee, M., and Hooker, S · 2023
Later among the works it cites.
Scaling data-constrained language models
Muennighoff, N., Rush, A. M., Barak, B., Scao, T. L., Tazi, N., Piktus, A., Pyysalo, S., Wolf, T., and Raffel, C · 2023
Later among the works it cites.
The refinedweb dataset for falcon LLM: Outperforming curated corpora with web data only
Penedo, G., Malartic, Q., Hesslow, D., Cojocaru, R., Alobeidli, H., Cappelli, A., Pannier, B., Almazrouei, E., and Launay, J · 2023
Later among the works it cites.
Large language models are effective text rankers with pairwise ranking prompting, 2023
Qin, Z., Jagerman, R., Hui, K., Zhuang, H., Wu, J., Shen, J., Liu, T., Liu, J., Metzler, D., Wang, X., and Bendersky, M · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2023
Later among the works it cites.
SlimPajama: A 627B token cleaned and deduplicated version of RedPajama
Soboleva, D., Al-Khateeb, F., Myers, R., Steeves, J. R., Hestness, J., and Dey, N · 2023
Later among the works it cites.
Dolma: An Open Corpus of Three Trillion Tokens for Language Model Pretraining Research
Soldaini, L., Kinney, R., Bhagia, A., Schwenk, D., Atkinson, D., Authur, R., Bogin, B., Chandu, K., Dumas, J., Elazar, Y., Hofmann, V., Jha, A. H., Kumar, S., Lucy, L., Lyu, X., Magnusson, I., Morrison, J., Muennighoff, N., Naik, A., Nam, C., Peters, M. E., Ravichander, A., Richardson, K., Shen, Z., Strubell, E., Subramani, N., Tafjord, O., Walsh, E. P., Hajishirzi, H., Smith, N. A., Zettlemoyer, L., Beltagy, I., Groeneveld, D., Dodge, J., and Lo, K · 2023
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding
Su, J., Ahmed, M., Lu, Y., Pan, S., Bo, W., and Liu, Y · 2023
Later among the works it cites.
Is ChatGPT good at search? investigating large language models as re-ranking agents
Sun, W., Yan, L., Ma, X., Wang, S., Ren, P., Chen, Z., Yin, D., and Ren, Z · 2023
Later among the works it cites.
D4: Improving LLM pretraining via document de-duplication and diversification
Tirumala, K., Simig, D., Aghajanyan, A., and Morcos, A · 2023
Later among the works it cites.
RedPajama: An open source recipe to reproduce llama training dataset, 2023
TogetherAI · 2023
Later among the works it cites.
Large language models are not fair evaluators
Wang, P., Li, L., Chen, L., Zhu, D., Lin, B., Cao, Y., Liu, Q., Liu, T., and Sui, Z · 2023
Later among the works it cites.
Training trajectories of language models across scales
Xia, M., Artetxe, M., Zhou, C., Lin, X. V., Pasunuru, R., Chen, D., Zettlemoyer, L., and Stoyanov, V · 2023
Later among the works it cites.
On-policy distillation of language models: Learning from self-generated mistakes
Agarwal, R., Vieillard, N., Zhou, Y., Stanczyk, P., Garea, S. R., Geist, M., and Bachem, O · 2024
Closest in time.
Understanding emergent abilities of language models from the loss perspective
Du, Z., Zeng, A., Dong, Y., and Tang, J · 2024
Closest in time.
Lucy, L., Gururangan, S., Soldaini, L., Strubell, E., Bamman, D., Klein, L., and Dodge, J · 2024
Closest in time.
A war of flags between guyana and venezuela
Ramalho, M · 2024
Closest in time.
Sheared LLaMA: Accelerating language model pre-training via structured pruning
Xia, M., Gao, T., Zeng, Z., and Chen, D · 2024
Closest in time.
Skill-mix: a flexible and expandable family of evaluations for AI models
Yu, D., Kaur, S., Gupta, A., Brown-Cohen, J., Goyal, A., and Arora, S · 2024
Closest in time.
Evaluating large language models at evaluating instruction following
Zeng, Z., Yu, J., Gao, T., Meng, Y., Goyal, T., and Chen, D · 2024
Closest in time.