Fetching the paper…
Reading the bibliography…
It is often advantageous to train models on a subset of the available train examples, because the examples are of variable quality or because one would like to train with fewer examples, without sacrificing performance.
On information and sufficiency
Solomon Kullback and Richard Liebler · 1951
Earlier work this paper cites.
Prior probabilities
Edwin T. Jaynes · 1968
Earlier work this paper cites.
Particle swarm optimization
J. Kennedy and R. Eberhart · 1995
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
k-means++: the advantages of careful seeding
David Arthur and Sergei Vassilvitskii · 2007
Earlier work this paper cites.
Divergence estimation for multidimensional densities via k k -nearest-neighbor distances
Qing Wang, Sanjeev R. Kulkarni, and Sergio Verdu · 2009
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2016
Earlier work this paper cites.
Annoy (approximate nearest neighbors oh yeah), 2017
Erik Bernhardsson · 2017
Earlier work this paper cites.
Deep bayesian active learning with image data
Yarin Gal, Riashat Islam, and Zoubin Ghahramani · 2017
Earlier work this paper cites.
Billion-scale similarity search with gpus
Jeff Johnson, Matthijs Douze, and Herve Jegou · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Spelling correction as a foreign language
Han Xiao, Kashif Rasul, and Roland Vollgraf · 2017
Earlier work this paper cites.
Understanding back-translation at scale
Sergey Edunov, Myle Ott, Michael Auli, and David Grangier · 2018
Earlier work this paper cites.
Active learning for convolutional neural networks: A core-set approach
Ozan Sener and Silvio Savarese · 2018
Cited alongside, same era.
The BEA-2019 shared task on grammatical error correction
Christopher Bryant, Mariano Felice, Øistein E. Andersen, and Ted Briscoe · 2019
Cited alongside, same era.
Batchbald: Efficient and diverse batch acquisition for deep bayesian active learning
Andreas Kirsch, Joost van Amersfoort, and Yarin Gal · 2019
Cited alongside, same era.
compound_split_bleu.sh, 2019
Myle Ott · 2019
Cited alongside, same era.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli · 2019
Cited alongside, same era.
Fashion mnist with pytorch (93% accuracy), 2019
Pankaj · 2019
Cited alongside, same era.
Scalable and generalizable social bot detection through data selection
Kai-Cheng Yang, Onur Varol, Pik-Mai Hui, and Filippo Menczer · 2020
Later among the works it cites.
Neuspell: A neural spelling correction toolkit, 2021
Sai Muralidhar Jayanthi · 2021
Later among the works it cites.
The perils of using Mechanical Turk to evaluate open-ended text generation
Marzena Karpinska, Nader Akoury, and Mohit Iyyer · 2021
Later among the works it cites.
What’s in the box? an analysis of undesirable content in the Common Crawl corpus
Alexandra Luccioni and Joseph Viviano · 2021
Later among the works it cites.
Deep learning on a data diet: Finding important examples early in training
Mansheej Paul, Surya Ganguli, and Gintare Karolina Dziugaite · 2021
Later among the works it cites.
Subword neural machine translation, 2021
Rico Sennrich · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Impact of data pruning on machine learning algorithm performance, 2019
Arun Thundyill Saseendran, Lovish Setia, Viren Chhabria, Debrup Chakraborty, and Aneek Barman Roy · 2019
Cited alongside, same era.
Spelling correction as a foreign language
Yingbo Zhou, Utkarsh Porwal, and Roberto Konow · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Cited alongside, same era.
Using resnet for fashion mnist in pytorch, 2020
Amithash K J · 2020
Cited alongside, same era.
NeuSpell: A neural spelling correction toolkit
Sai Muralidhar Jayanthi, Danish Pruthi, and Graham Neubig · 2020
Cited alongside, same era.
BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer · 2020
Cited alongside, same era.
Later among the works it cites.
Training data-efficient image transformers and distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herve Jegou · 2021
Later among the works it cites.
Nlp from scratch without large-scale pretraining, 2021
Xingcheng Yao and Zongmeng Zhang · 2021
Later among the works it cites.
Fairseq, 2022
FAIR · 2022
Later among the works it cites.
Submodlib: A submodular optimization library, 2022
Vishal Kaushal, Ganesh Ramakrishnan, and Rishabh Iyer · 2022
Later among the works it cites.
ColBERTv2: Effective and efficient retrieval via lightweight late interaction
Keshav Santhanam, Omar Khattab, Jon Saad-Falcon, Christopher Potts, and Matei Zaharia · 2022
Later among the works it cites.
Beyond neural scaling laws: beating power law scaling via data pruning
Ben Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli, and Ari Morcos · 2022
Later among the works it cites.
NLP from scratch without large-scale pretraining: A simple and efficient framework
Xingcheng Yao, Yanan Zheng, Xiaocong Yang, and Zhilin Yang · 2022
Later among the works it cites.
Data selection for language models via importance resampling
Sang Michael Xie, Shibani Santurkar, Tengyu Ma, and Percy Liang · 2023
Closest in time.
Dataset pruning: Reducing training data by examining generalization influence, 2023
Shuo Yang, Zeke Xie, Hanyu Peng, Min Xu, Mingming Sun, and Ping Li · 2023
Closest in time.