Fetching the paper…
Reading the bibliography…
Training deep networks and tuning hyperparameters on large datasets is computationally intensive.
Semantic redundancies in image-classification datasets: The 10% you don’t need
Vighnesh Birodkar, Hossein Mobahi, and Samy Bengio · 1901
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych · 1908
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing, 2019
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush · 1910
Earlier work this paper cites.
An analysis of approximations for maximizing submodular set functions—i
George L Nemhauser, Laurence A Wolsey, and Marshall L Fisher · 1978
Earlier work this paper cites.
Toward semantics-based answer pinpointing
Eduard Hovy, Laurie Gerber, Ulf Hermjakob, Chin-Yew Lin, and Deepak Ravichandran · 2001
Earlier work this paper cites.
Learning question classifiers
Xin Li and Dan Roth · 2002
Earlier work this paper cites.
Pre-trained models for natural language processing: A survey
Xipeng Qiu, Tianxiang Sun, Yige Xu, Yunfan Shao, Ning Dai, and Xuanjing Huang · 2003
Earlier work this paper cites.
Submodular functions and optimization
Satoru Fujishige · 2005
Earlier work this paper cites.
Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales
Bo Pang and Lillian Lee · 2005
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
A high-throughput screening approach to discovering good forms of biologically inspired visual representation
Nicolas Pinto, David Doukhan, James J. DiCarlo, and David D. Cox · 2009
Earlier work this paper cites.
Algorithms for hyper-parameter optimization
James Bergstra, Rémi Bardenet, Yoshua Bengio, and Balázs Kégl · 2011
Earlier work this paper cites.
Learning the easy things first: Self-paced visual category discovery
Yong Jae Lee and Kristen Grauman · 2011
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts · 2011
Earlier work this paper cites.
Development of a benchmark corpus to support the automatic extraction of drug-related adverse effects from medical case reports
Harsha Gurulingappa, Abdul Mateen Rajput, Angus Roberts, Juliane Fluck, Martin Hofmann-Apitius, and Luca Toldo · 2012
Earlier work this paper cites.
Summarization through submodularity and dispersion
Anirban Dasgupta, Ravi Kumar, and Sujith Ravi · 2013
Earlier work this paper cites.
Submodularity for data selection in machine translation
Katrin Kirchhoff and Jeff Bilmes · 2014
Earlier work this paper cites.
GloVe: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning · 2014
Earlier work this paper cites.
Submodular subset selection for large-scale speech training data
Kai Wei, Yuzong Liu, Katrin Kirchhoff, Chris Bartels, and Jeff Bilmes · 2014
Earlier work this paper cites.
Unsupervised submodular subset selection for speech data
Kai Wei, Yuzong Liu, Katrin Kirchhoff, and Jeff Bilmes · 2014
Earlier work this paper cites.
Sampling from probabilistic submodular models
Alkis Gotovos, Hamed Hassani, and Andreas Krause · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Tiny imagenet visual recognition challenge, 2015
Ya Le and Xuan S. Yang · 2015
Earlier work this paper cites.
Online batch selection for faster training of neural networks
Ilya Loshchilov and Frank Hutter · 2015
Cited alongside, same era.
Lazier than lazy greedy
Baharan Mirzasoleiman, Ashwinkumar Badanidiyuru, Amin Karbasi, Jan Vondrák, and Andreas Krause · 2015
Cited alongside, same era.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al · 2015
Cited alongside, same era.
An exploration of softmax alternatives belonging to the spherical loss family
Alexandre de Brébisson and Pascal Vincent · 2016
Cited alongside, same era.
Weighted Random Sampling , pages 2365–2367
Pavlos Efraimidis and Paul (Pavlos) Spirakis · 2016
Cited alongside, same era.
Ro{bert}a: A robustly optimized {bert} pretraining approach, 2020
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2020
Later among the works it cites.
Coresets for data-efficient training of machine learning models, 2020
Baharan Mirzasoleiman, Jeff Bilmes, and Jure Leskovec · 2020
Later among the works it cites.
Green ai
Roy Schwartz, Jesse Dodge, Noah Smith, and Oren Etzioni · 2020
Later among the works it cites.
Mpnet: Masked and permuted pre-training for language understanding
Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie-Yan Liu · 2020
Later among the works it cites.
Efficientnet: Rethinking model scaling for convolutional neural networks, 2020
Mingxing Tan and Quoc V. Le · 2020
Later among the works it cites.
Curriculum learning by dynamic instance hardness
Tianyi Zhou, Shengjie Wang, and Jeffrey Bilmes · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kaiming He, X. Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Mobilenets: Efficient convolutional neural networks for mobile vision applications, 2017
Andrew G. Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam · 2017
Cited alongside, same era.
Hyperband: A novel bandit-based approach to hyperparameter optimization
Lisha Li, Kevin Jamieson, Giulia DeSalvo, Afshin Rostamizadeh, and Ameet Talwalkar · 2017
Cited alongside, same era.
SGDR: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2017
Cited alongside, same era.
Not all samples are created equal: Deep learning with importance sampling
Angelos Katharopoulos and Francois Fleuret · 2018
Cited alongside, same era.
Tune: A research platform for distributed model selection and training
Richard Liaw, Eric Liang, Robert Nishihara, Philipp Moritz, Joseph E Gonzalez, and Ion Stoica · 2018
Cited alongside, same era.
Active learning for convolutional neural networks: A core-set approach
Ozan Sener and Silvio Savarese · 2018
Cited alongside, same era.
Later among the works it cites.
Emerging properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Herv’e J’egou, Julien Mairal, Piotr Bojanowski, and Armand Joulin · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Later among the works it cites.
Vishal Kaushal, Suraj Kothawade, Ganesh Ramakrishnan, Jeff Bilmes, and Rishabh Iyer · 2021
Later among the works it cites.
Submodular mutual information for targeted data subset selection
Suraj Kothawade, Vishal Kaushal, Ganesh Ramakrishnan, Jeff Bilmes, and Rishabh Iyer · 2021
Later among the works it cites.
Deep learning on a data diet: Finding important examples early in training
Mansheej Paul, Surya Ganguli, and Gintare Karolina Dziugaite · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Later among the works it cites.
Gcr: Gradient coreset based replay buffer selection for continual learning
Rishabh Tiwari, Krishnateja Killamsetty, Rishabh K. Iyer, and Pradeep Shenoy · 2021
Later among the works it cites.
Medmnist classification decathlon: A lightweight automl benchmark for medical image analysis
Jiancheng Yang, Rui Shi, and Bingbing Ni · 2021
Later among the works it cites.
Compute-efficient deep learning: Algorithmic trends and opportunities, 2022
Brian R. Bartoldson, Bhavya Kailkhura, and Davis Blalock · 2022
Later among the works it cites.
ORIENT: Submodular mutual information measures for data subset selection under distribution shift
Athresh Karanam, Krishnateja Killamsetty, Harsha Kokel, and Rishabh K Iyer · 2022
Later among the works it cites.
Submodlib: A submodular optimization library, 2022
Vishal Kaushal, Ganesh Ramakrishnan, and Rishabh Iyer · 2022
Later among the works it cites.
Transformers in vision: A survey
Salman Khan, Muzammal Naseer, Munawar Hayat, Syed Waqas Zamir, Fahad Shahbaz Khan, and Mubarak Shah · 2022
Later among the works it cites.
AUTOMATA: Gradient based data subset selection for compute-efficient hyper-parameter tuning
Krishnateja Killamsetty, Guttu Sai Abhishek, Aakriti Lnu, Ganesh Ramakrishnan, Alexandre V. Evfimievski, Lucian Popa, and Rishabh K Iyer · 2022
Later among the works it cites.
Adaptive second order coresets for data-efficient machine learning
Omead Pooladzandi, David Davini, and Baharan Mirzasoleiman · 2022
Later among the works it cites.
Beyond neural scaling laws: beating power law scaling via data pruning
Ben Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli, and Ari S. Morcos · 2022
Later among the works it cites.
Medmnist v2-a large-scale lightweight benchmark for 2d and 3d biomedical image classification
Jiancheng Yang, Rui Shi, Donglai Wei, Zequan Liu, Lin Zhao, Bilian Ke, Hanspeter Pfister, and Bingbing Ni · 2023
Closest in time.