Fetching the paper…
Reading the bibliography…
Data pruning aims to obtain lossless performances with less overall cost.
Large batch optimization for deep learning: Training bert in 76 minutes, 2019
Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh · 1904
Earlier work this paper cites.
Stochastic gradient learning in neural networks
Léon Bottou et al · 1991
Earlier work this paper cites.
On coresets for k-means and k-median clustering
Sariel Har-Peled and Soham Mazumdar · 2004
Earlier work this paper cites.
On coresets for k-median and k-means clustering in metric and euclidean spaces and their applications
Ke Chen · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Herding dynamical weights to learn
Max Welling · 2009
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale, 2020
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2010
Earlier work this paper cites.
Training data-efficient image transformers & distillation through attention, 2020
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2014
Diederik P. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Stochastic optimization with importance sampling, 2014
Peilin Zhao and Tong Zhang · 2014
Earlier work this paper cites.
Neural networks and back propagation algorithm
Mirza Cilimkovic · 2015
Earlier work this paper cites.
Tiny imagenet visual recognition challenge
Ya Le and Xuan S. Yang · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Importance sampling for minibatches, 2016
Dominik Csiba and Peter Richtárik · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts, 2016
Ilya Loshchilov and Frank Hutter · 2016
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Earlier work this paper cites.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Earlier work this paper cites.
Understanding black-box predictions via influence functions
Pang Wei Koh and Percy Liang · 2017
Earlier work this paper cites.
Super-convergence: Very fast training of neural networks using large learning rates, 2017
Leslie N. Smith and Nicholay Topin · 2017
Earlier work this paper cites.
Large batch training of convolutional networks, 2017
Yang You, Igor Gitman, and Boris Ginsburg · 2017
Cited alongside, same era.
Random erasing data augmentation, 2017
Zhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li, and Yi Yang · 2017
Cited alongside, same era.
Scene parsing through ade20k dataset
Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba · 2017
Cited alongside, same era.
Dataset quantization, 2023
Daquan Zhou, Kai Wang, Jianyang Gu, Xiangyu Peng, Dongze Lian, Yifan Zhang, Yang You, and Jiashi Feng · 2017
Cited alongside, same era.
Adversarial active learning for deep networks: a margin based approach
Melanie Ducoffe and Frederic Precioso · 2018
Cited alongside, same era.
Training deep models faster with robust, approximate importance sampling
Dataset meta-learning from kernel ridge-regression
Timothy Nguyen, Zhourong Chen, and Jaehoon Lee · 2020
Later among the works it cites.
pytorch-fid: FID Score for PyTorch
Maximilian Seitzer · 2020
Later among the works it cites.
Coresets for near-convex functions
Morad Tukan, Alaa Maalouf, and Dan Feldman · 2020
Later among the works it cites.
Submodular combinatorial information measures with applications in machine learning
Rishabh Iyer, Ninad Khargoankar, Jeff Bilmes, and Himanshu Asanani · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tyler B Johnson and Carlos Guestrin · 2018
Cited alongside, same era.
An empirical model of large-batch training
Sam McCandlish, Jared Kaplan, Dario Amodei, and OpenAI Dota Team · 2018
Cited alongside, same era.
Active learning for convolutional neural networks: A core-set approach
Ozan Sener and Silvio Savarese · 2018
Cited alongside, same era.
An empirical study of example forgetting during deep neural network learning, 2018
Mariya Toneva, Alessandro Sordoni, Remi Tachet des Combes, Adam Trischler, Yoshua Bengio, and Geoffrey J. Gordon · 2018
Cited alongside, same era.
Unified perceptual parsing for scene understanding, 2018
Tete Xiao, Yingcheng Liu, Bolei Zhou, Yuning Jiang, and Jian Sun · 2018
Cited alongside, same era.
mixup: Beyond empirical risk minimization, 2018
Hongyi Zhang, Moustapha Cisse, Yann N. Dauphin, and David Lopez-Paz · 2018
Cited alongside, same era.
Selection via proxy: Efficient data selection for deep learning
Cody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman, Peter Bailis, Percy Liang, Jure Leskovec, and Matei Zaharia · 2019
Cited alongside, same era.
Katerina Margatina, Giorgos Vernikos, Loïc Barrault, and Nikolaos Aletras · 2021
Later among the works it cites.
Dataset distillation with infinitely wide convolutional networks
Timothy Nguyen, Roman Novak, Lechao Xiao, and Jaehoon Lee · 2021
Later among the works it cites.
Deep learning on a data diet: Finding important examples early in training, 2021
Mansheej Paul, Surya Ganguli, and Gintare Karolina Dziugaite · 2021
Later among the works it cites.
Accelerating deep learning with dynamic data pruning, 2021
Ravi S Raju, Kyle Daruwalla, and Mikko Lipasti · 2021
Later among the works it cites.
High-resolution image synthesis with latent diffusion models, 2021
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2021
Later among the works it cites.
Core-set sampling for efficient neural architecture search
Jae-hun Shim, Kyeongbo Kong, and Suk-Ju Kang · 2021
Later among the works it cites.
Resnet strikes back: An improved training procedure in timm, 2021
Ross Wightman, Hugo Touvron, and Hervé Jégou · 2021
Later among the works it cites.
Dataset condensation with distribution matching
Bo Zhao and Hakan Bilen · 2021
Later among the works it cites.
Dataset distillation by matching training trajectories
George Cazenavette, Tongzhou Wang, Antonio Torralba, Alexei A. Efros, and Jun-Yan Zhu · 2022
Later among the works it cites.
Cafe: Learning to condense dataset by aligning features
Kai Wang, Bo Zhao, Xiangyu Peng, Zheng Zhu, Shuo Yang, Shuo Wang, Guan Huang, Hakan Bilen, Xinchao Wang, and Yang You · 2022
Later among the works it cites.
Scaling vision transformers to 22 billion parameters
Mostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski, Jonathan Heek, Justin Gilmer, Andreas Steiner, Mathilde Caron, Robert Geirhos, Ibrahim Alabdulmohsin, et al · 2023
Closest in time.
Large-scale dataset pruning with dynamic uncertainty, 2023
Muyang He, Shuo Yang, Tiejun Huang, and Bo Zhao · 2023
Closest in time.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Closest in time.
Llama: Open and efficient foundation language models, 2023
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample · 2023
Closest in time.
Dataset pruning: Reducing training data by examining generalization influence
Shuo Yang, Zeke Xie, Hanyu Peng, Min Xu, Mingming Sun, and Ping Li · 2023
Closest in time.