Fetching the paper…
Reading the bibliography…
Methods for carefully selecting or generating a small set of training data to learn from, i.e., data pruning, coreset selection, and data distillation, have been shown to be effective in reducing the ever-increasing cost of training neural networks.
A method of solving a convex programming problem with convergence rate O ( 1 / k 2 ) \text{O}(1/k^{2})
Yurii Evgen’evich Nesterov · 1983
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course
Yurii Nesterov · 2003
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Herding dynamical weights to learn
Max Welling · 2009
Earlier work this paper cites.
Super-samples from kernel herding
Yutian Chen, Max Welling, and Alex Smola · 2010
Earlier work this paper cites.
Better mini-batch algorithms via accelerated gradient methods
Andrew Cotter, Ohad Shamir, Nati Srebro, and Karthik Sridharan · 2011
Earlier work this paper cites.
Optimal distributed online prediction using mini-batches
Ofer Dekel, Ran Gilad-Bachrach, Ohad Shamir, and Lin Xiao · 2012
Earlier work this paper cites.
An optimal method for stochastic composite optimization
Guanghui Lan · 2012
Earlier work this paper cites.
Distributed balanced clustering via mapping coresets
MohammadHossein Bateni, Aditya Bhaskara, Silvio Lattanzi, and Vahab S Mirrokni · 2014
Earlier work this paper cites.
Coresets for nonparametric estimation-the case of dp-means
Olivier Bachem, Mario Lucic, and Andreas Krause · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Earlier work this paper cites.
Accelerated gradient methods for nonconvex nonlinear and stochastic programming
Saeed Ghadimi and Guanghui Lan · 2016
Earlier work this paper cites.
Train faster, generalize better: Stability of stochastic gradient descent
Moritz Hardt, Ben Recht, and Yoram Singer · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Earlier work this paper cites.
The lambada dataset: Word prediction requiring a broad discourse context
Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Quan Ngoc Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández · 2016
Earlier work this paper cites.
State-of-the-art speech recognition with sequence-to-sequence models
Chung-Cheng Chiu, Tara N Sainath, Yonghui Wu, Rohit Prabhavalkar, Patrick Nguyen, Zhifeng Chen, Anjuli Kannan, Ron J Weiss, Kanishka Rao, Ekaterina Gonina, et al · 2018
Earlier work this paper cites.
Adversarial active learning for deep networks: a margin based approach
Melanie Ducoffe and Frederic Precioso · 2018
Earlier work this paper cites.
On coresets for logistic regression
Alexander Munteanu, Chris Schwiegelshohn, Christian Sohler, and David P Woodruff · 2018
Earlier work this paper cites.
Active learning for convolutional neural networks: A core-set approach
Ozan Sener and Silvio Savarese · 2018
Earlier work this paper cites.
An empirical study of example forgetting during deep neural network learning
Mariya Toneva, Alessandro Sordoni, Remi Tachet des Combes, Adam Trischler, Yoshua Bengio, and Geoffrey J Gordon · 2018
Cited alongside, same era.
Tongzhou Wang, Jun-Yan Zhu, Antonio Torralba, and Alexei A Efros · 2018
Cited alongside, same era.
Teaching a black-box learner
Sanjoy Dasgupta, Daniel Hsu, Stefanos Poulis, and Xiaojin Zhu · 2019
Cited alongside, same era.
Openwebtext corpus
Aaron Gokaslan and Vanya Cohen · 2019
Cited alongside, same era.
Random shuffling beats sgd after finite epochs
Jeff Haochen and Suvrit Sra · 2019
Cited alongside, same era.
Using self-supervised learning can improve model robustness and uncertainty
Dan Hendrycks, Mantas Mazeika, Saurav Kadavath, and Dawn Song · 2019
A survey of deep active learning
Pengzhen Ren, Yun Xiao, Xiaojun Chang, Po-Yao Huang, Zhihui Li, Brij B Gupta, Xiaojiang Chen, and Xin Wang · 2021
Later among the works it cites.
Svp-cf: Selection via proxy for collaborative filtering data
Noveen Sachdeva, Carole-Jean Wu, and Julian McAuley · 2021
Later among the works it cites.
Dataset condensation with differentiable siamese augmentation
Bo Zhao and Hakan Bilen · 2021
Later among the works it cites.
Dataset condensation with gradient matching
Bo Zhao, Konda Reddy Mopuri, and Hakan Bilen · 2021
Later among the works it cites.
Dataset distillation by matching training trajectories
George Cazenavette, Tongzhou Wang, Antonio Torralba, Alexei A Efros, and Jun-Yan Zhu · 2022
Later among the works it cites.
Remember the past: Distilling datasets into addressable memories for neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Cited alongside, same era.
Stochastic nonconvex optimization with large minibatches
Weiran Wang and Nathan Srebro · 2019
Cited alongside, same era.
Contextual diversity for active learning
Sharat Agarwal, Himanshu Arora, Saket Anand, and Chetan Arora · 2020
Cited alongside, same era.
Flexible dataset distillation: Learn labels instead of images
Ondrej Bohdal, Yongxin Yang, and Timothy Hospedales · 2020
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Random reshuffling is not always better
Christopher M De Sa · 2020
Cited alongside, same era.
Zhiwei Deng and Olga Russakovsky · 2022
Later among the works it cites.
Deepcore: A comprehensive library for coreset selection in deep learning
Chengcheng Guo, Bo Zhao, and Yanbing Bai · 2022
Later among the works it cites.
Efficient adversarial training with data pruning
Maximilian Kaufmann, Yiren Zhao, Ilia Shumailov, Robert Mullins, and Nicolas Papernot · 2022
Later among the works it cites.
Dataset condensation via efficient synthetic-data parameterization
Jang-Hyun Kim, Jinuk Kim, Seong Joon Oh, Sangdoo Yun, Hwanjun Song, Joonhyun Jeong, Jung-Woo Ha, and Hyun Oh Song · 2022
Later among the works it cites.
Dataset condensation with contrastive signals
Saehyung Lee, Sanghyuk Chun, Sangwon Jung, Sangdoo Yun, and Sungroh Yoon · 2022
Later among the works it cites.
Grab: Finding provably better data permutations than random reshuffling
Yucheng Lu, Wentao Guo, and Christopher M De Sa · 2022
Later among the works it cites.
Prioritized training on points that are learnable, worth learning, and not yet learnt
Sören Mindermann, Jan M Brauner, Muhammed T Razzak, Mrinank Sharma, Andreas Kirsch, Winnie Xu, Benedikt Höltgen, Aidan N Gomez, Adrien Morisot, Sebastian Farquhar, et al · 2022
Later among the works it cites.
Active learning is a strong baseline for data subset selection
Dongmin Park, Dimitris Papailiopoulos, and Kangwook Lee · 2022
Later among the works it cites.
Beyond neural scaling laws: beating power law scaling via data pruning
Ben Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli, and Ari Morcos · 2022
Later among the works it cites.
Cafe: Learning to condense dataset by aligning features
Kai Wang, Bo Zhao, Xiangyu Peng, Zheng Zhu, Shuo Yang, Shuo Wang, Guan Huang, Hakan Bilen, Xinchao Wang, and Yang You · 2022
Later among the works it cites.
Data pruning and neural scaling laws: fundamental limitations of score-based algorithms
Fadhel Ayed and Soufiane Hayou · 2023
Closest in time.
Select without fear: Almost all mini-batch schedules generalize optimally
Konstantinos E Nikolakakis, Amin Karbasi, and Dionysis Kalogerias · 2023
Closest in time.
OpenAI · 2023
Closest in time.
Dataset distillation: A comprehensive review
Ruonan Yu, Songhua Liu, and Xinchao Wang · 2023
Closest in time.
Dataset condensation with distribution matching
Bo Zhao and Hakan Bilen · 2023
Closest in time.