Fetching the paper…
Reading the bibliography…
We define a notion of information that an individual sample provides to the training of a neural network, and we specialize it to measure both how much a sample informs the final weights and how much it informs the function computed by the weights.
Solution of the matrix equation ax + xb = c [f4]
R. H. Bartels and G. W. Stewart · 1972
Earlier work this paper cites.
Detection of influential observation in linear regression
R Dennis Cook · 1977
Earlier work this paper cites.
Downdating the singular value decomposition
Ming Gu and Stanley C Eisenstat · 1995
Earlier work this paper cites.
Fisher information and stochastic complexity
J. J. Rissanen · 1996
Earlier work this paper cites.
Wrappers for feature subset selection
Ron Kohavi, George H John, et al · 1997
Earlier work this paper cites.
Stability and generalization
Olivier Bousquet and André Elisseeff · 2002
Earlier work this paper cites.
Database-friendly random projections: Johnson-lindenstrauss with binary coins
Dimitris Achlioptas · 2003
Earlier work this paper cites.
Elements of information theory
Thomas M Cover and Joy A Thomas · 2006
Earlier work this paper cites.
Very sparse random projections
Ping Li, Trevor J Hastie, and Kenneth W Church · 2006
Earlier work this paper cites.
Nonnegative decomposition of multivariate information
Paul L. Williams and Randall D. Beer · 2010
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts · 2011
Earlier work this paper cites.
Unbiased look at dataset bias
Antonio Torralba and Alexei A Efros · 2011
Earlier work this paper cites.
Bayesian learning via stochastic gradient langevin dynamics
Max Welling and Yee Whye Teh · 2011
Earlier work this paper cites.
Dogs vs. Cats , 2013
Kaggle · 2013
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning · 2014
Earlier work this paper cites.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus · 2014
Cited alongside, same era.
Explaining and harnessing adversarial examples
Ian Goodfellow, Jonathon Shlens, and Christian Szegedy · 2015
Cited alongside, same era.
Train faster, generalize better: Stability of stochastic gradient descent
Moritz Hardt, Benjamin Recht, and Yoram Singer · 2015
Cited alongside, same era.
Algorithmic stability for adaptive data analysis
Raef Bassily, Kobbi Nissim, Adam Smith, Thomas Steinke, Uri Stemmer, and Jonathan Ullman · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Where is the information in a deep neural network?
Alessandro Achille, Giovanni Paolini, and Stefano Soatto · 2019
Later among the works it cites.
Data shapley: Equitable valuation of data for machine learning
Amirata Ghorbani and James Zou · 2019
Later among the works it cites.
Wide neural networks of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Later among the works it cites.
How complex is your classification problem? a survey on measuring classification complexity
Ana C Lorena, Luís PF Garcia, Jens Lehmann, Marcilio CP Souto, and Tin Kam Ho · 2019
Later among the works it cites.
icassava 2019fine-grained visual categorization challenge
Ernest Mwebaze, Timnit Gebru, Andrea Frome, Solomon Nsumba, and Jeremy Tusubira · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Information-theoretic analysis of stability and bias of learning algorithms
Maxim Raginsky, Alexander Rakhlin, Matthew Tsao, Yihong Wu, and Aolin Xu · 2016
Cited alongside, same era.
Emnist: an extension of mnist to handwritten letters
Gregory Cohen, Saeed Afshar, Jonathan Tapson, and André van Schaik · 2017
Cited alongside, same era.
Understanding black-box predictions via influence functions
Pang Wei Koh and Percy Liang · 2017
Cited alongside, same era.
Stochastic modified equations and adaptive stochastic gradient algorithms
Qianxiao Li, Cheng Tai, and Weinan E · 2017
Cited alongside, same era.
Stochastic gradient descent as approximate bayesian inference
Stephan Mandt, Matthew D Hoffman, and David M Blei · 2017
Cited alongside, same era.
Opening the black box of deep neural networks via information
Ravid Shwartz-Ziv and Naftali Tishby · 2017
Cited alongside, same era.
Emergence of invariance and disentanglement in deep representations
Alessandro Achille and Stefano Soatto · 2018
Cited alongside, same era.
Later among the works it cites.
On the information bottleneck theory of deep learning
Andrew M Saxe, Yamini Bansal, Joel Dapello, Madhu Advani, Artemy Kolchinsky, Brendan D Tracey, and David D Cox · 2019
Later among the works it cites.
An empirical study of example forgetting during deep neural network learning
Mariya Toneva, Alessandro Sordoni, Remi Tachet des Combes, Adam Trischler, Yoshua Bengio, and Geoffrey J. Gordon · 2019
Later among the works it cites.
Data valuation using reinforcement learning
Jinsung Yoon, Sercan O Arik, and Tomas Pfister · 2019
Later among the works it cites.
Influence functions in deep learning are fragile
Samyadeep Basu, Philip Pope, and Soheil Feizi · 2020
Later among the works it cites.
Forgetting outside the box: Scrubbing deep networks of information accessible from input-output observations
Aditya Golatkar, Alessandro Achille, and Stefano Soatto · 2020
Later among the works it cites.
Evaluation of neural architectures trained with square loss vs cross-entropy in classification tasks
Like Hui and Mikhail Belkin · 2020
Later among the works it cites.
Gradients as features for deep representation learning
Fangzhou Mu, Yingyu Liang, and Yin Li · 2020
Later among the works it cites.
Information in infinite ensembles of infinitely-wide neural networks
Ravid Shwartz-Ziv and Alexander A Alemi · 2020
Later among the works it cites.
Deltagrad: Rapid retraining of machine learning models
Yinjun Wu, Edgar Dobriban, and Susan B Davidson · 2020
Later among the works it cites.
Predicting training time without training
Luca Zancato, Alessandro Achille, Avinash Ravichandran, Rahul Bhotika, and Stefano Soatto · 2020
Later among the works it cites.