Fetching the paper…
Reading the bibliography…
Despite their ability to memorize large datasets, deep neural networks often achieve good generalization performance.
A Closer Look at Memorization in Deep Networks
Devansh Arpit, Stanisław Jastrzȩbski, Nicolas Ballas, David Krueger, Emmanuel Bengio, Maxinder S. Kanwal, Tegan Maharaj, Asja Fischer, Aaron Courville, Yoshua Bengio, and Simon Lacoste-Julien · 1938
Earlier work this paper cites.
The orientation and direction selectivity of cells in macaque visual cortex
Russell L De Valois, E William Yund, and Norva Hepler · 1982
Earlier work this paper cites.
The analysis of visual motion: a comparison of neuronal and psychophysical performance
Kenneth H Britten, Michael N Shadlen, William T Newsome, and J Anthony Movshon · 1992
Earlier work this paper cites.
Flat Minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Stability and generalization
Olivier Bousquet and André Elisseeff · 2002
Earlier work this paper cites.
Neural correlations, population coding and computation
Bruno B Averbeck, Peter E Latham, and Alexandre Pouget · 2006
Earlier work this paper cites.
Experience-dependent representation of visual categories in parietal cortex
David J Freedman and John a Assad · 2006
Earlier work this paper cites.
Visualizing Higher-Layer Features of a Deep Network Technical Report
Dumitru Erhan, Yoshua Bengio, Aaron Courville, and Pascal Vincent · 2009
Earlier work this paper cites.
Building high-level features using large scale unsupervised learning
Quoc V Le, Marc’Aurelio Ranzato, Rajat Monga, Matthieu Devin, Kai Chen, Greg S Corrado, Jeff Dean, and Andrew Y Ng · 2011
Earlier work this paper cites.
Emergence of Object-Selective Features in Unsupervised Feature Learning
Adam Coates, Andrej Karpathy, and Andrew Y Ng · 2012
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2013
Earlier work this paper cites.
Context-dependent computation by recurrent dynamics in prefrontal cortex
Valerio Mante, David Sussillo, Krishna V Shenoy, and William T Newsome · 2013
Earlier work this paper cites.
The importance of mixed selectivity in complex cognitive tasks
Mattia Rigotti, Omri Barak, Melissa R Warden, Xiao-Jing Wang, Nathaniel D Daw, Earl K Miller, and Stefano Fusi · 2013
Earlier work this paper cites.
Analyzing the Performance of Multilayer Neural Networks for Object Recognition
Pulkit Agrawal, Ross B Girshick, and Jitendra Malik · 2014
Earlier work this paper cites.
A category-free neural population supports evolving demands during decision-making
David Raposo, Matthew T Kaufman, and Anne K Churchland · 2014
Cited alongside, same era.
Dropout : A Simple Way to Prevent Neural Networks from Overfitting
Nitish Srivastava, Geoffrey E. Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Cited alongside, same era.
Visualizing and understanding convolutional networks
Matthew Zeiler and Rob Fergus · 2014
Cited alongside, same era.
Object Detectors Emerge in Deep Scene CNNs
Bolei Zhou, Aditya Khosla, Agata Lapedriza, Aude Oliva, and Antonio Torralba · 2014
Cited alongside, same era.
Structured Pruning of Deep Convolutional Neural Networks
Sajid Anwar, Kyuyeon Hwang, and Wonyong Sung · 2015
Cited alongside, same era.
Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
On the robustness of convolutional neural networks to internal architecture and weight perturbations
Nicholas Cheney, Martin Schrimpf, and Gabriel Kreiman · 2017
Later among the works it cites.
Sharp Minima Can Generalize For Deep Nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Later among the works it cites.
Gintare Karolina Dziugaite and Daniel M. Roy · 2017
Later among the works it cites.
On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikahail Smelyanskiy, and Ping Tak Peter Tang · 2017
Later among the works it cites.
Pruning filters for efficient ConvNets
Hao Li, Asim Kadav, Igor Durdanovic, Hanan Samet, and Hans Peter Graf · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
Understanding intermediate layers using linear classifier probes
Guillaume Alain and Yoshua Bengio · 2016
Cited alongside, same era.
Optimal compensation for neuron loss
David G.T. Barrett, Sophie Denève, and Christian K. Machens · 2016
Cited alongside, same era.
Population-Level Neural Codes Are Robust to Single-Neuron Variability from a Multidimensional Coding Perspective
Jorrit S. Montijn, Guido T. Meijer, Carien S. Lansink, and Cyriel M A Pennartz · 2016
Cited alongside, same era.
History-dependent variability in population dynamics during evidence accumulation in cortex
Ari S Morcos and Christopher D Harvey · 2016
Cited alongside, same era.
On the Expressive Power of Deep Neural Networks
Maithra Raghu, Ben Poole, Jon Kleinberg, Surya Ganguli, and Jascha Sohl-Dickstein · 2016
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al · 2016
Cited alongside, same era.
Pruning Convolutional Neural Networks for Resource Efficient Inference
Pavlo Molchanov, Stephen Tyree, Tero Karras, Timo Aila, and Jan Kautz · 2017
Later among the works it cites.
Exploring Generalization in Deep Learning
Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, and Nathan Srebro · 2017
Later among the works it cites.
Learning to Generate Reviews and Discovering Sentiment
Alec Radford, Rafal Jozefowicz, and Ilya Sutskever · 2017
Later among the works it cites.
SVCCA: Singular Vector Canonical Correlation Analysis for Deep Understanding and Improvement
Maithra Raghu, Justin Gilmer, Jason Yosinski, and Jascha Sohl-Dickstein · 2017
Later among the works it cites.
Opening the Black Box of Deep Neural Networks via Information
Ravid Shwartz-Ziv and Naftali Tishby · 2017
Later among the works it cites.
Understanding Generalization and Stochastic Gradient Descent
Samuel L. Smith and Quoc V. Le · 2017
Later among the works it cites.
The Marginal Value of Adaptive Gradient Methods in Machine Learning
Ashia C. Wilson, Rebecca Roelofs, Mitchell Stern, Nathan Srebro, and Benjamin Recht · 2017
Later among the works it cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Later among the works it cites.
Untuned but not irrelevant: The role of untuned neurons in sensory information coding
Joel Zylberberg · 2017
Later among the works it cites.