Fetching the paper…
Reading the bibliography…
This paper studies the curious phenomenon for machine learning models with Transformer architectures that their activation maps are sparse.
The principle of parsimony and some applications in psychology
Robert Epstein · 1984
Earlier work this paper cites.
Neural networks and the bias/variance dilemma
Stuart Geman, Elie Bienenstock, and René Doursat · 1992
Earlier work this paper cites.
Emergence of simple-cell receptive field properties by learning a sparse code for natural images
Bruno A Olshausen and David J Field · 1996
Earlier work this paper cites.
Regression shrinkage and selection via the lasso
Robert Tibshirani · 1996
Earlier work this paper cites.
The role of occam’s razor in knowledge discovery
Pedro Domingos · 1999
Earlier work this paper cites.
Imaging input and output of neocortical networks in vivo
Jason ND Kerr, David Greenberg, and Fritjof Helmchen · 2005
Earlier work this paper cites.
An introduction to compressive sampling
Emmanuel J Candès and Michael B Wakin · 2008
Earlier work this paper cites.
Odor representations in olfactory cortex:“sparse” coding, global inhibition, and oscillations
Cindy Poo and Jeffry S Isaacson · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning fast approximations of sparse coding
Karol Gregor and Yann LeCun · 2010
Earlier work this paper cites.
Experimental evidence for sparse firing in the neocortex
Alison L Barth and James FA Poulet · 2012
Earlier work this paper cites.
Low-rank approximations for conditional feedforward computation in deep neural networks
Andrew Davis and Itamar Arel · 2013
Earlier work this paper cites.
Asymmetric lsh (alsh) for sublinear time maximum inner product search (mips)
Anshumali Shrivastava and Ping Li · 2014
Earlier work this paper cites.
Sparse modeling for image and vision processing
Julien Mairal, Francis Bach, Jean Ponce, et al · 2014
Earlier work this paper cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Earlier work this paper cites.
Generalized principal component analysis
Rene Vidal, Yi Ma, and Shankar Sastry · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Obtaining well calibrated probabilities using bayesian binning
Mahdi Pakdaman Naeini, Gregory Cooper, and Milos Hauskrecht · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger · 2017
Earlier work this paper cites.
Robust loss functions under label noise for deep neural networks
Aritra Ghosh, Himanshu Kumar, and PS Sastry · 2017
Earlier work this paper cites.
Convolutional neural networks analyzed via convolutional sparse coding
Vardan Papyan, Yaniv Romano, and Michael Elad · 2017
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2017
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Earlier work this paper cites.
Compressing dma engine: Leveraging activation sparsity for training deep neural networks
Minsoo Rhu, Mike O’Connor, Niladrish Chatterjee, Jeff Pool, Youngeun Kwon, and Stephen W Keckler · 2018
Earlier work this paper cites.
Sparse dnns with improved adversarial robustness
Yiwen Guo, Chao Zhang, Changshui Zhang, and Yurong Chen · 2018
Earlier work this paper cites.
Theoretical foundations of deep learning via sparse representations: A multilayer sparse model and its connection to convolutional neural networks
Vardan Papyan, Yaniv Romano, Jeremias Sulam, and Michael Elad · 2018
Earlier work this paper cites.
Multilayer convolutional sparse modeling: Pursuit and dictionary learning
Jeremias Sulam, Vardan Papyan, Yaniv Romano, and Michael Elad · 2018
Earlier work this paper cites.
Supervised deep sparse coding networks
Xiaoxia Sun, Nasser M Nasrabadi, and Trac D Tran · 2018
Earlier work this paper cites.
How can we be so dense? the benefits of using highly sparse representations
Subutai Ahmad and Luiz Scheinkman · 2019
Earlier work this paper cites.
Billion-scale similarity search with GPUs
Jeff Johnson, Matthijs Douze, and Hervé Jégou · 2019
Earlier work this paper cites.
Benchmarking neural network robustness to common corruptions and perturbations
Dan Hendrycks and Thomas Dietterich · 2019
Earlier work this paper cites.
Accelerating convolutional neural networks via activation map compression
Georgios Georgiadis · 2019
Earlier work this paper cites.
Seernet: Predicting convolutional neural network feature-map sparsity through low-bit quantization
Shijie Cao, Lingxiao Ma, Wencong Xiao, Chen Zhang, Yunxin Liu, Lintao Zhang, Lanshun Nie, and Zhi Yang · 2019
Earlier work this paper cites.
What is one grain of sand in the desert? analyzing individual neurons in deep nlp models
Fahim Dalvi, Nadir Durrani, Hassan Sajjad, Yonatan Belinkov, Anthony Bau, and James Glass · 2019
Earlier work this paper cites.
Implicit regularization for optimal sparse recovery
Tomas Vaskevicius, Varun Kanade, and Patrick Rebeschini · 2019
Cited alongside, same era.
Peng Zhao, Yun Yang, and Qiao-Chu He · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Fixup initialization: Residual learning without normalization
Hongyi Zhang, Yann N Dauphin, and Tengyu Ma · 2019
Cited alongside, same era.
Natural questions: a benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al · 2019
Cited alongside, same era.
Mlp-mixer: An all-mlp architecture for vision
Ilya O Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Andreas Steiner, Daniel Keysers, Jakob Uszkoreit, et al · 2021
Later among the works it cites.
Rezero is all you need: Fast convergence at large depth
Thomas Bachlechner, Bodhisattwa Prasad Majumder, Henry Mao, Gary Cottrell, and Julian McAuley · 2021
Later among the works it cites.
Going deeper with image transformers
Hugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve, and Hervé Jégou · 2021
Later among the works it cites.
Moefication: Transformer feed-forward layers are mixtures of experts
Zhengyan Zhang, Yankai Lin, Zhiyuan Liu, Peng Li, Maosong Sun, and Jie Zhou · 2022
Closest in time.
Scenic: A jax library for computer vision research and beyond
Mostafa Dehghani, Alexey Gritsenko, Anurag Arnab, Matthias Minderer, and Yi Tay · 2022
Closest in time.
High-Dimensional Data Analysis with Low-Dimensional Models: Principles, Computation, and Applications
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hippocampal network reorganization underlies the formation of a temporal association memory
Mohsin S Ahmed, James B Priestley, Angel Castro, Fabio Stefanini, Ana Sofia Solis Canales, Elizabeth M Balough, Erin Lavoie, Luca Mazzucato, Stefano Fusi, and Attila Losonczy · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J Liu, et al · 2020
Cited alongside, same era.
Accelerating large-scale inference with anisotropic vector quantization
Ruiqi Guo, Philip Sun, Erik Lindgren, Quan Geng, David Simcha, Felix Chern, and Sanjiv Kumar · 2020
Cited alongside, same era.
Slide: In defense of smart algorithms over hardware acceleration for large-scale deep learning systems
Beidi Chen, Tharun Medini, James Farwell, Charlie Tai, Anshumali Shrivastava, et al · 2020
Cited alongside, same era.
Inducing and exploiting activation sparsity for fast inference on deep neural networks
Mark Kurtz, Justin Kopinsky, Rati Gelashvili, Alexander Matveev, John Carr, Michael Goin, William Leiserson, Sage Moore, Nir Shavit, and Dan Alistarh · 2020
Cited alongside, same era.
Robust recovery via implicit bias of discrepant learning rates for double over-parameterization
Chong You, Zhihui Zhu, Qing Qu, and Yi Ma · 2020
Cited alongside, same era.
Kernel and rich regimes in overparametrized models
Blake Woodworth, Suriya Gunasekar, Jason D Lee, Edward Moroshko, Pedro Savarese, Itay Golan, Daniel Soudry, and Nathan Srebro · 2020
Cited alongside, same era.
John Wright and Yi Ma · 2022
Closest in time.
Tpu-knn: K nearest neighbor search at peak flop/s
Felix Chern, Blake Hechtman, Andy Davis, Ruiqi Guo, David Majnemer, and Sanjiv Kumar · 2022
Closest in time.
Learning from noisy labels with deep neural networks: A survey
Hwanjun Song, Minseok Kim, Dongmin Park, Yooju Shin, and Jae-Gil Lee · 2022
Closest in time.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer · 2022
Closest in time.
Glam: Efficient scaling of language models with mixture-of-experts
Nan Du, Yanping Huang, Andrew M Dai, Simon Tong, Dmitry Lepikhin, Yuanzhong Xu, Maxim Krikun, Yanqi Zhou, Adams Wei Yu, Orhan Firat, et al · 2022
Closest in time.
Deepspeed-moe: Advancing mixture-of-experts inference and training to power next-generation ai scale
Samyam Rajbhandari, Conglong Li, Zhewei Yao, Minjia Zhang, Reza Yazdani Aminabadi, Ammar Ahmad Awan, Jeff Rasley, and Yuxiong He · 2022
Closest in time.
A review of sparse expert models in deep learning
William Fedus, Jeff Dean, and Barret Zoph · 2022
Closest in time.
A theoretical view on sparsely activated networks
Cenk Baykal, Nishanth Dikkala, Rina Panigrahy, Cyrus Rashtchian, and Xin Wang · 2022
Closest in time.
Sparsity winning twice: Better robust generalization from more efficient training
Tianlong Chen, Zhenyu Zhang, Santosh Balachandra, Haoyu Ma, Zehao Wang, Zhangyang Wang, et al · 2022
Closest in time.
Superior generalization of smaller models in the presence of significant label noise
Yihao Xue, Kyle Whitecross, and Baharan Mirzasoleiman · 2022
Closest in time.
Robust training under label noise by over-parameterization
Sheng Liu, Zhihui Zhu, Qing Qu, and Chong You · 2022
Closest in time.
Adversarial robustness of sparse local lipschitz predictors
Ramchandran Muthukumar and Jeremias Sulam · 2022
Closest in time.
Knowledge neurons in pretrained transformers
Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, Baobao Chang, and Furu Wei · 2022
Closest in time.
Softmax linear units
Nelson Elhage, Tristan Hume, Catherine Olsson, Neel Nanda, Tom Henighan, Scott Johnston, Sheer ElShowk, Nicholas Joseph, Nova DasSarma, Ben Mann, Danny Hernandez, Amanda Askell, Kamal Ndousse, Jones, , Dawn Drain, Anna Chen, Yuntao Bai, Deep Ganguli, Liane Lovitt, Zac Hatfield-Dodds, Jackson Kernion, Tom Conerly, Shauna Kravec, Stanislav Fort, Saurav Kadavath, Josh Jacobson, Eli Tran-Johnson, Jared Kaplan, Jack Clark, Tom Brown, Sam McCandlish, Dario Amodei, and Christopher Olah · 2022
Closest in time.
Self-conditioning pre-trained language models
Xavier Suau Cuadros, Luca Zappella, and Nicholas Apostoloff · 2022
Closest in time.
Revisiting sparse convolutional model for visual recognition
Xili Dai, Mingyang Li, Pengyuan Zhai, Shengbang Tong, Xingjian Gao, Shao-Lun Huang, Zhihui Zhu, Chong You, and Yi Ma · 2022
Closest in time.
Implicit bias of the step size in linear diagonal neural networks
Mor Shpigel Nacson, Kavya Ravichandran, Nathan Srebro, and Daniel Soudry · 2022
Closest in time.
Recovery and generalization in over-realized dictionary learning
Jeremias Sulam, Chong You, and Zhihui Zhu · 2022
Closest in time.
On the robustness of minimum norm interpolators and regularized empirical risk minimizers
Geoffrey Chinot, Matthias Löffler, and Sara van de Geer · 2022
Closest in time.
Fast rates for noisy interpolation require rethinking the effects of inductive bias
Konstantin Donhauser, Nicolo Ruggeri, Stefan Stojanovic, and Fanny Yang · 2022
Closest in time.
Deep learning meets nonparametric regression: Are weight-decayed dnns locally adaptive?
Kaiqi Zhang and Yu-Xiang Wang · 2022
Closest in time.
Sgd with large step sizes learns sparse features
Maksym Andriushchenko, Aditya Varre, Loucas Pillaud-Vivien, and Nicolas Flammarion · 2022
Closest in time.
Extended unconstrained features model for exploring deep neural collapse
Tom Tirer and Joan Bruna · 2022
Closest in time.
Imbalance trouble: Revisiting neural-collapse geometry
Christos Thrampoulidis, Ganesh Ramachandra Kini, Vala Vakilian, and Tina Behnia · 2022
Closest in time.
Pathways: Asynchronous distributed dataflow for ml
Paul Barham, Aakanksha Chowdhery, Jeff Dean, Sanjay Ghemawat, Steven Hand, Daniel Hurt, Michael Isard, Hyeontaek Lim, Ruoming Pang, Sudip Roy, et al · 2022
Closest in time.
A path towards autonomous machine intelligence
Yann LeCun · 2022
Closest in time.
On the principles of parsimony and self-consistency for the emergence of intelligence
Yi Ma, Doris Tsao, and Heung-Yeung Shum · 2022
Closest in time.
The application of artificial intelligence to biology and neuroscience
Blake Richards, Doris Tsao, and Anthony Zador · 2022
Closest in time.
Decoupled context processing for context augmented language modeling
Zonglin Li, Ruiqi Guo, and Sanjiv Kumar · 2022
Closest in time.
Sharper analysis of sparsely activated wide neural networks with trainable biases
Hongru Yang, Ziyu Jiang, Ruizhe Zhang, Zhangyang Wang, and Yingbin Liang · 2023
Closest in time.
Deep neural collapse is provably optimal for the deep unconstrained features model
Peter Súkeník, Marco Mondelli, and Christoph Lampert · 2023
Closest in time.