Fetching the paper…
Reading the bibliography…
Deep Learning has revolutionized the fields of computer vision, natural language understanding, speech recognition, information retrieval and more.
The State of Sparsity in Deep Neural Networks
Trevor Gale, Erich Elsen, and Sara Hooker. 2019 · 1902
Earlier work this paper cites.
On Bayesian methods for seeking the extremum. In
Jonas Močkus. 1975 · 1975
Earlier work this paper cites.
Introduction to VLSI systems
HT Kung and CE Leiserson. 1980 · 1980
Earlier work this paper cites.
Why systolic architectures?
Hsiang-Tsung Kung. 1982 · 1982
Earlier work this paper cites.
Neural network ensembles
Lars Kai Hansen and Peter Salamon. 1990 · 1990
Earlier work this paper cites.
Optimal brain damage. In
Yann LeCun, John S Denker, and Sara A Solla. 1990 · 1990
Earlier work this paper cites.
Optimal brain surgeon and general network pruning. In
Babak Hassibi, David G Stork, and Gregory J Wolff. 1993 · 1993
Earlier work this paper cites.
Neural network ensembles, cross validation and active learning
Anders Krogh and Jesper Vedelsby. 1994 · 1994
Earlier work this paper cites.
Combining labeled and unlabeled data with co-training. In
Avrim Blum and Tom Mitchell. 1998 · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann Lecun, Leon Bottou, Yoshua Bengio, and Patrick Haffner. 1998 · 1998
Earlier work this paper cites.
Ensemble methods in machine learning. In
Thomas G Dietterich. 2000 · 2000
Earlier work this paper cites.
Similarity estimation techniques from rounding algorithms. In
Moses S Charikar. 2002 · 2002
Earlier work this paper cites.
SMOTE: synthetic minority over-sampling technique
Nitesh V Chawla, Kevin W Bowyer, Lawrence O Hall, and W Philip Kegelmeyer. 2002 · 2002
Earlier work this paper cites.
Best practices for convolutional neural networks applied to visual document analysis.. In
Patrice Y Simard, David Steinkraus, John C Platt, et al · 2003
Earlier work this paper cites.
Training with quantization noise for extreme model compression
Angela Fan, Pierre Stock, Benjamin Graham, Edouard Grave, Rémi Gribonval, Hervé Jégou, and Armand Joulin. 2020 · 2004
Earlier work this paper cites.
Model compression. In
Cristian Buciluǎ, Rich Caruana, and Alexandru Niculescu-Mizil. 2006 · 2006
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database. In
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009 · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Large-scale deep unsupervised learning using graphics processors. In
Rajat Raina, Anand Madhavan, and Andrew Y Ng. 2009 · 2009
Earlier work this paper cites.
High-performance neural networks for visual object classification
Dan C Cireşan, Ueli Meier, Jonathan Masci, Luca M Gambardella, and Jürgen Schmidhuber. 2011 · 2011
Earlier work this paper cites.
Improving the speed of neural networks on CPUs
Vincent Vanhoucke, Andrew Senior, and Mark Z Mao. 2011 · 2011
Earlier work this paper cites.
Random search for hyper-parameter optimization
James Bergstra and Yoshua Bengio. 2012 · 2012
Earlier work this paper cites.
Towards unsupervised speech processing. In
James Glass. 2012 · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. 2012 · 2012
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Yoshua Bengio, Nicholas Léonard, and Aaron Courville. 2013 · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Deep speech: Scaling up end-to-end speech recognition
Awni Hannun, Carl Case, Jared Casper, Bryan Catanzaro, Greg Diamos, Erich Elsen, Ryan Prenger, Sanjeev Satheesh, Shubho Sengupta, Adam Coates, et al · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation. In
Jeffrey Pennington, Richard Socher, and Christopher D Manning. 2014 · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014 · 2014
Earlier work this paper cites.
Unsupervised visual representation learning by context prediction. In
Carl Doersch, Abhinav Gupta, and Alexei A Efros. 2015 · 2015
Earlier work this paper cites.
Song Han, Huizi Mao, and William J Dally. 2015a · 2015
Earlier work this paper cites.
Learning both weights and connections for efficient neural networks
Song Han, Jeff Pool, John Tran, and William J Dally. 2015b · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015 · 2015
Earlier work this paper cites.
Going deeper with convolutions. In
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. 2015 · 2015
Earlier work this paper cites.
Tensorflow: A system for large-scale machine learning. In
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al · 2016
Earlier work this paper cites.
Quasi-recurrent neural networks
James Bradbury, Stephen Merity, Caiming Xiong, and Richard Socher. 2016 · 2016
Earlier work this paper cites.
Adaptive data augmentation for image classification. In
Alhussein Fawzi, Horst Samulowitz, Deepak Turaga, and Pascal Frossard. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition. In
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Binarized neural networks. In
Itay Hubara, Matthieu Courbariaux, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio. 2016 · 2016
Earlier work this paper cites.
Non-stochastic best arm identification and hyperparameter optimization. In
Kevin Jamieson and Ameet Talwalkar. 2016 · 2016
Earlier work this paper cites.
Fengfu Li, Bo Zhang, and Bin Liu. 2016 · 2016
Earlier work this paper cites.
Pruning Filters for Efficient ConvNets. In
Hao Li, Asim Kadav, Igor Durdanovic, Hanan Samet, and Hans Peter Graf. 2016 · 2016
Earlier work this paper cites.
Pruning Convolutional Neural Networks for Resource Efficient Transfer Learning
Pavlo Molchanov, Stephen Tyree, Tero Karras, Timo Aila, and Jan Kautz. 2016 · 2016
Earlier work this paper cites.
Xnor-net: Imagenet classification using binary convolutional neural networks. In
Mohammad Rastegari, Vicente Ordonez, Joseph Redmon, and Ali Farhadi. 2016 · 2016
Earlier work this paper cites.
Edinburgh neural machine translation systems for wmt 16
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Do deep convolutional nets really need to be deep and convolutional?
Gregor Urban, Krzysztof J Geras, Samira Ebrahimi Kahou, Ozlem Aslan, Shengjie Wang, Rich Caruana, Abdelrahman Mohamed, Matthai Philipose, and Matt Richardson. 2016 · 2016
Earlier work this paper cites.
Sergey Zagoruyko and Nikos Komodakis. 2016 · 2016
Earlier work this paper cites.
Neural architecture search with reinforcement learning
Barret Zoph and Quoc V Le. 2016 · 2016
Earlier work this paper cites.
Structured pruning of deep convolutional neural networks
Sajid Anwar, Kyuyeon Hwang, and Wonyong Sung. 2017 · 2017
Earlier work this paper cites.
Xception: Deep learning with depthwise separable convolutions. In
François Chollet. 2017 · 2017
Earlier work this paper cites.
Learning to prune deep neural networks via layer-wise optimal brain surgeon
Xin Dong, Shangyu Chen, and Sinno Jialin Pan. 2017 · 2017
Earlier work this paper cites.
Google vizier: A service for black-box optimization. In
Daniel Golovin, Benjamin Solnik, Subhodeep Moitra, Greg Kochanski, John Karro, and D Sculley. 2017 · 2017
Cited alongside, same era.
Population based training of neural networks
Max Jaderberg, Valentin Dalibard, Simon Osindero, Wojciech M Czarnecki, Jeff Donahue, Ali Razavi, Oriol Vinyals, Tim Green, Iain Dunning, Karen Simonyan, et al · 2017
Cited alongside, same era.
In-datacenter performance analysis of a tensor processing unit. In
Norman P Jouppi, Cliff Young, Nishant Patil, David Patterson, Gaurav Agrawal, Raminder Bajwa, Sarah Bates, Suresh Bhatia, Nan Boden, Al Borchers, et al · 2017
Cited alongside, same era.
Hyperband: A novel bandit-based approach to hyperparameter optimization
Lisha Li, Kevin Jamieson, Giulia DeSalvo, Afshin Rostamizadeh, and Ameet Talwalkar. 2017 · 2017
Cited alongside, same era.
Advances in pre-training distributed word representations
Tomas Mikolov, Edouard Grave, Piotr Bojanowski, Christian Puhrsch, and Armand Joulin. 2017 · 2017
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, and Michael Auli. 2020 · 2020
Later among the works it cites.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Later among the works it cites.
Big self-supervised models are strong semi-supervised learners
Ting Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi, and Geoffrey Hinton. 2020b · 2020
Later among the works it cites.
Randaugment: Practical automated data augmentation with a reduced search space. In
Ekin D Cubuk, Barret Zoph, Jonathon Shlens, and Quoc V Le. 2020 · 2020
Later among the works it cites.
Fast sparse convnets. In
Erich Elsen, Marat Dukhan, Trevor Gale, and Karen Simonyan. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Projectionnet: Learning efficient on-device deep networks using neural projections
Sujith Ravi. 2017 · 2017
Cited alongside, same era.
Revisiting unreasonable effectiveness of data in deep learning era. In
Chen Sun, Abhinav Shrivastava, Saurabh Singh, and Abhinav Gupta. 2017 · 2017
Cited alongside, same era.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
mixup: Beyond empirical risk minimization
Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz. 2017 · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Cited alongside, same era.
QNNPACK: Open source library for optimized mobile deep learning - Facebook Engineering
Marat Dukhan, Yiming Wu Wu, and Hao Lu. 2020 · 2018
Cited alongside, same era.
The lottery ticket hypothesis: Training pruned neural networks
Jonathan Frankle and Michael Carbin. 2018 · 2018
Cited alongside, same era.
Later among the works it cites.
Rigging the lottery: Making all tickets winners. In
Utku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro, and Erich Elsen. 2020 · 2020
Later among the works it cites.
Characterising bias in compressed models
Sara Hooker, Nyalleng Moorosi, Gregory Clark, Samy Bengio, and Emily Denton. 2020 · 2020
Later among the works it cites.
Setting the learning rate of your neural network
Jeremy Jordan. 2020 · 2020
Later among the works it cites.
TensorFlow 2 MLPerf submissions demonstrate best-in-class performance on Google Cloud
Pankaj Kanwar, Peter Brandt, and Zongwei Zhou. 2021 · 2020
Later among the works it cites.
Automating Data Augmentation: Practice, Theory and New Direction
Sharon Y. Li. 2020 · 2020
Later among the works it cites.
Multi-modal self-supervision from generalized data transformations
Mandela Patrick, Yuki M Asano, Polina Kuznetsova, Ruth Fong, João F Henriques, Geoffrey Zweig, and Andrea Vedaldi. 2020 · 2020
Later among the works it cites.
Amazon SageMaker Automatic Model Tuning: Scalable Black-box Optimization
Valerio Perrone, Huibin Shen, Aida Zolic, Iaroslav Shcherbatyi, Amr Ahmed, Tanya Bansal, Michele Donini, Fela Winkelmolen, Rodolphe Jenatton, Jean Baptiste Faddoul, et al · 2020
Later among the works it cites.
ProFormer: Towards On-Device LSH Projection Based Transformers
Chinnadhurai Sankar, Sujith Ravi, and Zornitsa Kozareva. 2020 · 2020
Later among the works it cites.
Mastering atari, go, chess and shogi by planning with a learned model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, et al · 2020
Later among the works it cites.
Mobilebert: a compact task-agnostic bert for resource-limited devices
Zhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu, Yiming Yang, and Denny Zhou. 2020 · 2020
Later among the works it cites.
Efficient transformers: A survey
Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler. 2020 · 2020
Later among the works it cites.
Towards Accurate Post-training Network Quantization via Bit-Split and Stitching
Peisong Wang, Qiang Chen, Xiangyu He, and Jian Cheng. 2020 · 2020
Later among the works it cites.
Self-training with noisy student improves imagenet classification. In
Qizhe Xie, Minh-Thang Luong, Eduard Hovy, and Quoc V Le. 2020 · 2020
Later among the works it cites.
Hyper-parameter optimization: A review of algorithms and applications
Tong Yu and Hong Zhu. 2020 · 2020
Later among the works it cites.
The Illustrated Transformer
Jay Alammar. 2021 · 2021
Closest in time.
Neural Networks API
Android Developers. 2021 · 2021
Closest in time.
Accelerate
Apple Authors. 2021 · 2021
Closest in time.
Automatic Mixed Precision examples — PyTorch 1.8.1 documentation
PyTorch Authors. 2021a · 2021
Closest in time.
Performance Tuning Guide — PyTorch Tutorials 1.8.1+cu102 documentation
PyTorch Authors. 2021b · 2021
Closest in time.
PyTorch Mobile
PyTorch Authors. 2021c · 2021
Closest in time.
Quantization Recipe — PyTorch Tutorials 1.8.1+cu102 documentation
PyTorch Authors. 2021d · 2021
Closest in time.
torch.jit.script — PyTorch 1.8.1 documentation
PyTorch Authors. 2021e · 2021
Closest in time.
TensorFlow Model Optimization
Tensorflow Authors. 2020 · 2021
Closest in time.
TensorFlow
Tensorflow Authors. 2021g · 2021
Closest in time.
TensorFlow Lite converter
Tensorflow Authors. 2021h · 2021
Closest in time.
XNNPACK backend for TensorFlow Lite
XNNPACK Authors. 2021l · 2021
Closest in time.
The Keras Blog
Francois Chollet. 2020 · 2021
Closest in time.
AVX-512 - Wikipedia
Contributors to Wikimedia projects. 2021a · 2021
Closest in time.
CUDA - Wikipedia
Contributors to Wikimedia projects. 2021b · 2021
Closest in time.
Hyperparameter optimization - Wikipedia
Contributors to Wikimedia projects. 2021c · 2021
Closest in time.
Multiply–accumulate operation - Wikipedia
Contributors to Wikimedia projects. 2021d · 2021
Closest in time.
SSE4 - Wikipedia
Contributors to Wikimedia projects. 2021e · 2021
Closest in time.
WebGL - Wikipedia
Contributors to Wikimedia projects. 2021f · 2021
Closest in time.
google-research
google research. 2021 · 2021
Closest in time.
Neural structured learning: training neural networks with structured signals. In
Arjun Gopalan, Da-Cheng Juan, Cesar Ilharco Magalhaes, Chun-Sung Ferng, Allan Heydon, Chun-Ta Lu, Philip Pham, George Yu, Yicheng Fan, and Yueqi Wang. 2021 · 2021
Closest in time.
Distilling Large Language Models into Tiny and Effective Students using pQRNN
Prabhu Kaliamoorthi, Aditya Siddhant, Edward Li, and Melvin Johnson. 2021 · 2021
Closest in time.
Yann LeCun @EPFL - "Self-supervised learning: could machines learn like humans?"
Yann LeCun. 2018 · 2021
Closest in time.
SIMD ISAs
Arm Ltd. 2021 · 2021
Closest in time.
GTC 2020: Accelerating Sparsity in the NVIDIA Ampere Architecture
NVIDIA. 2020a · 2021
Closest in time.
NVIDIA Embedded Systems for Next-Gen Autonomous Machines
NVIDIA. 2021 · 2021
Closest in time.
Matrix Compression Operator
Rina Panigrahy. 2021 · 2021
Closest in time.
Papers with Code - The latest in Machine Learning
PapersWithCode.com. 2021 · 2021
Closest in time.
Neural Network Intelligence - Microsoft Research
Microsoft Research. 2019 · 2021
Closest in time.
What makes TPUs fine-tuned for deep learning?
Kaz Sato. 2021 · 2021
Closest in time.
Training Neural Networks with Tensor Cores - Dusan Stosic, NVIDIA
Dusan Stosic. 2020 · 2021
Closest in time.
Model optimization
TensorFlow. 2021 · 2021
Closest in time.
BFloat16: The secret to high performance on Cloud TPUs
Shibo Wang and Pankaj Kanwar. 2021 · 2021
Closest in time.