Fetching the paper…
Reading the bibliography…
Efficient machine learning implementations optimized for inference in hardware have wide-ranging benefits, depending on the application, from lower inference latency to higher data throughput and reduced energy consumption.
A mathematical theory of communication
Shannon, C. E. (1948) · 1948
Earlier work this paper cites.
Curve fitting and optimal design for prediction
O’Hagan, A. (1978) · 1978
Earlier work this paper cites.
Efficient global optimization of expensive black-box functions
Jones, D. R., Schonlau, M., and Welch, W. J. (1998) · 1998
Earlier work this paper cites.
Feature selection, L1 vs. L2 regularization, and rotational invariance
Ng, A. Y. (2004) · 2004
Earlier work this paper cites.
Assessing intelligence in artificial neural networks
Schaub, N. J. and Hotaling, N. (2020) · 2006
Earlier work this paper cites.
Rectified linear units improve restricted Boltzmann machines
Nair, V. and Hinton, G. E. (2010) · 2010
Earlier work this paper cites.
Bayesian Gaussian processes for sequential prediction, optimisation and quadrature
Osborne, M. A. (2010) · 2010
Earlier work this paper cites.
Deep sparse rectifier neural networks
Glorot, X., Bordes, A., and Bengio, Y. (2011) · 2011
Earlier work this paper cites.
Improving the speed of neural networks on CPUs
Vanhoucke, V., Senior, A., and Mao, M. Z. (2011) · 2011
Earlier work this paper cites.
Compressing deep convolutional networks using vector quantization
Gong, Y., Liu, L., Yang, M., and Bourdev, L. D. (2014) · 2014
Earlier work this paper cites.
BinaryConnect: Training deep neural networks with binary weights during propagations
Courbariaux, M., Bengio, Y., and David, J.-P. (2015) · 2015
Earlier work this paper cites.
Deep learning with limited numerical precision
Gupta, S., Agrawal, A., Gopalakrishnan, K., and Narayanan, P. (2015) · 2015
Earlier work this paper cites.
Learning both weights and connections for efficient neural networks
Han, S., Pool, J., Tran, J., and Dally, W. J. (2015) · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C. (2015) · 2015
Earlier work this paper cites.
Deep compression: Compressing deep neural networks with pruning, trained quantization and Huffman coding
Han, S., Mao, H., and Dally, W. J. (2016) · 2016
Earlier work this paper cites.
Binarized neural networks
Hubara, I., Courbariaux, M., Soudry, D., El-Yaniv, R., and Bengio, Y. (2016) · 2016
Earlier work this paper cites.
Li, F. and Liu, B. (2016) · 2016
Earlier work this paper cites.
Deep neural networks are robust to weight binarization and other non-linear distortions
Merolla, P., Appuswamy, R., Arthur, J. V., Esser, S. K., and Modha, D. S. (2016) · 2016
Earlier work this paper cites.
Xnor-net: Imagenet classification using binary convolutional neural networks
Rastegari, M., Ordonez, V., Redmon, J., and Farhadi, A. (2016b) · 2016
Earlier work this paper cites.
Quantized convolutional neural networks for mobile devices
Wu, J., Leng, C., Wang, Y., Hu, Q., and Cheng, J. (2016) · 2016
Earlier work this paper cites.
DoReFa-Net: Training low bitwidth convolutional neural networks with low bitwidth gradients
Zhou, S., Wu, Y., Ni, Z., Zhou, X., Wen, H., and Zou, Y. (2016) · 2016
Earlier work this paper cites.
A survey of model compression and acceleration for deep neural networks
Cheng, Y., Wang, D., Zhou, P., and Zhang, T. (2018) · 2017
Cited alongside, same era.
Minimum energy quantized neural networks
Moons, B., Goetschalckx, K., Berckelaer, N. V., and Verhelst, M. (2017) · 2017
Cited alongside, same era.
SVCCA: Singular vector canonical correlation analysis for deep learning dynamics and interpretability
Raghu, M., Gilmer, J., Yosinski, J., and Sohl-Dickstein, J. (2017) · 2017
Cited alongside, same era.
FINN: A framework for fast, scalable binarized neural network inference
Umuroglu, Y., Fraser, N. J., Gambardella, G., Blott, M., Leong, P., Jahre, M., et al. (2017) · 2017
Cited alongside, same era.
FINN-R: An end-to-end deep-learning framework for fast exploration of quantized neural networks
Blott, M., Preusser, T., Fraser, N., Gambardella, G., O’Brien, K., and Umuroglu, Y. (2018) · 2018
Cited alongside, same era.
Improving neural network quantization without retraining using outlier channel splitting
Zhao, R., Hu, Y., Dotzel, J., Sa, C. D., and Zhang, Z. (2019) · 2019
Later among the works it cites.
Deconstructing lottery tickets: Zeros, signs, and the supermask
Zhou, H., Lan, J., Liu, R., and Yosinski, J. (2019) · 2019
Later among the works it cites.
BoTorch: Programmable Bayesian optimization in PyTorch
Balandat, M., Karrer, B., Jiang, D., Daulton, S., Letham, B., Wilson, A. G., et al. (2020) · 2020
Later among the works it cites.
What is the state of neural network pruning?
Blalock, D., Ortiz, J. J. G., Frankle, J., and Guttag, J. (2020) · 2020
Later among the works it cites.
A comprehensive survey on model compression and acceleration
Choudhary, T., Mishra, V., Goswami, A., and Sarangapani, J. (2020) · 2020
Later among the works it cites.
Differentiable expected hypervolume improvement for parallel multi-objective Bayesian optimization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Coleman, E., Freytsis, M., Hinzmann, A., Narain, M., Thaler, J., Tran, N., et al. (2018) · 2018
Cited alongside, same era.
Fast inference of deep neural networks in FPGAs for particle physics
Duarte, J., Han, S., Harris, P., Jindariani, S., Kreinar, E., Kreis, B., et al. (2018) · 2018
Cited alongside, same era.
Quantized neural networks: Training neural networks with low precision weights and activations
Hubara, I., Courbariaux, M., Soudry, D., El-Yaniv, R., and Bengio, Y. (2018) · 2018
Cited alongside, same era.
Quantization and training of neural networks for efficient integer-arithmetic-only inference
Jacob, B., Kligys, S., Chen, B., Zhu, M., Tang, M., Howard, A., et al. (2018) · 2018
Cited alongside, same era.
Learning Sparse Neural Networks through L 0 L_{0} Regularization
Louizos, C., Welling, M., and Kingma, D. P. (2018) · 2018
Cited alongside, same era.
Mixed precision training
Micikevicius, P., Narang, S., Alben, J., Diamos, G. F., Elsen, E., García, D., et al. (2018) · 2018
Cited alongside, same era.
How does batch normalization help optimization?
Santurkar, S., Tsipras, D., Ilyas, A., and Madry, A. (2018) · 2018
Cited alongside, same era.
Daulton, S., Balandat, M., and Bakshy, E. (2020) · 2020
Later among the works it cites.
Model compression and hardware acceleration for neural networks: A comprehensive survey
Deng, L., Li, G., Han, S., Shi, L., and Xie, Y. (2020) · 2020
Later among the works it cites.
HAWQ-V2: Hessian aware trace-weighted quantization of neural networks
Dong, Z., Yao, Z., Cai, Y., Arfeen, D., Gholami, A., Mahoney, M. W., et al. (2020) · 2020
Later among the works it cites.
Quantized guided pruning for efficient hardware implementations of convolutional neural networks
Hacene, G. B., Gripon, V., Arzel, M., Farrugia, N., and Bengio, Y. (2020) · 2020
Later among the works it cites.
Trained quantization thresholds for accurate and efficient fixed-point inference of deep neural networks 2, 112
Jain, S. R., Gural, A., Wu, M., and Dick, C. H. (2020) · 2020
Later among the works it cites.
Compressing deep neural networks on FPGAs to binary and ternary precision with hls4ml
Loncar, V. et al. (2020) · 2020
Later among the works it cites.
JEDI-net: a jet identification algorithm based on interaction networks
Moreno, E. A., Cerri, O., Duarte, J. M., Newman, H. B., Nguyen, T. Q., Periwal, A., et al. (2020) · 2020
Later among the works it cites.
Compressing deep neural networks on FPGAs to binary and ternary precision with hls4ml
Ngadiuba, J., Guglielmo, G. D., Duarte, J., Harris, P., Hoang, D., Jindariani, S., et al. (2020) · 2020
Later among the works it cites.
brevitas
Pappalardo, A. (2020) · 2020
Later among the works it cites.
Comparing rewinding and fine-tuning in neural network pruning
Renda, A., Frankle, J., and Carbin, M. (2020) · 2020
Later among the works it cites.
Efficient processing of deep neural networks
Sze, V., Chen, Y.-H., Yang, T.-J., and Emer, J. S. (2020) · 2020
Later among the works it cites.
Bayesian bits: Unifying quantization and pruning
van Baalen, M., Louizos, C., Nagel, M., Amjad, R. A., Wang, Y., Blankevoort, T., et al. (2020) · 2020
Later among the works it cites.
UNIQ: Uniform noise injection for the quantization of neural networks
Baskin, C., Liss, N., Schwartz, E., Zheltonozhskii, E., Giryes, R., Bronstein, A. M., et al. (2021) · 2021
Closest in time.
Mix and match: A novel FPGA-centric deep neural network quantization framework
Chang, S.-E., Li, Y., Sun, M., Shi, R., So, H. K. H., Qian, X., et al. (2021) · 2021
Closest in time.
Automatic heterogeneous quantization of deep neural networks for low-latency inference on the edge for particle detectors
Coelho, C. N., Kuusela, A., Li, S., Zhuang, H., Ngadiuba, J., Aarrestad, T. K., et al. (2021) · 2021
Closest in time.
Early-stage neural network hardware performance analysis
Karbachevsky, A., Baskin, C., Zheltonozhskii, E., Yermolin, Y., Gabbay, F., Bronstein, A. M., et al. (2021) · 2021
Closest in time.