Fetching the paper…
Reading the bibliography…
Conventional hardware-friendly quantization methods, such as fixed-point or integer, tend to perform poorly at very low word sizes as their shrinking dynamic ranges cannot adequately capture the wide data distributions commonly seen in sequence transduction models.
Efficient 8-bit quantization of transformer neural machine language translation model
Bhandare, A., Sripathi, V., Karkada, D., Menon, V., Choi, S., Datta, K., and Saletore, V · 1906
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Fixed-point feedforward deep neural network design using weights +1, 0, and −1
Hwang, K. and Sung, W · 2014
Earlier work this paper cites.
Attention-based models for speech recognition
Chorowski, J., Bahdanau, D., Serdyuk, D., Cho, K., and Bengio, Y · 2015
Earlier work this paper cites.
Binaryconnect: Training deep neural networks with binary weights during propagations
Courbariaux, M., Bengio, Y., and David, J · 2015
Earlier work this paper cites.
Deep learning with limited numerical precision
Gupta, S., Agrawal, A., Gopalakrishnan, K., and Narayanan, P · 2015
Earlier work this paper cites.
Deep compression: Compressing deep neural network with pruning, trained quantization and huffman coding
Han, S., Mao, H., and Dally, W. J · 2015
Earlier work this paper cites.
Fixed point quantization of deep convolutional networks
Lin, D. D., Talathi, S. S., and Annapureddy, V. S · 2015
Earlier work this paper cites.
Quantized convolutional neural networks for mobile devices
Wu, J., Leng, C., Wang, Y., Hu, Q., and Cheng, J · 2015
Earlier work this paper cites.
Ba, L. J., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Binarynet: Training deep neural networks with weights and activations constrained to +1 or -1
Courbariaux, M. and Bengio, Y · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Convolutional neural networks using logarithmic data representation
Miyashita, D., Lee, E. H., and Murmann, B · 2016
Earlier work this paper cites.
Minerva: Enabling low-power, highly-accurate deep neural network accelerators
Reagen, B., Whatmough, P., Adolf, R., Rama, S., Lee, H., Lee, S. K., Hernández-Lobato, J. M., Wei, G., and Brooks, D · 2016
Earlier work this paper cites.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Salimans, T. and Kingma, D. P · 2016
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Wu, Y., Schuster, M., Chen, Z., Le, Q. V., Norouzi, M., Macherey, W., Krikun, M., Cao, Y., Gao, Q., Macherey, K., Klingner, J., Shah, A., Johnson, M., Liu, X., Kaiser, L., Gouws, S., Kato, Y., Kudo, T., Kazawa, H., Stevens, K., Kurian, G., Patil, N., Wang, W., Young, C., Smith, J., Riesa, J., Rudnick, A., Vinyals, O., Corrado, G., Hughes, M., and Dean, J · 2016
Cited alongside, same era.
Dorefa-net: Training low bitwidth convolutional neural networks with low bitwidth gradients
Zhou, S., Ni, Z., Zhou, X., Wen, H., Wu, Y., and Zou, Y · 2016
Cited alongside, same era.
Zhu, C., Han, S., Mao, H., and Dally, W. J · 2016
Cited alongside, same era.
Regularizing deep neural networks by noise: Its interpretation and optimization
Noh, H., You, T., Mun, J., and Han, B · 2017
Later among the works it cites.
Weighted-entropy-based quantization for deep neural networks
Park, E., Ahn, J., and Yoo, S · 2017
Later among the works it cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Later among the works it cites.
Effective quantization approaches for recurrent neural networks
Alom, M. Z., Moody, A. T., Maruyama, N., Essen, B. C. V., and Taha, T. M · 2018
Later among the works it cites.
Deep positron: A deep neural network using the posit number system
Carmichael, Z., Langroudi, S. H. F., Khazanov, C., Lillie, J., Gustafson, J. L., and Kudithipudi, D · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cai, Z., He, X., Sun, J., and Vasconcelos, N · 2017
Cited alongside, same era.
State-of-the-art speech recognition with sequence-to-sequence models
Chiu, C., Sainath, T. N., Wu, Y., Prabhavalkar, R., Nguyen, P., Chen, Z., Kannan, A., Weiss, R. J., Rao, K., Gonina, K., Jaitly, N., Li, B., Chorowski, J., and Bacchiani, M · 2017
Cited alongside, same era.
Beating floating point at its own game: Posit arithmetic
Gustafson and Yonemoto · 2017
Cited alongside, same era.
Quantization and training of neural networks for efficient integer-arithmetic-only inference
Jacob, B., Kligys, S., Chen, B., Zhu, M., Tang, M., Howard, A. G., Adam, H., and Kalenichenko, D · 2017
Cited alongside, same era.
In-datacenter performance analysis of a tensor processing unit
Jouppi, N. P., Young, C., Patil, N., Patterson, D., Agrawal, G., Bajwa, R., Bates, S., Bhatia, S., Boden, N., Borchers, A., Boyle, R., Cantin, P., Chao, C., Clark, C., Coriell, J., Daley, M., Dau, M., Dean, J., Gelb, B., Ghaemmaghami, T. V., Gottipati, R., Gulland, W., Hagmann, R., Ho, C. R., Hogberg, D., Hu, J., Hundt, R., Hurt, D., Ibarz, J., Jaffey, A., Jaworski, A., Kaplan, A., Khaitan, H., Killebrew, D., Koch, A., Kumar, N., Lacy, S., Laudon, J., Law, J., Le, D., Leary, C., Liu, Z., Lucke, K., Lundin, A., MacKean, G., Maggiore, A., Mahony, M., Miller, K., Nagarajan, R., Narayanaswami, R., Ni, R., Nix, K., Norrie, T., Omernick, M., Penukonda, N., Phelps, A., Ross, J., Ross, M., Salek, A., Samadiani, E., Severn, C., Sizikov, G., Snelham, M., Souter, J., Steinberg, D., Swing, A., Tan, M., Thorson, G., Tian, B., Toma, H., Tuttle, E., Vasudevan, V., Walter, R., Wang, W., Wilcox, E., and Yoon, D. H · 2017
Cited alongside, same era.
Opennmt: Open-source toolkit for neural machine translation
Klein, G., Kim, Y., Deng, Y., Senellart, J., and Rush, A. M · 2017
Cited alongside, same era.
Flexpoint: An adaptive numerical format for efficient training of deep neural networks
Köster, U., Webb, T., Wang, X., Nassar, M., Bansal, A. K., Constable, W., Elibol, O., Hall, S., Hornof, L., Khosrowshahi, A., Kloss, C., Pai, R. J., and Rao, N · 2017
Cited alongside, same era.
Lognet: Energy-efficient neural networks using logarithmic computation
Lee, E. H., Miyashita, D., Chai, E., Murmann, B., and Wong, S. S · 2017
Cited alongside, same era.
Later among the works it cites.
End-to-end DNN training with block floating point arithmetic
Drumond, M., Lin, T., Jaggi, M., and Falsafi, B · 2018
Later among the works it cites.
A configurable cloud-scale dnn processor for real-time ai
Fowers, J., Ovtcharov, K., Papamichael, M., Massengill, T., Liu, M., Lo, D., Alkalay, S., Haselman, M., Adams, L., Ghandi, M., et al · 2018
Later among the works it cites.
Rethinking floating point for deep learning
Johnson, J · 2018
Later among the works it cites.
A modular digital vlsi flow for high-productivity soc design
Khailany, B., Khmer, E., Venkatesan, R., Clemons, J., Emer, J. S., Fojtik, M., Klinefelter, A., Pellauer, M., Pinckney, N., Shao, Y. S., Srinath, S., Torng, C., Xi, S. L., Zhang, Y., and Zimmer, B · 2018
Later among the works it cites.
Energy-efficient neural network accelerator based on outlier-aware low-precision computation
Park, E., Kim, D., and Yoo, S · 2018
Later among the works it cites.
Efficient hardware acceleration of cnns using logarithmic data representation with arbitrary log-base
Vogel, S., Liang, M., Guntoro, A., Stechele, W., and Ascheid, G · 2018
Later among the works it cites.
Lq-nets: Learned quantization for highly accurate and compact deep neural networks
Zhang, D., Yang, J., Ye, D., and Hua, G · 2018
Later among the works it cites.
Accurate and efficient 2-bit quantized neural networks
Choi, J., Venkataramani, S., Srinivasan, V., Gopalakrishnan, K., Wang, Z., and Chuang, P · 2019
Closest in time.