Fetching the paper…
Reading the bibliography…
The amounts of data that need to be transmitted, processed, and stored by the modern deep neural networks have reached truly enormous volumes in the last few years calling for the invention of new paradigms both in hardware and software development.
S. N. Bernstein, “On a modification of Chebyshev’s inequality and of the error formula of Laplace,” Ann. Sci. Inst. Sav. Ukraine, Sect. Math , vol. 1, no. 4, pp. 38–49, 1924
1924
Earlier work this paper cites.
M. Fréchet, “Sur la loi de probabilité de l’écart maximum,” Ann. Soc. Math. Polon. , vol. 6, pp. 93–116, 1927
1927
Earlier work this paper cites.
R. A. Fisher and L. H. C. Tippett, “Limiting forms of the frequency distribution of the largest or smallest member of a sample,” Mathematical Proceedings of the Cambridge Philosophical Society , vol. 24, no. 2, pp. 180–190, 1928
1928
Earlier work this paper cites.
R. Mises, “La distribution de la plus grande de n valeurs,” Rev. Math. Union Interbalcanique , vol. 1, pp. 141–160, 1936
1936
Earlier work this paper cites.
B. Gnedenko, “Sur la distribution limite du terme maximum d’une serie aleatoire,” Annals of Mathematics , pp. 423–453, 1943
1943
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in Neural Information Processing Systems , vol. 30, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
R. Vershynin, “High-dimensional probability: An introduction with applications in data science,” Cambridge University Press , vol. 47, 2018
2018
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” North American Chapter of the Association for Computational Linguistics: Human Language Technologies , 2019
2019
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI Blog , vol. 1, no. 8, p. 9, 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
O. Zafrir, G. Boudoukh, P. Izsak, and M. Wasserblat, “Q8bert: Quantized 8bit BERT,” Workshop on Energy Efficient Machine Learning and Cognitive Computing-NeurIPS Edition , pp. 36–39, 2019
2019
Cited alongside, same era.
A. H. Zadeh, I. Edo, O. M. Awad, and A. Moshovos, “GOBO: Quantizing attention-based NLP models for low latency and energy efficient inference,” IEEE/ACM International Symposium on Microarchitecture , pp. 811–824, 2020
2020
Cited alongside, same era.
T. Lin, Y. Wang, X. Liu, and X. Qiu, “A survey of transformers,” arXiv:2106.04554 , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
A. Srinivas, T.-Y. Lin, N. Parmar, J. Shlens, P. Abbeel, and A. Vaswani, “Bottleneck transformers for visual recognition,” IEEE/CVF Conference on Computer Vision and Pattern Recognition , pp. 16 519–16 529, 2021
2021
Later among the works it cites.
S. Dai, R. Venkatesan, M. Ren, B. Zimmer, W. Dally, and B. Khailany, “VS-Quant: Per-vector scaled quantization for accurate low-precision neural network inference,” Proceedings of Machine Learning and Systems , vol. 3, pp. 873–884, 2021
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Shen, Z. Dong, J. Ye, L. Ma, Z. Yao, A. Gholami, M. W. Mahoney, and K. Keutzer, “Q-bert: Hessian based ultra low precision quantization of BERT,” Proceedings of the AAAI Conference on Artificial Intelligence , vol. 34, no. 05, pp. 8815–8821, 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
B. Rouhani, D. Lo, R. Zhao, M. Liu, J. Fowers et al. , “Pushing the limits of narrow precision inferencing at cloud scale with Microsoft floating point,” Advances in Neural Information Processing Systems , vol. 33, pp. 10 271–10 281, 2020
2020
Cited alongside, same era.
G. Franchi, A. Bursuc, E. Aldea, S. Dubuisson, and I. Bloch, “TRADI: Tracking deep neural network weight distributions,” European Conference on Computer Vision , pp. 105–121, 2020
2020
Cited alongside, same era.
A. C. Elster and T. A. Haugdahl, “NVIDIA Hopper GPU and Grace CPU highlights,” Computing in Science and Engineering , vol. 24, no. 2, pp. 95–100, 2022
2022
Closest in time.
G. Pang, “The AI chip race,” IEEE Intelligent Systems , vol. 37, no. 2, pp. 111–112, 2022
2022
Closest in time.
S. Huang, E. Tang, S. Li, X. Ping, and R. Chen, “Hardware-friendly compression and hardware acceleration for transformer: A survey,” Electronic Research Archive , vol. 30, no. 10, pp. 3755–3785, 2022
2022
Closest in time.
I. Lyubomirsky and X. Wang, “Block floating point (BFP) for efficient deep neural net inference,” IEEE P3109 Working Group, June 6 , 2022
2022
Closest in time.