Fetching the paper…
Reading the bibliography…
There has been a rapid advance of custom hardware (HW) for accelerating the inference speed of deep neural networks (DNNs).
Efficient hardware architecture of softmax layer in deep neural network
B. Yuan · 2016
Earlier work this paper cites.
OpenNMT: Open-source toolkit for neural machine translation
Guillaume Klein, Yoon Kim, Yuntian Deng, Jean Senellart, and Alexander M. Rush · 2017
Earlier work this paper cites.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Tvt: Two-view transformer network for video captioning
Ming Chen, Yingming Li, Zhongfei Zhang, and Siyu Huang · 2018
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Bottom-up abstractive summarization
Sebastian Gehrmann, Yuntian Deng, and Alexander Rush · 2018
Earlier work this paper cites.
Streaming end-to-end speech recognition for mobile devices, 2018
Yanzhang He, Tara N. Sainath, Rohit Prabhavalkar, Ian McGraw, Raziel Alvarez, Ding Zhao, David Rybach, Anjuli Kannan, Yonghui Wu, Ruoming Pang, Qiao Liang, Deepti Bhatia, Yuan Shangguan, Bo Li, Golan Pundak, Khe Chai Sim, Tom Bagby, Shuo yiin Chang, Kanishka Rao, and Alexander Gruenstein · 2018
Earlier work this paper cites.
Efficient hardware architecture of softmax layer in deep neural network
R. Hu, B. Tian, S. Yin, and S. Wei · 2018
Earlier work this paper cites.
AI benchmark: Running deep neural networks on android smartphones
Andrey Ignatov, Radu Timofte, William Chou, Ke Wang, Max Wu, Tim Hartley, and Luc Van Gool · 2018
Earlier work this paper cites.
Efficient fpga implementation of softmax function for dnn applications
Z. Li, H. Li, X. Jiang, B. Chen, Y. Zhang, and G. Du · 2018
Earlier work this paper cites.
Mcdram: Low latency and energy-efficient matrix computations in dram
Hyunsung Shin, Dongyoung Kim, Eunhyeok Park, Sungho Park, Yongsik Park, and Sungjoo Yoo · 2018
Earlier work this paper cites.
A high speed softmax vlsi architecture based on basic-split
Q. Sun, Z. Di, Z. Lv, F. Song, Q. Xiang, Q. Feng, Y. Fan, X. Yu, and W. Wang · 2018
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman · 2018
Earlier work this paper cites.
A high-speed and low-complexity architecture for softmax function in deep learning
M. Wang, S. Lu, D. Zhu, J. Lin, and Z. Wang · 2018
Cited alongside, same era.
Efficient 8-bit quantization of transformer neural machine language translation model
Aishwarya Bhandare, Vamsi Sripathi, Deepthi Karkada, Vivek Menon, Sun Choi, Kushal Datta, and Vikram Saletore · 2019
Cited alongside, same era.
Softmax optimizations for intel xeon processor-based platforms
Jacek Czaja, Michal Gallus, Tomasz Patejko, and Jian Tang · 2019
Cited alongside, same era.
Hardware-aware softmax approximation for deep neural networks
Xue Geng, Jie Lin, Bin Zhao, Anmin Kong, Mohamed M. Sabry Aly, and Vijay Chandrasekhar · 2019
Cited alongside, same era.
Deep learning for symbolic mathematics
Guillaume Lample and François Charton · 2019
Q8bert: Quantized 8bit bert, 2019
Ofir Zafrir, Guy Boudoukh, Peter Izsak, and Moshe Wasserblat · 2019
Later among the works it cites.
Openei: An open framework for edge intelligence
Xingzhou Zhang, Yifan Wang, Sidi Lu, Liangkai Liu, Lanyu Xu, and Weisong Shi · 2019
Later among the works it cites.
End-to-end object detection with transformers, 2020
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale, 2020
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2020
Later among the works it cites.
Hardware implementation of a softmax-like function for deep learning
Ioannis Kouretas and Vassilis Paliouras · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Albert: A lite bert for self-supervised learning of language representations, 2019
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut · 2019
Cited alongside, same era.
Roberta: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Cited alongside, same era.
Machine learning at the network edge: A survey, 2019
M. G. Sarwar Murshed, Christopher Murphy, Daqing Hou, Nazar Khan, Ganesh Ananthanarayanan, and Faraz Hussain · 2019
Cited alongside, same era.
Fully quantized transformer for improved translation, 2019
Gabriele Prato, Ella Charlaix, and Mehdi Rezagholizadeh · 2019
Cited alongside, same era.
Q-bert: Hessian based ultra low precision quantization of bert, 2019
Sheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma, Zhewei Yao, Amir Gholami, Michael W. Mahoney, and Kurt Keutzer · 2019
Cited alongside, same era.
A customized convolutional neural network design using improved softmax layer for real-time human emotion recognition
K. Wang, Y. Huang, Y. Ho, and W. Fang · 2019
Cited alongside, same era.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime G. Carbonell, Ruslan Salakhutdinov, and Quoc V. Le · 2019
Cited alongside, same era.
Towards fast and accurate streaming end-to-end asr, 2020
Bo Li, Shuo yiin Chang, Tara N. Sainath, Ruoming Pang, Yanzhang He, Trevor Strohman, and Yonghui Wu · 2020
Later among the works it cites.
Molecule attention transformer
Lukasz Maziarka, Tomasz Danel, Slawomir Mucha, Krzysztof Rataj, Jacek Tabor, and Stanislaw Jastrzebski · 2020
Later among the works it cites.
A streaming on-device end-to-end model surpassing server-side conventional model quality and latency, 2020
Tara N. Sainath, Yanzhang He, Bo Li, Arun Narayanan, Ruoming Pang, Antoine Bruguier, Shuo yiin Chang, Wei Li, Raziel Alvarez, Zhifeng Chen, Chung-Cheng Chiu, David Garcia, Alex Gruenstein, Ke Hu, Minho Jin, Anjuli Kannan, Qiao Liang, Ian McGraw, Cal Peyser, Rohit Prabhavalkar, Golan Pundak, David Rybach, Yuan Shangguan, Yash Sheth, Trevor Strohman, Mirko Visontai, Yonghui Wu, Yu Zhang, and Ding Zhao · 2020
Later among the works it cites.
Neural network inference on mobile socs
Siqi Wang, Anuj Pathania, and Tulika Mitra · 2020
Later among the works it cites.
Softermax: Hardware/software co-design of an efficient softmax for transformers
Jacob R. Stevens, Rangharajan Venkatesan, Steve Dai, Brucek Khailany, and Anand Raghunathan · 2021
Closest in time.
Training data-efficient image transformers & distillation through attention, 2021
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou · 2021
Closest in time.
Lightweight approximation of softmax layer for on-device inference
Ihor Vasyltsov and Wooseok Chang · 2021
Closest in time.