Fetching the paper…
Reading the bibliography…
Transformer based architectures have become de-facto models used for a range of Natural Language Processing tasks.
“Optimal brain damage”
Yann LeCun, John Denker and Sara Solla · 1990
Earlier work this paper cites.
“Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition”
Erik Sang and Fien De · 2003
Earlier work this paper cites.
“Estimating or propagating gradients through stochastic neurons for conditional computation”
Yoshua Bengio, Nicholas L\’eonard and Aaron Courville · 2013
Earlier work this paper cites.
“Recursive deep models for semantic compositionality over a sentiment treebank”
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher Manning, Andrew Ng and Christopher Potts · 2013
Earlier work this paper cites.
“Do deep nets really need to be deep?”
Jimmy Ba and Rich Caruana · 2014
Earlier work this paper cites.
“Binaryconnect: Training deep neural networks with binary weights during propagations”
Matthieu Courbariaux, Yoshua Bengio and Jean-Pierre David · 2015
Earlier work this paper cites.
“Distilling the knowledge in a neural network”
Geoffrey Hinton, Oriol Vinyals and Jeff Dean · 2015
Earlier work this paper cites.
“SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and< 0.5 MB model size”
Forrest Iandola, Song Han, Matthew Moskewicz, Khalid Ashraf, William Dally and Kurt Keutzer · 2016
Earlier work this paper cites.
Fengfu Li, Bo Zhang and Bin Liu · 2016
Earlier work this paper cites.
“Pruning filters for efficient convnets”
Hao Li, Asim Kadav, Igor Durdanovic, Hanan Samet and Hans Graf · 2016
Earlier work this paper cites.
“Squad: 100,000+ questions for machine comprehension of text”
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev and Percy Liang · 2016
Earlier work this paper cites.
“Xnor-net: Imagenet classification using binary convolutional neural networks”
Mohammad Rastegari, Vicente Ordonez, Joseph Redmon and Ali Farhadi · 2016
Earlier work this paper cites.
“Dorefa-net: Training low bitwidth convolutional neural networks with low bitwidth gradients”
Shuchang Zhou, Yuxin Wu, Zekun Ni, Xinyu Zhou, He Wen and Yuheng Zou · 2016
Earlier work this paper cites.
“Mobilenets: Efficient convolutional neural networks for mobile vision applications”
Andrew Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto and Hartwig Adam · 2017
Earlier work this paper cites.
“Quantized neural networks: Training neural networks with low precision weights and activations”
Itay Hubara, Matthieu Courbariaux, Daniel Soudry, Ran El-Yaniv and Yoshua Bengio · 2017
Earlier work this paper cites.
“Attention is all you need”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan Gomez, ukasz Kaiser and Illia Polosukhin · 2017
Cited alongside, same era.
“A broad-coverage challenge corpus for sentence understanding through inference”
Adina Williams, Nikita Nangia and Samuel Bowman · 2017
Cited alongside, same era.
“Incremental network quantization: Towards lossless cnns with low-precision weights”
Aojun Zhou, Anbang Yao, Yiwen Guo, Lin Xu and Yurong Chen · 2017
Cited alongside, same era.
“Pact: Parameterized clipping activation for quantized neural networks”
Jungwook Choi, Zhuo Wang, Swagath Venkataramani, Pierce I-Jen Chuang, Vijayalakshmi Srinivasan and Kailash Gopalakrishnan · 2018
Cited alongside, same era.
“Squeezenext: Hardware-aware neural network design”
Amir Gholami, Kiseok Kwon, Bichen Wu, Zizheng Tai, Xiangyu Yue, Peter Jin, Sicheng Zhao and Kurt Keutzer · 2018
“Efficient 8-Bit Quantization of Transformer Neural Machine Language Translation Model”
Aishwarya Bhandare, Vamsi Sripathi, Deepthi Karkada, Vivek Menon, Sun Choi, Kushal Datta and Vikram Saletore · 2019
Closest in time.
“What Does BERT Look At? An Analysis of BERT’s Attention”
Kevin Clark, Urvashi Khandelwal, Omer Levy and Christopher Manning · 2019
Closest in time.
“BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”
Jacob Devlin, Ming-Wei Chang, Kenton Lee and Kristina Toutanova · 2019
Closest in time.
“HAWQ: Hessian AWare Quantization of Neural Networks with Mixed-Precision”
Zhen Dong, Zhewei Yao, Amir Gholami, Michael Mahoney and Kurt Keutzer · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference”
Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew Howard, Hartwig Adam and Dmitry Kalenichenko · 2018
Cited alongside, same era.
“Quantizing deep convolutional networks for efficient inference: A whitepaper”
Raghuraman Krishnamoorthi · 2018
Cited alongside, same era.
“Glue: A multi-task benchmark and analysis platform for natural language understanding”
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy and Samuel Bowman · 2018
Cited alongside, same era.
“HitNet: hybrid ternary recurrent neural network”
Peiqi Wang, Xinfeng Xie, Lei Deng, Guoqi Li, Dongsheng Wang and Yuan Xie · 2018
Cited alongside, same era.
“Mixed Precision Quantization of ConvNets via Differentiable Neural Architecture Search”
Bichen Wu, Yanghan Wang, Peizhao Zhang, Yuandong Tian, Peter Vajda and Kurt Keutzer · 2018
Cited alongside, same era.
“Alternating multi-bit quantization for recurrent neural networks”
Chen Xu, Jianqiang Yao, Zhouchen Lin, Wenwu Ou, Yuanbin Cao, Zhirong Wang and Hongbin Zha · 2018
Cited alongside, same era.
“Hessian-based Analysis of Large Batch Training and Robustness to Adversaries”
Zhewei Yao, Amir Gholami, Qi Lei, Kurt Keutzer and Michael. Mahoney · 2018
Cited alongside, same era.
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer and Veselin Stoyanov · 2019
Closest in time.
“A Tensorized Transformer for Language Modeling”
Xindian Ma, Peng Zhang, Shuai Zhang, Nan Duan, Yuexian Hou, Dawei Song and Ming Zhou · 2019
Closest in time.
“Are Sixteen Heads Really Better than One?”
Paul Michel, Omer Levy and Graham Neubig · 2019
Closest in time.
“Language Models are Unsupervised Multitask Learners”, 2019
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei and Ilya Sutskever · 2019
Closest in time.
“Patient Knowledge Distillation for BERT Model Compression”
Siqi Sun, Cheng Yu, Gan Zhe and Liu Jingjing · 2019
Closest in time.
“Distilling Task-Specific Knowledge from BERT into Simple Neural Networks”
Raphael Tang, Yao Lu, Linqing Liu, Lili Mou, Olga Vechtomova and Jimmy Lin · 2019
Closest in time.
“Lightweight and Efficient Neural Natural Language Processing with Quaternion Networks”
Yi Tay, Aston Zhang, Luu Tuan, Jinfeng Rao, Shuai Zhang, Shuohang Wang, Jie Fu and Siu Hui · 2019
Closest in time.
“HAQ: Hardware-Aware Automated Quantization with Mixed Precision”
Kuan Wang, Zhijian Liu, Yujun Lin, Ji Lin and Song Han · 2019
Closest in time.
“Fbnet: Hardware-aware efficient convnet design via differentiable neural architecture search”
Bichen Wu, Xiaoliang Dai, Peizhao Zhang, Yanghan Wang, Fei Sun, Yiming Wu, Yuandong Tian, Peter Vajda, Yangqing Jia and Kurt Keutzer · 2019
Closest in time.
“XLNet: Generalized Autoregressive Pretraining for Language Understanding”
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov and Quoc Le · 2019
Closest in time.