Fetching the paper…
Reading the bibliography…
As NLP models become larger, executing a trained model requires significant computational resources incurring monetary and environmental costs.
The state of sparsity in deep neural networks
Trevor Gale, Erich Elsen, and Sara Hooker. 2019 · 1902
Earlier work this paper cites.
Distilling task-specific knowledge from BERT into simple neural networks
Raphael Tang, Yao Lu, Linqing Liu, Lili Mou, Olga Vechtomova, and Jimmy Lin. 2019 · 1903
Earlier work this paper cites.
WinoGrande: An adversarial winograd schema challenge at scale
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2019 · 1907
Earlier work this paper cites.
Roy Schwartz, Jesse Dodge, Noah A. Smith, and Oren Etzioni. 2019 · 1907
Earlier work this paper cites.
TinyBERT: Distilling BERT for natural language understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2019 · 1909
Earlier work this paper cites.
Q-BERT: Hessian based ultra low precision quantization of BERT
Sheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma, Zhewei Yao, Amir Gholami, Michael W. Mahoney, and Kurt Keutzer. 2019 · 1909
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2019 · 1910
Earlier work this paper cites.
The comparison and evaluation of forecasters
Morris H. DeGroot and Stephen E. Fienberg. 1983 · 1983
Earlier work this paper cites.
Optimal brain damage
Yann LeCun, John S. Denker, and Sara A. Solla. 1990 · 1990
Earlier work this paper cites.
Stacked generalization
David H. Wolpert. 1992 · 1992
Earlier work this paper cites.
Efficient parsing for bilexical context-free grammars and head automaton grammars
Jason Eisner and Giorgio Satta. 1999 · 1999
Earlier work this paper cites.
Controlling computation versus quality for neural sequence models
Ankur Bapna, Naveen Arivazhagan, and Orhan Firat. 2020 · 2002
Earlier work this paper cites.
Balancing cost and benefit with tied-multi transformers
Raj Dabre, Raphael Rubino, and Atsushi Fujita. 2020 · 2002
Earlier work this paper cites.
Calibration of pre-trained transformers
Shrey Desai and Greg Durrett. 2020 · 2003
Earlier work this paper cites.
Measures of diversity in classifier ensembles and their relationship with the ensemble accuracy
Ludmila I. Kuncheva and Christopher J. Whitaker. 2003 · 2003
Earlier work this paper cites.
The PASCAL recognising textual entailment challenge
Ido Dagan, Oren Glickman, and Bernardo Magnini. 2005 · 2005
Earlier work this paper cites.
An efficient algorithm for easy-first non-directional dependency parsing
Yoav Goldberg and Michael Elhadad. 2010 · 2010
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011 · 2011
Cited alongside, same era.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013 · 2013
Cited alongside, same era.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2014 · 2014
Cited alongside, same era.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Cited alongside, same era.
Adaptive computation time for recurrent neural networks
Alex Graves. 2016 · 2016
Cited alongside, same era.
Sequence-level knowledge distillation
Hypothesis Only Baselines in Natural Language Inference
Adam Poliak, Jason Naradowsky, Aparajita Haldar, Rachel Rudinger, and Benjamin Van Durme. 2018 · 2018
Later among the works it cites.
Neural speed reading via skim-RNN
Minjoon Seo, Sewon Min, Ali Farhadi, and Hannaneh Hajishirzi. 2018 · 2018
Later among the works it cites.
Multi-mention learning for reading comprehension with neural cascades
Swabha Swayamdipta, Ankur P. Parikh, and Tom Kwiatkowski. 2018 · 2018
Later among the works it cites.
SkipNet: Learning dynamic routing in convolutional networks
Xin Wang, Fisher Yu, Zi-Yi Dou, Trevor Darrell, and Joseph E. Gonzalez. 2018 · 2018
Later among the works it cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel R Bowman. 2018 · 2018
Later among the works it cites.
Convolutional neural network compression for natural language processing
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yoon Kim and Alexander M. Rush. 2016 · 2016
Cited alongside, same era.
Adaptive neural networks for efficient inference
Tolga Bolukbasi, Joseph Wang, Ofer Dekel, and Venkatesh Saligrama. 2017 · 2017
Cited alongside, same era.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. 2017 · 2017
Cited alongside, same era.
On fairness and calibration
Geoff Pleiss, Manish Raghavan, Felix Wu, Jon Kleinberg, and Kilian Q. Weinberger. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
AI and compute
Dario Amodei and Danny Hernandez. 2018 · 2018
Cited alongside, same era.
AllenNLP: A deep semantic natural language processing platform
Matt Gardner, Joel Grus, Mark Neumann, Oyvind Tafjord, Pradeep Dasigi, Nelson F. Liu, Matthew Peters, Michael Schmitz, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
Krzysztof Wróbel, Marcin Pietroń, Maciej Wielgosz, Michał Karwatowski, and Kazimierz Wiatr. 2018 · 2018
Later among the works it cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015 · 2018
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Are sixteen heads really better than one?
Paul Michel, Omer Levy, and Graham Neubig. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Later among the works it cites.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 2019
Later among the works it cites.
An empirical study of example forgetting during deep neural network learning
Mariya Toneva, Alessandro Sordoni, Remi Tachet des Combes, Adam Trischler, Yoshua Bengio, and Geoffrey J. Gordon. 2019 · 2019
Later among the works it cites.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019 · 2019
Later among the works it cites.
Q8BERT: Quantized 8bit BERT
Ofir Zafrir, Guy Boudoukh, Peter Izsak, and Moshe Wasserblat. 2019 · 2019
Later among the works it cites.
Maha Elbayad, , Jiatao Gu, Edouard Grave, and Michael Auli. 2020 · 2020
Closest in time.
Reducing transformer depth on demand with structured dropout
Angela Fan, Edouard Grave, and Armand Joulin. 2020 · 2020
Closest in time.
FastBERT: a self-distilling BERT with adaptive inference time
Weijie Liu, Peng Zhou, Zhe Zhao, Zhiruo Wang, Haotang Deng, and Qi Ju. 2020 · 2020
Closest in time.