Fetching the paper…
Reading the bibliography…
Transformer models have achieved promising results on natural language processing (NLP) tasks including extractive question answering (QA).
Revealing the dark secrets of BERT
Kovaleva, O.; Romanov, A.; Rogers, A.; and Rumshisky, A. 2019 · 1908
Earlier work this paper cites.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Sanh, V.; Debut, L.; Chaumond, J.; and Wolf, T. 2019 · 1910
Earlier work this paper cites.
HuggingFace’s Transformers: State-of-the-art Natural Language Processing
Wolf, T.; Debut, L.; Sanh, V.; Chaumond, J.; Delangue, C.; Moi, A.; Cistac, P.; Rault, T.; Louf, R.; Funtowicz, M.; Davison, J.; Shleifer, S.; von Platen, P.; Ma, C.; Jernite, Y.; Plu, J.; Xu, C.; Scao, T. L.; Gugger, S.; Drame, M.; Lhoest, Q.; and Rush, A. M. 2019 · 1910
Earlier work this paper cites.
Learning representations by back-propagating errors
Rumelhart, D. E.; Hinton, G. E.; and Williams, R. J. 1986 · 1986
Earlier work this paper cites.
Long short-term memory
Hochreiter, S.; and Schmidhuber, J. 1997 · 1997
Earlier work this paper cites.
PoWER-BERT: Accelerating BERT inference for Classification Tasks
Goyal, S.; Choudhary, A. R.; Chakaravarthy, V.; ManishRaje, S.; Sabharwal, Y.; and Verma, A. 2020 · 2001
Earlier work this paper cites.
Permitted and forbidden sets in symmetric threshold-linear networks
Hahnloser, R. H.; and Seung, H. S. 2001 · 2001
Earlier work this paper cites.
Longformer: The long-document transformer
Beltagy, I.; Peters, M. E.; and Cohan, A. 2020 · 2004
Earlier work this paper cites.
Deformer: Decomposing pre-trained transformers for faster question answering
Cao, Q.; Trivedi, H.; Balasubramanian, A.; and Balasubramanian, N. 2020 · 2005
Earlier work this paper cites.
Funnel-Transformer: Filtering out Sequential Redundancy for Efficient Language Processing
Dai, Z.; Lai, G.; Yang, Y.; and Le, Q. V. 2020 · 2006
Earlier work this paper cites.
BERT Loses Patience: Fast and Robust Inference with Early Exit
Zhou, W.; Xu, C.; Ge, T.; McAuley, J.; Xu, K.; and Wei, F. 2020 · 2006
Earlier work this paper cites.
Big bird: Transformers for longer sequences
Zaheer, M.; Guruganesh, G.; Dubey, A.; Ainslie, J.; Alberti, C.; Ontanon, S.; Pham, P.; Ravula, A.; Wang, Q.; Yang, L.; et al. 2020 · 2007
Earlier work this paper cites.
Accelerating Sparse DNN Models without Hardware-Support via Tile-Wise Sparsity
Guo, C.; Hsueh, B. Y.; Leng, J.; Qiu, Y.; Guan, Y.; Wang, Z.; Jia, X.; Li, X.; Guo, M.; and Zhu, Y. 2020 · 2008
Earlier work this paper cites.
Efficient Transformers: A Survey
Tay, Y.; Dehghani, M.; Bahri, D.; and Metzler, D. 2020 · 2009
Earlier work this paper cites.
Length-Adaptive Transformer: Train Once with Length Drop, Use Anytime with Search
Kim, G.; and Cho, K. 2020 · 2010
Earlier work this paper cites.
How Far Does BERT Look At: Distance-based Clustering and Analysis of BERT ′
Guan, Y.; Leng, J.; Li, C.; Chen, Q.; and Guo, M. 2020 · 2011
Cited alongside, same era.
Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
Ioffe, S.; and Szegedy, C. 2015 · 2015
Cited alongside, same era.
SQuAD: 100, 000+ Questions for Machine Comprehension of Text
Rajpurkar, P.; Zhang, J.; Lopyrev, K.; and Liang, P. 2016 · 2016
Cited alongside, same era.
NEWSQA: A MACHINE COMPREHENSION DATASET
Trischler, A.; Wang, T.; Yuan, X.; Harris, J.; Sordoni, A.; Bachman, P.; and Suleman, K. 2016 · 2016
Cited alongside, same era.
Skip RNN: Learning to Skip State Updates in Recurrent Neural Networks
Campos, V.; Jou, B.; Giró-i-Nieto, X.; Torres, J.; and Chang, S. 2017 · 2017
Cited alongside, same era.
HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering
Yang, Z.; Qi, P.; Zhang, S.; Bengio, Y.; Cohen, W. W.; Salakhutdinov, R.; and Manning, C. D. 2018 · 2018
Later among the works it cites.
What Does BERT Look at? An Analysis of BERT’s Attention
Clark, K.; Khandelwal, U.; Levy, O.; and Manning, C. D. 2019 · 2019
Later among the works it cites.
Reformer: The Efficient Transformer
Kitaev, N.; Kaiser, L.; and Levskaya, A. 2019 · 2019
Later among the works it cites.
Natural questions: a benchmark for question answering research
Kwiatkowski, T.; Palomaki, J.; Redfield, O.; Collins, M.; Parikh, A.; Alberti, C.; Epstein, D.; Polosukhin, I.; Devlin, J.; Lee, K.; et al. 2019 · 2019
Later among the works it cites.
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
Lan, Z.; Chen, M.; Goodman, S.; Gimpel, K.; Sharma, P.; and Soricut, R. 2019 · 2019
Later among the works it cites.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dunn, M.; Sagun, L.; Higgins, M.; Guney, V. U.; Cirik, V.; and Cho, K. 2017 · 2017
Cited alongside, same era.
TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
Joshi, M.; Choi, E.; Weld, D. S.; and Zettlemoyer, L. 2017 · 2017
Cited alongside, same era.
A structured self-attentive sentence embedding
Lin, Z.; Feng, M.; Santos, C. N. d.; Yu, M.; Xiang, B.; Zhou, B.; and Bengio, Y. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
Learning to Skim Text
Yu, A. W.; Lee, H.; and Le, Q. 2017 · 2017
Cited alongside, same era.
Universal Transformers
Dehghani, M.; Gouws, S.; Vinyals, O.; Uszkoreit, J.; and Kaiser, L. 2018 · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2018 · 2018
Cited alongside, same era.
Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; and Stoyanov, V. 2019 · 2019
Later among the works it cites.
Are Sixteen Heads Really Better than One?
Michel, P.; Levy, O.; and Neubig, G. 2019 · 2019
Later among the works it cites.
Adversarial Defense Through Network Profiling Based Path Extraction
Qiu, Y.; Leng, J.; Guo, C.; Chen, Q.; Li, C.; Guo, M.; and Zhu, Y. 2019 · 2019
Later among the works it cites.
Lite Transformer with Long-Short Range Attention
Wu, Z.; Liu, Z.; Lin, J.; Lin, Y.; and Han, S. 2019 · 2019
Later among the works it cites.
Ptolemy: Architecture Support for Robust Deep Learning
Gan, Y.; Qiu, Y.; Leng, J.; Guo, M.; and Zhu, Y. 2020 · 2020
Later among the works it cites.
Recent Trends in Deep Learning Based Open-Domain Textual Question Answering Systems
Huang, Z.; Xu, S.; Hu, M.; Wang, X.; Qiu, J.; Fu, Y.; Zhao, Y.; Peng, Y.; and Wang, C. 2020 · 2020
Later among the works it cites.
Torchprofile
Liu, Z. 2020 · 2020
Later among the works it cites.
Q-BERT: A BERT-based Framework for Computing SPARQL Similarity in Natural Language
Wang, C.; and Zhang, X. 2020 · 2020
Later among the works it cites.
SQuant: On-the-Fly Data-Free Quantization via Diagonal Hessian Approximation
Guo, C.; Qiu, Y.; Leng, J.; Gao, X.; Zhang, C.; Liu, Y.; Yang, F.; Zhu, Y.; and Guo, M. 2022 · 2022
Closest in time.