Fetching the paper…
Reading the bibliography…
Transformer-based models are popularly used in natural language processing (NLP).
Automatically Constructing a Corpus of Sentential Paraphrases
Dolan, W. and Brockett, C · 2005
Earlier work this paper cites.
The Fifth PASCAL Recognizing Textual Entailment Challenge
Bentivogli, L., Dagan, I., Hoa, D., Giampiccolo, D., and Magnini, B · 2009
Earlier work this paper cites.
Recursive Deep Models for Semantic Compositionality Over a Sentiment Treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C., Ng, A., and Potts, C · 2013
Earlier work this paper cites.
Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books
Zhu, Y., Kiros, R., Zemel, R., Salakhutdinov, R., Urtasun, R., Torralba, A., and Fidler, S · 2015
Earlier work this paper cites.
SQuAD: 100,000+ Questions for Machine Comprehension of Text
Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P · 2016
Earlier work this paper cites.
Google’s Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
Wu, Y., Schuster, M., Chen, Z., Le, Q., Norouzi, M., Macherey, W., Krikun, M., Cao, Y., Gao, Q., Macherey, K., et al · 2016
Earlier work this paper cites.
SemEval-2017 Task 1: Semantic Textual Similarity Multilingual and Crosslingual Focused Evaluation
Cer, D., Diab, M., Agirre, E., Lopez-Gazpio, I., and Specia, L · 2017
Earlier work this paper cites.
The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables
Maddison, C., Mnih, A., and Teh, Y · 2017
Earlier work this paper cites.
Attention Is All You Need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Quora question pairs
Chen, Z., Zhang, H., Zhang, X., and Zhao, L · 2018
Earlier work this paper cites.
Scaling Neural Machine Translation
Ott, M., Edunov, S., Grangier, D., and Auli, M · 2018
Earlier work this paper cites.
Know What You Don’t Know: Unanswerable Questions for SQuAD
Rajpurkar, P., Jia, R., and Liang, P · 2018
Earlier work this paper cites.
A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference
Williams, A., Nangia, N., and Bowman, S · 2018
Cited alongside, same era.
SNAS: Stochastic Neural Architecture Search
Xie, S., Zheng, H., Liu, C., and Lin, L · 2018
Cited alongside, same era.
SWAG: A Large-Scale Adversarial Dataset for Grounded Commonsense Inference
Zellers, R., Bisk, Y., Schwartz, R., and Choi, Y · 2018
Cited alongside, same era.
Generating Long Sequences with Sparse Transformers
Child, R., Gray, S., Radford, A., and Sutskever, I · 2019
Cited alongside, same era.
What Does BERT Look at? An Analysis of BERT’s Attention
Clark, K., Khandelwal, U., Levy, O., and Manning, C · 2019
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Are Sixteen Heads Really Better than One?
Michel, P., Levy, O., and Neubig, G · 2019
Later among the works it cites.
SANVis: Visual Analytics for Understanding Self-Attention Networks
Park, C., Na, I., Jo, Y., Shin, S., Yoo, J., Kwon, B., Zhao, J., Noh, H., Lee, Y., and Choo, J · 2019
Later among the works it cites.
Neural Network Acceptability Judgments
Warstadt, A., Singh, A., and Bowman, S · 2019
Later among the works it cites.
Are Transformers universal approximators of sequence-to-sequence functions?
Yun, C., Bhojanapalli, S., Rawat, A., Reddi, S., and Kumar, S · 2019
Later among the works it cites.
Longformer: The Long-Document Transformer
Beltagy, I., Peters, M., and Cohan, A · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Devlin, J., Chang, M., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Neural Architecture Search: A Survey
Elsken, T., Metzen, J., and Hutter, F · 2019
Cited alongside, same era.
Efficient Training of BERT by Progressively Stacking
Gong, L., He, D., Li, Z., Qin, T., Wang, L., and Liu, T · 2019
Cited alongside, same era.
Star-Transformer
Guo, Q., Qiu, X., Liu, P., Shao, Y., Xue, X., and Zhang, Z · 2019
Cited alongside, same era.
Revealing the Dark Secrets of BERT
Kovaleva, O., Romanov, A., Rogers, A., and Rumshisky, A · 2019
Cited alongside, same era.
Enhancing the Locality and Breaking the Memory Bottleneck of Transformer on Time Series Forecasting
Li, S., Jin, X., Xuan, Y., Zhou, X., Chen, W., Wang, Y., and Yan, X · 2019
Cited alongside, same era.
DARTS: Differentiable Architecture Search
Liu, H., Simonyan, K., and Yang, Y · 2019
Cited alongside, same era.
Bi, K., Xie, L., Chen, X., Wei, L., and Tian, Q · 2020
Later among the works it cites.
Language Models are Few-Shot Learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Later among the works it cites.
End-to-End Object Detection with Transformers
Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., and Zagoruyko, S · 2020
Later among the works it cites.
Blockwise Self-Attention for Long Document Understanding
Qiu, J., Ma, H., Levy, O., Yih, W., Wang, S., and Tang, J · 2020
Later among the works it cites.
O ( n ) {O}(n) Connections are Expressive Enough: Universal Approximability of Sparse Transformers
Yun, C., Chang, Y., Bhojanapalli, S., Rawat, A., Reddi, S., and Kumar, S · 2020
Later among the works it cites.
Big Bird: Transformers for Longer Sequences
Zaheer, M., Guruganesh, G., Dubey, K., Ainslie, J., Alberti, C., Ontanon, S., Pham, P., Ravula, A., Wang, Q., Yang, L., et al · 2020
Later among the works it cites.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2021
Closest in time.