Fetching the paper…
Reading the bibliography…
Recently over-smoothing phenomenon of Transformer-based models is observed in both vision and language fields.
Remarks on Some Nonparametric Estimates of a Density Function
M. Rosenblatt · 1956
Earlier work this paper cites.
Spectral Graph Theory
F. Chung and F. Graham · 1997
Earlier work this paper cites.
The Theory of Matrices, Volume 2
F. Gantmakher · 2000
Earlier work this paper cites.
Automatically Constructing a Corpus of Sentential Paraphrases
W. Dolan and C. Brockett · 2005
Earlier work this paper cites.
A Tutorial on Spectral Clustering
U. Von Luxburg · 2007
Earlier work this paper cites.
The Fifth PASCAL Recognizing Textual Entailment Challenge
L. Bentivogli, I. Dagan, D. Hoa, D. Giampiccolo, and B. Magnini · 2009
Earlier work this paper cites.
Recursive Deep Models for Semantic Compositionality Over a Sentiment Treebank
R. Socher, A. Perelygin, J. Wu, J. Chuang, C. Manning, A. Ng, and C. Potts · 2013
Earlier work this paper cites.
Learning Phrase Representations using RNN Encoder–Decoder for Statistical Machine Translation
K. Cho, B. van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio · 2014
Earlier work this paper cites.
Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books
Y. Zhu, R. Kiros, R. Zemel, R. Salakhutdinov, R. Urtasun, A. Torralba, and S. Fidler · 2015
Earlier work this paper cites.
J. Ba, J. Kiros, and G. Hinton · 2016
Earlier work this paper cites.
Neural Text Generation from Structured Data with Application to the Biography Domain
R. Lebret, D. Grangier, and M. Auli · 2016
Earlier work this paper cites.
SQuAD: 100,000+ Questions for Machine Comprehension of Text
P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang · 2016
Earlier work this paper cites.
Google’s Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
Y. Wu, M. Schuster, Z. Chen, Q. Le, M. Norouzi, W. Macherey, M. Krikun, Y. Cao, Q. Gao, K. Macherey, et al · 2016
Earlier work this paper cites.
SemEval-2017 Task 1: Semantic Textual Similarity Multilingual and Crosslingual Focused Evaluation
D. Cer, M. Diab, E. Agirre, I. Lopez-Gazpio, and L. Specia · 2017
Earlier work this paper cites.
Semi-Supervised Classification with Graph Convolutional Networks
T. Kipf and M. Welling · 2017
Earlier work this paper cites.
Attention Is All You Need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Quora Question Pairs
Z. Chen, H. Zhang, X. Zhang, and L. Zhao · 2018
Cited alongside, same era.
Deeper Insights into Graph Convolutional Networks for Semi-Supervised Learning
Q. Li, Z. Han, and X. Wu · 2018
Cited alongside, same era.
Scaling Neural Machine Translation
M. Ott, S. Edunov, D. Grangier, and M. Auli · 2018
Cited alongside, same era.
Know What You Don’t Know: Unanswerable Questions for SQuAD
P. Rajpurkar, R. Jia, and P. Liang · 2018
Cited alongside, same era.
A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference
A. Williams, N. Nangia, and S. Bowman · 2018
Cited alongside, same era.
Representation Learning on Graphs with Jumping Knowledge Networks
K. Xu, C. Li, Y. Tian, T. Sonobe, K. Kawarab, and S. Jegelka · 2018
Cited alongside, same era.
End-to-End Object Detection with Transformers
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko · 2020
Later among the works it cites.
Tackling Over-Smoothing for General Graph Convolutional Networks
W. Huang, Y. Rong, T. Xu, F. Sun, and J. Huang · 2020
Later among the works it cites.
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
Z. Lan, M. Chen, S. Goodman, K. Gimpel, P. Sharma, and R. Soricut · 2020
Later among the works it cites.
Graph Neural Networks Exponentially Lose Expressive Power for Node Classification
K. Oono and T. Suzuki · 2020
Later among the works it cites.
DropEdge: Towards Deep Graph Convolutional Networks on Node Classification
Y. Rong, W. Huang, T. Xu, and J. Huang · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
SWAG: A Large-Scale Adversarial Dataset for Grounded Commonsense Inference
R. Zellers, Y. Bisk, R. Schwartz, and Y. Choi · 2018
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
J. Devlin, M. Chang, K. Lee, and K. Toutanova · 2019
Cited alongside, same era.
Efficient Training of BERT by Progressively Stacking
L. Gong, D. He, Z. Li, T. Qin, L. Wang, and T. Liu · 2019
Cited alongside, same era.
Shallow-Deep Networks: Understanding and Mitigating Network Overthinking
Y. Kaya, S. Hong, and T. Dumitras · 2019
Cited alongside, same era.
Revealing the Dark Secrets of BERT
O. Kovaleva, A. Romanov, A. Rogers, and A. Rumshisky · 2019
Cited alongside, same era.
DeepGCNs: Can GCNs Go as Deep as CNNs?
G. Li, M. Müller, A. Thabet, and B. Ghanem · 2019
Cited alongside, same era.
T. Wolf, J. Chaumond, L. Debut, V. Sanh, C. Delangue, A. Moi, P. Cistac, M. Funtowicz, J. Davison, S. Shleifer, et al · 2020
Later among the works it cites.
O ( n ) {O}(n) Connections are Expressive Enough: Universal Approximability of Sparse Transformers
C. Yun, Y. Chang, S. Bhojanapalli, A. Rawat, S. Reddi, and S. Kumar · 2020
Later among the works it cites.
PairNorm: Tackling Oversmoothing in GNNs
L. Zhao and L. Akoglu · 2020
Later among the works it cites.
BERT Loses Patience: Fast and Robust Inference with Early Exit
W. Zhou, C. Xu, T. Ge, J. McAuley, K. Xu, and F. Wei · 2020
Later among the works it cites.
Attention is Not All You Need: Pure Attention Loses Rank Doubly Exponentially with Depth
Y. Dong, J. Cordonnier, and A. Loukas · 2021
Later among the works it cites.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2021
Later among the works it cites.
Improve Vision Transformers Training by Suppressing Over-smoothing
C. Gong, D. Wang, M. Li, V. Chandra, and Q. Liu · 2021
Later among the works it cites.
Realformer: Transformer likes residual attention
R. He, A. Ravula, B. Kanagal, and J. Ainslie · 2021
Later among the works it cites.
Training Graph Neural Networks with 1000 Layers
G. Li, M. Müller, B. Ghanem, and V. Koltun · 2021
Later among the works it cites.
SparseBERT: Rethinking the Importance Analysis in Self-attention
H. Shi, J. Gao, X. Ren, H. Xu, X. Liang, Z. Li, and J. Kwok · 2021
Later among the works it cites.
Segmenter: Transformer for Semantic Segmentation
R. Strudel, R. Garcia, I. Laptev, and C. Schmid · 2021
Later among the works it cites.