Fetching the paper…
Reading the bibliography…
The self-attention mechanism in transformers and the message-passing mechanism in graph neural networks are repeatedly applied within deep learning architectures.
Citeseer: An automatic citation indexing system
C Lee Giles, Kurt D Bollacker, and Steve Lawrence · 1998
Earlier work this paper cites.
Automating the construction of internet portals with machine learning
Andrew Kachites McCallum, Kamal Nigam, Jason Rennie, and Kristie Seymore · 2000
Earlier work this paper cites.
The pascal visual object classes (voc) challenge
Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman · 2010
Earlier work this paper cites.
Semantic contours from inverse detectors
Bharath Hariharan, Pablo Arbeláez, Lubomir Bourdev, Subhransu Maji, and Jitendra Malik · 2011
Earlier work this paper cites.
Steps toward deep kernel methods from infinite neural networks, 2015
Tamir Hazan and Tommi Jaakkola · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Semi-supervised classification with graph convolutional networks
Thomas N. Kipf and Max Welling · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Earlier work this paper cites.
Deep neural networks as gaussian processes
Jaehoon Lee, Yasaman Bahri, Roman Novak, Samuel S Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Earlier work this paper cites.
On regularized losses for weakly-supervised cnn segmentation
Meng Tang, Federico Perazzi, Abdelaziz Djelouah, Ismail Ben Ayed, Christopher Schroers, and Yuri Boykov · 2018
Earlier work this paper cites.
Graph attention networks
Petar Veličković, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Liò, and Yoshua Bengio · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Attention guided graph convolutional networks for relation extraction
Zhijiang Guo, Yan Zhang, and Wei Lu · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach, 2019
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Earlier work this paper cites.
Transformers without tears: Improving the normalization of self-attention
Toan Q Nguyen and Julian Salazar · 2019
Earlier work this paper cites.
Self-supervised difference detection for weakly-supervised semantic segmentation
Wataru Shimoda and Keiji Yanai · 2019
Earlier work this paper cites.
Pixel-adaptive convolutional neural networks
Hang Su, Varun Jampani, Deqing Sun, Orazio Gallo, Erik Learned-Miller, and Jan Kautz · 2019
Earlier work this paper cites.
Single-stage semantic segmentation from image labels
Nikita Araslanov and Stefan Roth · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, et al · 2020
Earlier work this paper cites.
A note on over-smoothing for graph neural networks
Chen Cai and Yusu Wang · 2020
Earlier work this paper cites.
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut · 2020
Cited alongside, same era.
Understanding the difficulty of training transformers
Liyuan Liu, Xiaodong Liu, Jianfeng Gao, Weizhu Chen, and Jiawei Han · 2020
Cited alongside, same era.
Improving transformer models by reordering their sublayers
Ofir Press, Noah A. Smith, and Omer Levy · 2020
Cited alongside, same era.
On layer normalization in the transformer architecture
Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tieyan Liu · 2020
Cited alongside, same era.
On layer normalization in the transformer architecture
Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tieyan Liu · 2020
Cited alongside, same era.
How neural networks extrapolate: From feedforward to graph neural networks
Weakly supervised semantic segmentation for large-scale point cloud
Yachao Zhang, Zonghao Li, Yuan Xie, Yanyun Qu, Cuihua Li, and Tao Mei · 2021
Later among the works it cites.
{IOT}: Instance-wise layer reordering for transformer structures
Jinhua Zhu, Lijun Wu, Yingce Xia, Shufang Xie, Tao Qin, Wengang Zhou, Houqiang Li, and Tie-Yan Liu · 2021
Later among the works it cites.
PaLM: Scaling language modeling with pathways, 2022
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, et al · 2022
Later among the works it cites.
Weakly supervised semantic segmentation by pixel-to-prototype contrast
Ye Du, Zehua Fu, Qingjie Liu, and Yunhong Wang · 2022
Later among the works it cites.
L2g: A simple local-to-global knowledge transfer framework for weakly supervised semantic segmentation
Peng-Tao Jiang, Yuqi Yang, Qibin Hou, and Yunchao Wei · 2022
Later among the works it cites.
Not too little, not too much: a theoretical analysis of graph (over)smoothing
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Keyulu Xu, Mozhi Zhang, Jingling Li, Simon S Du, Ken-ichi Kawarabayashi, and Stefanie Jegelka · 2020
Cited alongside, same era.
Reliability does matter: An end-to-end weakly supervised semantic segmentation approach
Bingfeng Zhang, Jimin Xiao, Yunchao Wei, Mingjie Sun, and Kaizhu Huang · 2020
Cited alongside, same era.
Pairnorm: Tackling oversmoothing in gnns
Lingxiao Zhao and Leman Akoglu · 2020
Cited alongside, same era.
Grand: Graph neural diffusion
Ben Chamberlain, James Rowbottom, Maria I Gorinova, Michael Bronstein, Stefan Webb, and Emanuele Rossi · 2021
Cited alongside, same era.
Lipschitz normalization for self-attention layers with application to graph neural networks
George Dasoulas, Kevin Scaman, and Aladin Virmaux · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Cited alongside, same era.
Improve vision transformers training by suppressing over-smoothing
Chengyue Gong, Dilin Wang, Meng Li, Vikas Chandra, and Qiang Liu · 2021
Cited alongside, same era.
Nicolas Keriven · 2022
Later among the works it cites.
Weakly supervised semantic segmentation using out-of-distribution data
Jungbeom Lee, Seong Joon Oh, Sangdoo Yun, Junsuk Choe, Eunji Kim, and Sungroh Yoon · 2022
Later among the works it cites.
Expansion and shrinkage of localization for weakly-supervised semantic segmentation
JINLONG LI, ZEQUN JIE, Xu Wang, xiaolin wei, and Lin Ma · 2022
Later among the works it cites.
Learning self-supervised low-rank network for single-stage weakly and semi-supervised semantic segmentation
Junwen Pan, Pengfei Zhu, Kaihua Zhang, Bing Cao, Yu Wang, Dingwen Zhang, Junwei Han, and Qinghua Hu · 2022
Later among the works it cites.
Weakly-supervised semantic segmentation with visual words learning and hybrid pooling
Lixiang Ru, Bo Du, Yibing Zhan, and Chen Wu · 2022
Later among the works it cites.
Learning affinity from attention: end-to-end weakly-supervised semantic segmentation with transformers
Lixiang Ru, Yibing Zhan, Baosheng Yu, and Bo Du · 2022
Later among the works it cites.
Revisiting over-smoothing in BERT from the perspective of graph
Han Shi, JIAHUI GAO, Hang Xu, Xiaodan Liang, Zhenguo Li, Lingpeng Kong, Stephen M. S. Lee, and James Kwok · 2022
Later among the works it cites.
Anti-oversmoothing in deep vision transformers via the fourier domain analysis: From theory to practice
Peihao Wang, Wenqing Zheng, Tianlong Chen, and Zhangyang Wang · 2022
Later among the works it cites.
Scaled relu matters for training vision transformers
Pichao Wang, Xue Wang, Hao Luo, Jingkai Zhou, Zhipeng Zhou, Fan Wang, Hao Li, and Rong Jin · 2022
Later among the works it cites.
Multi-class token transformer for weakly supervised semantic segmentation
Lian Xu, Wanli Ouyang, Mohammed Bennamoun, Farid Boussaid, and Dan Xu · 2022
Later among the works it cites.
Regional semantic contrast and aggregation for weakly supervised semantic segmentation
Tianfei Zhou, Meijie Zhang, Fang Zhao, and Jianwu Li · 2022
Later among the works it cites.
Contranorm: A contrastive learning perspective on oversmoothing and beyond
Xiaojun Guo, Yifei Wang, Tianqi Du, and Yisen Wang · 2023
Closest in time.
Token contrast for weakly-supervised semantic segmentation
Lixiang Ru, Heliang Zheng, Yibing Zhan, and Bo Du · 2023
Closest in time.
A survey on oversmoothing in graph neural networks
T Konstantin Rusch, Michael M Bronstein, and Siddhartha Mishra · 2023
Closest in time.
A non-asymptotic analysis of oversmoothing in graph neural networks
Xinyi Wu, Zhengdao Chen, William Wei Wang, and Ali Jadbabaie · 2023
Closest in time.
Residual: Transformer with dual residual connections, 2023
Shufang Xie, Huishuai Zhang, Junliang Guo, Xu Tan, Jiang Bian, Hany Hassan Awadalla, Arul Menezes, Tao Qin, and Rui Yan · 2023
Closest in time.