Fetching the paper…
Reading the bibliography…
Transformers, renowned for their self-attention mechanism, have achieved state-of-the-art performance across various tasks in natural language processing, computer vision, time-series modeling, etc.
Building a large annotated corpus of english: The penn treebank
Mitchell Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz · 1993
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
Bill Dolan and Chris Brockett · 2005
Earlier work this paper cites.
Linformer: Self-attention with linear complexity
Sinong Wang, Belinda Z Li, Madian Khabsa, Han Fang, and Hao Ma · 2006
Earlier work this paper cites.
Functions of matrices: theory and computation
Nicholas J Higham · 2008
Earlier work this paper cites.
The fifth pascal recognizing textual entailment challenge
Luisa Bentivogli, Peter Clark, Ido Dagan, and Danilo Giampiccolo · 2009
Earlier work this paper cites.
Zinc: a free tool to discover chemistry for biology
John J Irwin, Teague Sterling, Michael M Mysinger, Erin S Bolstad, and Ryan G Coleman · 2012
Earlier work this paper cites.
Discrete signal processing on graphs
Aliaksei Sandryhaila and José MF Moura · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts · 2013
Earlier work this paper cites.
Towards end-to-end speech recognition with recurrent neural networks
Alex Graves and Navdeep Jaitly · 2014
Earlier work this paper cites.
Discrete signal processing on graphs: Frequency analysis
Aliaksei Sandryhaila and Jose MF Moura · 2014
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
Convolutional neural networks on graphs with fast localized spectral filtering
Michaël Defferrard, Xavier Bresson, and Pierre Vandergheynst · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Earlier work this paper cites.
Towards end-to-end speech recognition with deep convolutional neural networks
Ying Zhang, Mohammad Pezeshki, Philémon Brakel, Saizheng Zhang, César Laurent, Yoshua Bengio, and Aaron Courville · 2016
Earlier work this paper cites.
Semeval-2017 task 1: Semantic textual similarity-multilingual and cross-lingual focused evaluation
Daniel Cer, Mona Diab, Eneko Agirre, Inigo Lopez-Gazpio, and Lucia Specia · 2017
Earlier work this paper cites.
Densely connected convolutional networks
Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger · 2017
Earlier work this paper cites.
Semi-supervised classification with graph convolutional networks
Thomas N. Kipf and Max Welling · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel R Bowman · 2017
Earlier work this paper cites.
Quora question pairs, 2018
Zihan Chen, Hongbo Zhang, Xiaoji Zhang, and Leqi Zhao · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
Graph Attention Networks
Petar Veličković, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Liò, and Yoshua Bengio · 2018
Earlier work this paper cites.
Wei Wang, Ming Yan, and Chen Wu · 2018
Earlier work this paper cites.
Representation learning on graphs with jumping knowledge networks
Keyulu Xu, Chengtao Li, Yonglong Tian, Tomohiro Sonobe, Ken-ichi Kawarabayashi, and Stefanie Jegelka · 2018
Earlier work this paper cites.
GRU-ODE-Bayes: Continuous modeling of sporadically-observed time series
Edward De Brouwer, Jaak Simm, Adam Arany, and Yves Moreau · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Diffusion improves graph learning
Johannes Gasteiger, Stefan Weißenberger, and Stephan Günnemann · 2019
Earlier work this paper cites.
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut · 2019
Earlier work this paper cites.
RoBERTa: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Earlier work this paper cites.
Specaugment: A simple data augmentation method for automatic speech recognition
Daniel S Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D Cubuk, and Quoc V Le · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Neural network acceptability judgments
Alex Warstadt, Amanpreet Singh, and Samuel R Bowman · 2019
Cited alongside, same era.
Pytorch image models
Ross Wightman · 2019
Cited alongside, same era.
Devign: Effective vulnerability identification by learning comprehensive program semantics via graph neural networks
Yaqin Zhou, Shangqing Liu, Jingkai Siow, Xiaoning Du, and Yang Liu · 2019
Cited alongside, same era.
A note on over-smoothing for graph neural networks
Chen Cai and Yusu Wang · 2020
Cited alongside, same era.
CodeBERT: A pre-trained model for programming and natural languages
Zhangyin Feng, Daya Guo, Duyu Tang, Nan Duan, Xiaocheng Feng, Ming Gong, Linjun Shou, Bing Qin, Ting Liu, Daxin Jiang, et al · 2020
Cited alongside, same era.
Conformer: Convolution-augmented transformer for speech recognition
Improving vision transformers by revisiting high-frequency components
Jiawang Bai, Li Yuan, Shu-Tao Xia, Shuicheng Yan, Zhifeng Li, and Wei Liu · 2022
Later among the works it cites.
The principle of diversity: Training stronger vision transformers calls for reducing all levels of redundancy
Tianlong Chen, Zhenyu Zhang, Yu Cheng, Ahmed Awadallah, and Zhangyang Wang · 2022
Later among the works it cites.
Long range graph benchmark
Vijay Prakash Dwivedi, Ladislav Rampášek, Mikhail Galkin, Ali Parviz, Guy Wolf, Anh Tuan Luu, and Dominique Beaini · 2022
Later among the works it cites.
Not too little, not too much: a theoretical analysis of graph (over) smoothing
Nicolas Keriven · 2022
Later among the works it cites.
Signal propagation in transformers: Theoretical perspectives and the role of rank collapse
Lorenzo Noci, Sotiris Anagnostidis, Luca Biggio, Antonio Orvieto, Sidak Pal Singh, and Aurelien Lucchi · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, et al · 2020
Cited alongside, same era.
Open graph benchmark: Datasets for machine learning on graphs
Weihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong, Hongyu Ren, Bowen Liu, Michele Catasta, and Jure Leskovec · 2020
Cited alongside, same era.
Transformers are rnns: Fast autoregressive transformers with linear attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret · 2020
Cited alongside, same era.
BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer · 2020
Cited alongside, same era.
Signal processing on directed graphs: The role of edge directionality when processing and learning from network data
Antonio G. Marques, Santiago Segarra, and Gonzalo Mateos · 2020
Cited alongside, same era.
Graph neural networks exponentially lose expressive power for node classification
Kenta Oono and Taiji Suzuki · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Cited alongside, same era.
Branchformer: Parallel mlp-attention architectures to capture local and global context for speech recognition and understanding
Yifan Peng, Siddharth Dalmia, Ian Lane, and Shinji Watanabe · 2022
Later among the works it cites.
Recipe for a general, powerful, scalable graph transformer
Ladislav Rampášek, Michael Galkin, Vijay Prakash Dwivedi, Anh Tuan Luu, Guy Wolf, and Dominique Beaini · 2022
Later among the works it cites.
Revisiting over-smoothing in bert from the perspective of graph
Han Shi, Jiahui Gao, Hang Xu, Xiaodan Liang, Zhenguo Li, Lingpeng Kong, Stephen Lee, and James T Kwok · 2022
Later among the works it cites.
Anti-oversmoothing in deep vision transformers via the fourier domain analysis: From theory to practice
Peihao Wang, Wenqing Zheng, Tianlong Chen, and Zhangyang Wang · 2022
Later among the works it cites.
Addressing token uniformity in transformers via singular value transformation
Hanqi Yan, Lin Gui, Wenjie Li, and Yulan He · 2022
Later among the works it cites.
Zhewei Yao, Xiaoxia Wu, Conglong Li, Connor Holmes, Minjia Zhang, Cheng Li, and Yuxiong He · 2022
Later among the works it cites.
Jump self-attention: Capturing high-order statistics in transformers
Haoyi Zhou, Siyang Xiao, Shanghang Zhang, Jieqi Peng, Shuai Zhang, and Jianxin Li · 2022
Later among the works it cites.
A simple yet effective svd-gcn for directed graphs
Chunya Zou, Andi Han, Lequan Lin, and Junbin Gao · 2022
Later among the works it cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Closest in time.
Centered self-attention layers
Ameen Ali, Tomer Galanti, and Lior Wolf · 2023
Closest in time.
Alleviating over-smoothing for unsupervised sentence representation
Nuo Chen, Linjun Shou, Ming Gong, Jian Pei, Bowen Cao, Jianhui Chang, Daxin Jiang, and Jia Li · 2023
Closest in time.
Gread: Graph neural reaction-diffusion networks
Jeongwhan Choi, Seoyoung Hong, Noseong Park, and Sung-Bae Cho · 2023
Closest in time.
Benchmarking graph neural networks
Vijay Prakash Dwivedi, Chaitanya K Joshi, Anh Tuan Luu, Thomas Laurent, Yoshua Bengio, and Xavier Bresson · 2023
Closest in time.
ContraNorm: A contrastive learning perspective on oversmoothing and beyond
Xiaojun Guo, Yifei Wang, Tianqi Du, and Yisen Wang · 2023
Closest in time.
A generalization of vit/mlp-mixer to graphs
Xiaoxin He, Bryan Hooi, Thomas Laurent, Adam Perold, Yann LeCun, and Xavier Bresson · 2023
Closest in time.
Transformers in speech processing: A survey
Siddique Latif, Aun Zaidi, Heriberto Cuayahuitl, Fahad Shamshad, Moazzam Shoukat, and Junaid Qadir · 2023
Closest in time.
A fractional graph laplacian approach to oversmoothing
Sohir Maskey, Raffaele Paolino, Aras Bacho, and Gitta Kutyniok · 2023
Closest in time.
Mitigating over-smoothing in transformers via regularized nonlocal functionals
Tam Nguyen, Tan Nguyen, and Richard Baraniuk · 2023
Closest in time.
Scattering vision transformer: Spectral mixing matters
Badri Patro and Vijay Agneeswaran · 2023
Closest in time.
Spectformer: Frequency and attention is what you need in a vision transformer
Badri N Patro, Vinay P Namboodiri, and Vijay Srinivas Agneeswaran · 2023
Closest in time.
A survey on oversmoothing in graph neural networks
T. Konstantin Rusch, Michael M. Bronstein, and Siddhartha Mishra · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Closest in time.
Stabilizing transformer training by preventing attention entropy collapse
Shuangfei Zhai, Tatiana Likhomanenko, Etai Littwin, Dan Busbridge, Jason Ramapuram, Yizhe Zhang, Jiatao Gu, and Joshua M. Susskind · 2023
Closest in time.
Setting the record straight on transformer oversmoothing
Gbètondji JS Dovonon, Michael M Bronstein, and Matt J Kusner · 2024
Closest in time.
Polynomial-based self-attention for table representation learning
Jayoung Kim, Yehjin Shin, Jeongwhan Choi, Hyowon Wi, and Noseong Park · 2024
Closest in time.
Attending to graph transformers
Luis Müller, Mikhail Galkin, Christopher Morris, and Ladislav Rampášek · 2024
Closest in time.
An attentive inductive bias for sequential recommendation beyond the self-attention
Yehjin Shin, Jeongwhan Choi, Hyowon Wi, and Noseong Park · 2024
Closest in time.
Learning flexible body collision dynamics with hierarchical contact mesh transformer
Youn-Yeol Yu, Jeongwhan Choi, Woojin Cho, Kookjin Lee, Nayong Kim, Kiseok Chang, ChangSeung Woo, ILHO KIM, SeokWoo Lee, Joon Young Yang, SOOYOUNG YOON, and Noseong Park · 2024
Closest in time.