Fetching the paper…
Reading the bibliography…
The design choices in Transformer feed-forward neural networks have resulted in significant computational and parameter overhead.
Understanding and improving transformer from a multi-particle dynamic system point of view
Yiping Lu, Zhuohan Li, Di He, Zhiqing Sun, Bin Dong, Tao Qin, Liwei Wang, and Tie-Yan Liu. 2019 · 1906
Earlier work this paper cites.
Define: Deep factorized input word embeddings for neural sequence modeling
Sachin Mehta, Rik Koncel-Kedziorski, Mohammad Rastegari, and Hannaneh Hajishirzi. 2019 · 1911
Earlier work this paper cites.
Swarm Intelligence: From Natural to Artificial Systems
Eric Bonabeau, Marco Dorigo, and Guy Theraulaz. 1999 · 1999
Earlier work this paper cites.
Low-rank bottleneck in multi-head attention models
Srinadh Bhojanapalli, Chulhee Yun, Ankit Singh Rawat, Sashank J. Reddi, and Sanjiv Kumar. 2020 · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Glu variants improve transformer
Noam Shazeer. 2020 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Consensus decision making in animals
Larissa Conradt and Timothy J Roper. 2005 · 2005
Earlier work this paper cites.
Multi-branch attentive transformer
Yang Fan, Shufang Xie, Yingce Xia, Lijun Wu, Tao Qin, Xiang-Yang Li, and Tie-Yan Liu. 2020 · 2006
Earlier work this paper cites.
Collective cognition in animal groups
Iain D Couzin. 2009 · 2009
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V. Le. 2014 · 2014
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Weighted transformer network for machine translation
Karim Ahmed, Nitish Shirish Keskar, and Richard Socher. 2017 · 2017
Earlier work this paper cites.
Language modeling with gated convolutional networks
Yann N. Dauphin, Angela Fan, Michael Auli, and David Grangier. 2017 · 2017
Earlier work this paper cites.
Convolutional sequence to sequence learning
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N. Dauphin. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Training deeper neural machine translation models with transparent attention
Ankur Bapna, Mia Chen, Orhan Firat, Yuan Cao, and Yonghui Wu. 2018 · 2018
Earlier work this paper cites.
Multi-head attention with disagreement regularization
Jian Li, Zhaopeng Tu, Baosong Yang, Michael R. Lyu, and Tong Zhang. 2018 · 2018
Earlier work this paper cites.
Shufflenet v2: Practical guidelines for efficient cnn architecture design
Ningning Ma, Xiangyu Zhang, Hai-Tao Zheng, and Jian Sun. 2018 · 2018
Cited alongside, same era.
A call for clarity in reporting BLEU scores
Matt Post. 2018 · 2018
Cited alongside, same era.
Self-attention with relative position representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. 2018 · 2018
Cited alongside, same era.
Adaptive input representations for neural language modeling
Alexei Baevski and Michael Auli. 2019 · 2019
Cited alongside, same era.
What does BERT look at? an analysis of BERT’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019 · 2019
Cited alongside, same era.
Universal transformers
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Lukasz Kaiser. 2019 · 2019
Cited alongside, same era.
Attention is not all you need: pure attention loses rank doubly exponentially with depth
Yihe Dong, Jean-Baptiste Cordonnier, and Andreas Loukas. 2021 · 2021
Later among the works it cites.
Mask attention networks: Rethinking and strengthen transformer
Zhihao Fan, Yeyun Gong, Dayiheng Liu, Zhongyu Wei, Siyuan Wang, Jian Jiao, Nan Duan, Ruofei Zhang, and Xuanjing Huang. 2021 · 2021
Later among the works it cites.
Transformer feed-forward layers are key-value memories
Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. 2021 · 2021
Later among the works it cites.
RealFormer: Transformer likes residual attention
Ruining He, Anirudh Ravula, Bhargav Kanagal, and Joshua Ainslie. 2021 · 2021
Later among the works it cites.
Delight: Deep and light-weight transformer
Sachin Mehta, Marjan Ghazvininejad, Srinivasan Iyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2021 · 2021
Later among the works it cites.
Subformer: Exploring weight sharing for parameter efficiency in generative transformers
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Are sixteen heads really better than one?
Paul Michel, Omer Levy, and Graham Neubig. 2019 · 2019
Cited alongside, same era.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
Cited alongside, same era.
The evolved transformer
David R. So, Quoc V. Le, and Chen Liang. 2019 · 2019
Cited alongside, same era.
Efficientnet: Rethinking model scaling for convolutional neural networks
Mingxing Tan and Quoc Le. 2019 · 2019
Cited alongside, same era.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019 · 2019
Cited alongside, same era.
Learning deep transformer models for machine translation
Qiang Wang, Bei Li, Tong Xiao, Jingbo Zhu, Changliang Li, Derek F. Wong, and Lidia S. Chao. 2019 · 2019
Cited alongside, same era.
Machel Reid, Edison Marrese-Taylor, and Yutaka Matsuo. 2021 · 2021
Later among the works it cites.
Facebook AI’s WMT21 news translation task submission
Chau Tran, Shruti Bhosale, James Cross, Philipp Koehn, Sergey Edunov, and Angela Fan. 2021 · 2021
Later among the works it cites.
R-drop: Regularized dropout for neural networks
Lijun Wu, Juntao Li, Yue Wang, Qi Meng, Tao Qin, Wei Chen, Min Zhang, Tie-Yan Liu, et al. 2021 · 2021
Later among the works it cites.
EdgeFormer: A parameter-efficient transformer for on-device seq2seq generation
Tao Ge, Si-Qing Chen, and Furu Wei. 2022 · 2022
Later among the works it cites.
Multi-path transformer is better: A case study on neural machine translation
Ye Lin, Shuhan Zhou, Yanyang Li, Anxiang Ma, Tong Xiao, and Jingbo Zhu. 2022 · 2022
Later among the works it cites.
Mega: Moving average equipped gated attention
Xuezhe Ma, Chunting Zhou, Xiang Kong, Junxian He, Liangke Gui, Graham Neubig, Jonathan May, and Luke Zettlemoyer. 2022 · 2022
Later among the works it cites.
Improving transformer with an admixture of attention heads
Tan Minh Nguyen, Tam Minh Nguyen, Hai Ngoc Do, Khai Nguyen, Vishwanath Saragadam, Minh Pham, Nguyen Duy Khuong, Nhat Ho, and Stanley Osher. 2022 · 2022
Later among the works it cites.
COMET-22: Unbabel-IST 2022 submission for the metrics shared task
Ricardo Rei, José G. C. de Souza, Duarte Alves, Chrysoula Zerva, Ana C Farinha, Taisiya Glushkova, Alon Lavie, Luisa Coheur, and André F. T. Martins. 2022 · 2022
Later among the works it cites.
Anti-oversmoothing in deep vision transformers via the fourier domain analysis: From theory to practice
Peihao Wang, Wenqing Zheng, Tianlong Chen, and Zhangyang Wang. 2022 · 2022
Later among the works it cites.
MoEfication: Transformer feed-forward layers are mixtures of experts
Zhengyan Zhang, Yankai Lin, Zhiyuan Liu, Peng Li, Maosong Sun, and Jie Zhou. 2022 · 2022
Later among the works it cites.
TranSFormer: Slow-fast transformer for machine translation
Bei Li, Yi Jing, Xu Tan, Zhen Xing, Tong Xiao, and Jingbo Zhu. 2023 · 2023
Closest in time.
Eit: Enhanced interactive transformer
Tong Zheng, Bei Li, Huiwen Bao, Tong Xiao, and Jingbo Zhu. 2024 · 2024
Closest in time.