Fetching the paper…
Reading the bibliography…
Transformers have achieved great success in many artificial intelligence fields, such as natural language processing, computer vision, and audio processing.
Generating Long Sequences with Sparse Transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. 2019 · 1904
Earlier work this paper cites.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019a · 1907
Earlier work this paper cites.
Augmenting Self-attention with Persistent Memory
Sainbayar Sukhbaatar, Edouard Grave, Guillaume Lample, Herve Jegou, and Armand Joulin. 2019b · 1907
Earlier work this paper cites.
R-Transformer: Recurrent Neural Network Enhanced Transformer
Zhiwei Wang, Yao Ma, Zitao Liu, and Jiliang Tang. 2019b · 1907
Earlier work this paper cites.
VisualBERT: A Simple and Performant Baseline for Vision and Language
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. 2019c · 1908
Earlier work this paper cites.
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2020 · 1909
Earlier work this paper cites.
Transformers without Tears: Improving the Normalization of Self-Attention
Toan Q. Nguyen and Julian Salazar. 2019 · 1910
Earlier work this paper cites.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 1910
Earlier work this paper cites.
Fast Transformer Decoding: One Write-Head is All You Need
Noam Shazeer. 2019 · 1911
Earlier work this paper cites.
BP-Transformer: Modelling Long-Range Context via Binary Partitioning
Zihao Ye, Qipeng Guo, Quan Gan, Xipeng Qiu, and Zheng Zhang. 2019 · 1911
Earlier work this paper cites.
Axial Attention in Multidimensional Transformers
Jonathan Ho, Nal Kalchbrenner, Dirk Weissenborn, and Tim Salimans. 2019 · 1912
Earlier work this paper cites.
Controlling Computation versus Quality for Neural Sequence Models
Ankur Bapna, Naveen Arivazhagan, and Orhan Firat. 2020 · 2002
Earlier work this paper cites.
GLU Variants Improve Transformer
Noam Shazeer. 2020 · 2002
Earlier work this paper cites.
ReZero is All You Need: Fast Convergence at Large Depth
Thomas Bachlechner, Bodhisattwa Prasad Majumder, Huanru Henry Mao, Garrison W. Cottrell, and Julian J. McAuley. 2020 · 2003
Earlier work this paper cites.
Efficient Content-Based Sparse Attention with Routing Transformers
Aurko Roy, Mohammad Saffar, Ashish Vaswani, and David Grangier. 2020 · 2003
Earlier work this paper cites.
Noam Shazeer, Zhenzhong Lan, Youlong Cheng, Nan Ding, and Le Hou. 2020 · 2003
Earlier work this paper cites.
Longformer: The Long-Document Transformer
Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020 · 2004
Earlier work this paper cites.
Synthesizer: Rethinking Self-Attention in Transformer Models
Yi Tay, Dara Bahri, Donald Metzler, Da-Cheng Juan, Zhe Zhao, and Che Zheng. 2020a · 2005
Earlier work this paper cites.
Masked Language Modeling for Proteins via Linearly Scalable Long-Context Transformers
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, David Belanger, Lucy Colwell, and Adrian Weller. 2020a · 2006
Earlier work this paper cites.
Multi-Head Attention: Collaborate Instead of Concatenate
Jean-Baptiste Cordonnier, Andreas Loukas, and Martin Jaggi. 2020 · 2006
Earlier work this paper cites.
DeBERTa: Decoding-enhanced BERT with Disentangled Attention
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2020a · 2006
Earlier work this paper cites.
Rethinking Positional Encoding in Language Pre-training
Guolin Ke, Di He, and Tie-Yan Liu. 2020 · 2006
Earlier work this paper cites.
GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen. 2020 · 2006
Earlier work this paper cites.
Linformer: Self-Attention with Linear Complexity
Sinong Wang, Belinda Z. Li, Madian Khabsa, Han Fang, and Hao Ma. 2020a · 2006
Earlier work this paper cites.
BERT Loses Patience: Fast and Robust Inference with Early Exit
Wangchunshu Zhou, Canwen Xu, Tao Ge, Julian McAuley, Ke Xu, and Furu Wei. 2020 · 2006
Earlier work this paper cites.
Random Features for Large-Scale Kernel Machines. In Proceedings of NeurIPS . 1177–1184
Ali Rahimi and Benjamin Recht. 2007 · 2007
Earlier work this paper cites.
Fast Transformers with Clustered Attention
Apoorv Vyas, Angelos Katharopoulos, and François Fleuret. 2020 · 2007
Earlier work this paper cites.
Big Bird: Transformers for Longer Sequences
Manzil Zaheer, Guru Guruganesh, Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, and Amr Ahmed. 2020 · 2007
Earlier work this paper cites.
DeLighT: Very Deep and Light-weight Transformer
Sachin Mehta, Marjan Ghazvininejad, Srinivasan Iyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2020 · 2008
Earlier work this paper cites.
Adding Recurrence to Pretrained Transformers for Improved Efficiency and Context Size
Davis Yoshida, Allyson Ettinger, and Kevin Gimpel. 2020 · 2008
Earlier work this paper cites.
Rethinking Attention with Performers
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, David Belanger, Lucy Colwell, and Adrian Weller. 2020b · 2009
Earlier work this paper cites.
Developing Real-time Streaming Transformer Transducer for Speech Recognition on Large-scale Dataset
Xie Chen, Yu Wu, Zhenghao Wang, Shujie Liu, and Jinyu Li. 2021 · 2010
Earlier work this paper cites.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2020 · 2010
Earlier work this paper cites.
Memformer: The Memory-Augmented Transformer
Qingyang Wu, Zhenzhong Lan, Jing Gu, and Zhou Yu. 2020a · 2010
Earlier work this paper cites.
Deformable DETR: Deformable Transformers for End-to-End Object Detection
Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. 2020 · 2010
Earlier work this paper cites.
End-to-End Object Detection with Adaptive Clustering Transformer
Minghang Zheng, Peng Gao, Xiaogang Wang, Hongsheng Li, and Hao Dong. 2020a · 2011
Earlier work this paper cites.
ERNIE-DOC: The Retrospective Long-Document Modeling Transformer
Siyu Ding, Junyuan Shang, Shuohuan Wang, Yu Sun, Hao Tian, Hua Wu, and Haifeng Wang. 2020 · 2012
Earlier work this paper cites.
A Survey on Visual Transformer
Kai Han, Yunhe Wang, Hanting Chen, Xinghao Chen, Jianyuan Guo, Zhenhua Liu, Yehui Tang, An Xiao, Chunjing Xu, Yixing Xu, Zhaohui Yang, Yiman Zhang, and Dacheng Tao. 2021a · 2012
Earlier work this paper cites.
RealFormer: Transformer Likes Residual Attention
Ruining He, Anirudh Ravula, Bhargav Kanagal, and Joshua Ainslie. 2020b · 2012
Earlier work this paper cites.
Sequence to Sequence Learning with Neural Networks. In Proceedings of NeurIPS . 3104–3112
Ilya Sutskever, Oriol Vinyals, and Quoc V. Le. 2014 · 2014
Earlier work this paper cites.
Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift. In Proceedings of ICML . 448–456
Sergey Ioffe and Christian Szegedy. 2015 · 2015
Earlier work this paper cites.
Lei Jimmy Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. 2016 · 2016
Earlier work this paper cites.
Adaptive Computation Time for Recurrent Neural Networks
Alex Graves. 2016 · 2016
Earlier work this paper cites.
Deep Residual Learning for Image Recognition. In Proceedings CVPR . 770–778
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
WaveNet: A Generative Model for Raw Audio. In Proceedings of ISCA . 125
Aäron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew W. Senior, and Koray Kavukcuoglu. 2016 · 2016
Earlier work this paper cites.
Language Modeling with Gated Convolutional Networks. In Proceedings of ICML . 933–941
Yann N. Dauphin, Angela Fan, Michael Auli, and David Grangier. 2017 · 2017
Earlier work this paper cites.
Convolutional Sequence to Sequence Learning. In Proceedings of ICML . 1243–1252
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N. Dauphin. 2017 · 2017
Earlier work this paper cites.
OpenNMT: Open-Source Toolkit for Neural Machine Translation. In Proceedings of ACL . 67–72
Guillaume Klein, Yoon Kim, Yuntian Deng, Jean Senellart, and Alexander Rush. 2017 · 2017
Earlier work this paper cites.
Learning multiple visual domains with residual adapters. In Proceedings of NeurIPS . 506–516
Sylvestre-Alvise Rebuffi, Hakan Bilen, and Andrea Vedaldi. 2017 · 2017
Earlier work this paper cites.
Dynamic Routing Between Capsules. In Proceedings of NeurIPS . 3856–3866
Sara Sabour, Nicholas Frosst, and Geoffrey E. Hinton. 2017 · 2017
Earlier work this paper cites.
Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer. In Proceedings of ICLR
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc V. Le, Geoffrey E. Hinton, and Jeff Dean. 2017 · 2017
Earlier work this paper cites.
Attention is All you Need. In Proceedings of NeurIPS . 5998–6008
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Training Deeper Neural Machine Translation Models with Transparent Attention. In Proceedings of EMNLP . Brussels, Belgium, 3028–3033
Ankur Bapna, Mia Chen, Orhan Firat, Yuan Cao, and Yonghui Wu. 2018 · 2018
Cited alongside, same era.
Relational inductive biases, deep learning, and graph networks
Peter W. Battaglia, Jessica B. Hamrick, Victor Bapst, Alvaro Sanchez-Gonzalez, Vinicius Zambaldi, Mateusz Malinowski, Andrea Tacchetti, David Raposo, Adam Santoro, Ryan Faulkner, Caglar Gulcehre, Francis Song, Andrew Ballard, Justin Gilmer, George Dahl, Ashish Vaswani, Kelsey Allen, Charles Nash, Victoria Langston, Chris Dyer, Nicolas Heess, Daan Wierstra, Pushmeet Kohli, Matt Botvinick, Oriol Vinyals, Yujia Li, and Razvan Pascanu. 2018 · 2018
Cited alongside, same era.
Speech-Transformer: A No-Recurrence Sequence-to-Sequence Model for Speech Recognition. In Proceedings of ICASSP . 5884–5888
Linhao Dong, Shuang Xu, and Bo Xu. 2018 · 2018
Cited alongside, same era.
Matrix capsules with EM routing. In Proceedings of ICLR
Geoffrey E. Hinton, Sara Sabour, and Nicholas Frosst. 2018 · 2018
Cited alongside, same era.
How much Position Information Do Convolutional Neural Networks Encode?. In Proceedings of ICLR
Md. Amirul Islam, Sen Jia, and Neil D. B. Bruce. 2020 · 2020
Later among the works it cites.
Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention. In Proceedings of ICML . 5156–5165
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret. 2020 · 2020
Later among the works it cites.
T-GSA: Transformer with Gaussian-Weighted Self-Attention for Speech Enhancement. In 2020 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2020, Barcelona, Spain, May 4-8, 2020 . IEEE, 6649–6653
Jaeyoung Kim, Mostafa El-Khamy, and Jungwon Lee. 2020 · 2020
Later among the works it cites.
Reformer: The Efficient Transformer. In Proceedings of ICLR
Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya. 2020 · 2020
Later among the works it cites.
BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. In Proceedings of ACL . 7871–7880
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Multi-Head Attention with Disagreement Regularization. In Proceedings of EMNLP . Brussels, Belgium, 2897–2903
Jian Li, Zhaopeng Tu, Baosong Yang, Michael R. Lyu, and Tong Zhang. 2018 · 2018
Cited alongside, same era.
Generating Wikipedia by Summarizing Long Sequences. In Proceedings of ICLR
Peter J. Liu, Mohammad Saleh, Etienne Pot, Ben Goodrich, Ryan Sepassi, Lukasz Kaiser, and Noam Shazeer. 2018 · 2018
Cited alongside, same era.
Document-Level Neural Machine Translation with Hierarchical Attention Networks. In Proceedings of EMNLP . Brussels, Belgium, 2947–2954
Lesly Miculicich, Dhananjay Ram, Nikolaos Pappas, and James Henderson. 2018 · 2018
Cited alongside, same era.
Image Transformer. In Proceedings of ICML . 4052–4061
Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Noam Shazeer, Alexander Ku, and Dustin Tran. 2018 · 2018
Cited alongside, same era.
FiLM: Visual Reasoning with a General Conditioning Layer. In Proceedings of AAAI . 3942–3951
Ethan Perez, Florian Strub, Harm de Vries, Vincent Dumoulin, and Aaron C. Courville. 2018 · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Cited alongside, same era.
Searching for Activation Functions. In Proceedings of ICLR
Prajit Ramachandran, Barret Zoph, and Quoc V. Le. 2018 · 2018
Cited alongside, same era.
Representation Learning with Contrastive Predictive Coding
Aäron van den Oord, Yazhe Li, and Oriol Vinyals. 2018 · 2018
Cited alongside, same era.
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020 · 2020
Later among the works it cites.
UNIMO: Towards Unified-Modal Understanding and Generation via Cross-Modal Contrastive Learning
Wei Li, Can Gao, Guocheng Niu, Xinyan Xiao, Hao Liu, Jiachen Liu, Hua Wu, and Haifeng Wang. 2020a · 2020
Later among the works it cites.
SAC: Accelerating and Structuring Self-Attention via Sparse Adaptive Connection. In Proceedings of NeurIPS
Xiaoya Li, Yuxian Meng, Mingxin Zhou, Qinghong Han, Fei Wu, and Jiwei Li. 2020b · 2020
Later among the works it cites.
FLAT: Chinese NER Using Flat-Lattice Transformer. In Proceedings of ACL . 6836–6842
Xiaonan Li, Hang Yan, Xipeng Qiu, and Xuanjing Huang. 2020c · 2020
Later among the works it cites.
Understanding the Difficulty of Training Transformers. In Proceedings of EMNLP . 5747–5763
Liyuan Liu, Xiaodong Liu, Jianfeng Gao, Weizhu Chen, and Jiawei Han. 2020a · 2020
Later among the works it cites.
Understanding and Improving Transformer From a Multi-Particle Dynamic System Point of View
Yiping Lu, Zhuohan Li, Di He, Zhiqing Sun, Bin Dong, Tao Qin, Liwei Wang, and Tie-Yan Liu. 2020 · 2020
Later among the works it cites.
Improving Transformer Models by Reordering their Sublayers. In Proceedings of ACL . Online, 2996–3005
Ofir Press, Noah A. Smith, and Omer Levy. 2020 · 2020
Later among the works it cites.
Pre-trained Models for Natural Language Processing: A Survey
Xipeng Qiu, TianXiang Sun, Yige Xu, Yunfan Shao, Ning Dai, and Xuanjing Huang. 2020 · 2020
Later among the works it cites.
Compressive Transformers for Long-Range Sequence Modelling. In Proceedings of ICLR
Jack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier, and Timothy P. Lillicrap. 2020 · 2020
Later among the works it cites.
PowerNorm: Rethinking Batch Normalization in Transformers. In Proceedings of ICML . 8741–8751
Sheng Shen, Zhewei Yao, Amir Gholami, Michael W. Mahoney, and Kurt Keutzer. 2020 · 2020
Later among the works it cites.
VL-BERT: Pre-training of Generic Visual-Linguistic Representations. In Proceedings of ICLR
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. 2020 · 2020
Later among the works it cites.
A Comprehensive Survey on Graph Neural Networks
Zonghan Wu, Shirui Pan, Fengwen Chen, Guodong Long, Chengqi Zhang, and Philip S. Yu. 2021a · 2020
Later among the works it cites.
DeeBERT: Dynamic Early Exiting for Accelerating BERT Inference. In Proceedings of ACL . 2246–2251
Ji Xin, Raphael Tang, Jaejun Lee, Yaoliang Yu, and Jimmy Lin. 2020 · 2020
Later among the works it cites.
On Layer Normalization in the Transformer Architecture. In Proceedings of ICML . 10524–10533
Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tie-Yan Liu. 2020 · 2020
Later among the works it cites.
On the Sub-layer Functionalities of Transformer Decoder. In Findings of EMNLP . Online, 4799–4811
Yilin Yang, Longyue Wang, Shuming Shi, Prasad Tadepalli, Stefan Lee, and Zhaopeng Tu. 2020 · 2020
Later among the works it cites.
Hard-Coded Gaussian Attention for Neural Machine Translation. In Proceedings of ACL . Online, 7689–7700
Weiqiu You, Simeng Sun, and Mohit Iyyer. 2020 · 2020
Later among the works it cites.
Improving End-to-End Speech Synthesis with Local Recurrent Neural Network Enhanced Transformer. In Proceedings of ICASSP . 6734–6738
Yibin Zheng, Xinhui Li, Fenglong Xie, and Li Lu. 2020b · 2020
Later among the works it cites.
ViViT: A Video Vision Transformer
Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, and Cordelia Schmid. 2021 · 2021
Closest in time.
Conditional Positional Encodings for Vision Transformers
Xiangxiang Chu, Zhi Tian, Bo Zhang, Xinlong Wang, Xiaolin Wei, Huaxia Xia, and Chunhua Shen. 2021 · 2021
Closest in time.
CogView: Mastering Text-to-Image Generation via Transformers
Ming Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng, Chang Zhou, Da Yin, Junyang Lin, Xu Zou, Zhou Shao, Hongxia Yang, and Jie Tang. 2021 · 2021
Closest in time.
Attention is Not All You Need: Pure Attention Loses Rank Doubly Exponentially with Depth
Yihe Dong, Jean-Baptiste Cordonnier, and Andreas Loukas. 2021 · 2021
Closest in time.
Mask Attention Networks: Rethinking and Strengthen Transformer. In Proceedings of NAACL . 1692–1701
Zhihao Fan, Yeyun Gong, Dayiheng Liu, Zhongyu Wei, Siyuan Wang, Jian Jiao, Nan Duan, Ruofei Zhang, and Xuanjing Huang. 2021a · 2021
Closest in time.
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
William Fedus, Barret Zoph, and Noam Shazeer. 2021 · 2021
Closest in time.
TransGAN: Two Transformers Can Make One Strong GAN
Yifan Jiang, Shiyu Chang, and Zhangyang Wang. 2021 · 2021
Closest in time.
Transformers in Vision: A Survey
Salman Khan, Muzammal Naseer, Munawar Hayat, Syed Waqas Zamir, Fahad Shahbaz Khan, and Mubarak Shah. 2021 · 2021
Closest in time.
Accelerating BERT Inference for Sequence Labeling via Early-Exit
Xiaonan Li, Yunfan Shao, Tianxiang Sun, Hang Yan, Xipeng Qiu, and Xuanjing Huang. 2021 · 2021
Closest in time.
M6: A Chinese Multimodal Pretrainer
Junyang Lin, Rui Men, An Yang, Chang Zhou, Ming Ding, Yichang Zhang, Peng Wang, Ang Wang, Le Jiang, Xianyan Jia, Jie Zhang, Jianwei Zhang, Xu Zou, Zhikang Li, Xiaodong Deng, Jie Liu, Jinbao Xue, Huiling Zhou, Jianxin Ma, Jin Yu, Yong Li, Wei Lin, Jingren Zhou, Jie Tang, and Hongxia Yang. 2021 · 2021
Closest in time.
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. 2021 · 2021
Closest in time.
Luna: Linear Unified Nested Attention
Xuezhe Ma, Xiang Kong, Sinong Wang, Chunting Zhou, Jonathan May, Hao Ma, and Luke Zettlemoyer. 2021 · 2021
Closest in time.
Random Feature Attention. In Proceedings of ICLR
Hao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz, Noah Smith, and Lingpeng Kong. 2021 · 2021
Closest in time.
Conditionally Adaptive Multi-Task Learning: Improving Transfer Learning in NLP Using Fewer Parameters & Less Data. In Proceedings of ICLR
Jonathan Pilault, Amine El hattami, and Christopher Pal. 2021 · 2021
Closest in time.
Zero-Shot Text-to-Image Generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. 2021 · 2021
Closest in time.
Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences
Alexander Rives, Joshua Meier, Tom Sercu, Siddharth Goyal, Zeming Lin, Jason Liu, Demi Guo, Myle Ott, C. Lawrence Zitnick, Jerry Ma, and Rob Fergus. 2021 · 2021
Closest in time.
Hash Layers For Large Sparse Models
Stephen Roller, Sainbayar Sukhbaatar, Arthur Szlam, and Jason Weston. 2021 · 2021
Closest in time.
Linear Transformers Are Secretly Fast Weight Memory Systems
Imanol Schlag, Kazuki Irie, and Jürgen Schmidhuber. 2021 · 2021
Closest in time.
Temporal Context Aggregation for Video Retrieval With Contrastive Learning. In Proceedings of WACV . 3268–3278
Jie Shao, Xin Wen, Bingchen Zhao, and Xiangyang Xue. 2021 · 2021
Closest in time.
RoFormer: Enhanced Transformer with Rotary Position Embedding
Jianlin Su, Yu Lu, Shengfeng Pan, Bo Wen, and Yunfeng Liu. 2021 · 2021
Closest in time.
Early Exiting with Ensemble Internal Classifiers
Tianxiang Sun, Yunhua Zhou, Xiangyang Liu, Xinyu Zhang, Hao Jiang, Zhao Cao, Xuanjing Huang, and Xipeng Qiu. 2021 · 2021
Closest in time.
On Position Embeddings in BERT, url = https://openreview.net/forum?id=onxoVA9FxMw, year = 2021. In Proceedings of ICLR
Benyou Wang, Lifeng Shang, Christina Lioma, Xin Jiang, Hao Yang, Qun Liu, and Jakob Grue Simonsen. [n.d.] · 2021
Closest in time.
Predictive Attention Transformer: Improving Transformer with Attention Map Prediction
Yujing Wang, Yaming Yang, Jiangang Bai, Mingliang Zhang, Jing Bai, Jing Yu, Ce Zhang, and Yunhai Tong. 2021 · 2021
Closest in time.
Nyströmformer: A Nyström-based Algorithm for Approximating Self-Attention
Yunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan, Glenn Fung, Yin Li, and Vikas Singh. 2021 · 2021
Closest in time.
Exploring Sparse Expert Models and Beyond
An Yang, Junyang Lin, Rui Men, Chang Zhou, Le Jiang, Xianyan Jia, Ang Wang, Jie Zhang, Jiamang Wang, Yong Li, Di Zhang, Wei Lin, Lin Qu, Jingren Zhou, and Hongxia Yang. 2021 · 2021
Closest in time.
LazyFormer: Self Attention with Lazy Update
Chengxuan Ying, Guolin Ke, Di He, and Tie-Yan Liu. 2021 · 2021
Closest in time.
SETransformer: Speech Enhancement Transformer
Weiwei Yu, Jian Zhou, HuaBin Wang, and Liang Tao. 2021 · 2021
Closest in time.
Poolingformer: Long Document Modeling with Pooling Attention
Hang Zhang, Yeyun Gong, Yelong Shen, Weisheng Li, Jiancheng Lv, Nan Duan, and Weizhu Chen. 2021 · 2021
Closest in time.
Memory-Efficient Differentiable Transformer Architecture Search
Yuekai Zhao, Li Dong, Yelong Shen, Zhihua Zhang, Furu Wei, and Weizhu Chen. 2021 · 2021
Closest in time.
Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting. In Proceedings of AAAI
Haoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang, Jianxin Li, Hui Xiong, and Wancai Zhang. 2021 · 2021
Closest in time.
Self-Attention with Relative Position Representations. In Proceedings of HLT-NAACL . New Orleans, Louisiana, 464–468
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. 2018 · 2074
Closest in time.