Fetching the paper…
Reading the bibliography…
Transformer architecture has become the de-facto model for many machine learning tasks from natural language processing and computer vision.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 1910
Earlier work this paper cites.
Learning representations by back-propagating errors
David E Rumelhart, Geoffrey E Hinton, and Ronald J Williams. 1986 · 1986
Earlier work this paper cites.
Newsweeder: Learning to filter netnews
Ken Lang. 1995 · 1995
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Longformer: The long-document transformer
Iz Beltagy, Matthew E Peters, and Arman Cohan. 2020 · 2004
Earlier work this paper cites.
Big bird: Transformers for longer sequences
Manzil Zaheer, Guru Guruganesh, Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et al. 2020 · 2007
Earlier work this paper cites.
Efficient transformers: A survey
Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler. 2020 · 2009
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011 · 2011
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Jake Zhao, and Yann LeCun. 2015 · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. 2016 · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
Dan Hendrycks and Kevin Gimpel. 2016 · 2016
Earlier work this paper cites.
Conditional deep learning for energy-efficient and enhanced pattern recognition
Priyadarshini Panda, Abhronil Sengupta, and Kaushik Roy. 2016 · 2016
Earlier work this paper cites.
Branchynet: Fast inference via early exiting from deep neural networks
Surat Teerapittayanon, Bradley McDanel, and Hsiang-Tsung Kung. 2016 · 2016
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole. 2017 · 2017
Earlier work this paper cites.
A structured self-attentive sentence embedding
Zhouhan Lin, Minwei Feng, Cícero Nogueira dos Santos, Mo Yu, Bing Xiang, Bowen Zhou, and Yoshua Bengio. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Learning to skim text
Adams Wei Yu, Hongrae Lee, and Quoc Le. 2017 · 2017
Cited alongside, same era.
Skip RNN: learning to skip state updates in recurrent neural networks
Víctor Campos, Brendan Jou, Xavier Giró-i-Nieto, Jordi Torres, and Shih-Fu Chang. 2018 · 2018
Cited alongside, same era.
Neural speed reading via skim-rnn
Min Joon Seo, Sewon Min, Ali Farhadi, and Hannaneh Hajishirzi. 2018 · 2018
Cited alongside, same era.
Universal transformers
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Lukasz Kaiser. 2019 · 2019
Cited alongside, same era.
FastBERT: a self-distilling BERT with adaptive inference time
Weijie Liu, Peng Zhou, Zhiruo Wang, Zhe Zhao, Haotang Deng, and Qi Ju. 2020 · 2020
Later among the works it cites.
Q-bert: A bert-based framework for computing sparql similarity in natural language
Chunpei Wang and Xiaowang Zhang. 2020 · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020 · 2020
Later among the works it cites.
Lite transformer with long-short range attention
Zhanghao Wu, Zhijian Liu, Ji Lin, Yujun Lin, and Song Han. 2020 · 2020
Later among the works it cites.
DeeBERT: Dynamic early exiting for accelerating BERT inference
Ji Xin, Raphael Tang, Jaejun Lee, Yaoliang Yu, and Jimmy Lin. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Neural speed reading with structural-jump-lstm
Christian Hansen, Casper Hansen, Stephen Alstrup, Jakob Grue Simonsen, and Christina Lioma. 2019 · 2019
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019 · 2019
Cited alongside, same era.
DeFormer: Decomposing pre-trained transformers for faster question answering
Qingqing Cao, Harsh Trivedi, Aruna Balasubramanian, and Niranjan Balasubramanian. 2020 · 2020
Cited alongside, same era.
Power-bert: Accelerating BERT inference via progressive word-vector elimination
Saurabh Goyal, Anamitra Roy Choudhury, Saurabh Raje, Venkatesan T. Chakaravarthy, Yogish Sabharwal, and Ashish Verma. 2020 · 2020
Cited alongside, same era.
How far does BERT look at: Distance-based clustering and analysis of BERT’s attention
Yue Guan, Jingwen Leng, Chao Li, Quan Chen, and Minyi Guo. 2020 · 2020
Cited alongside, same era.
Bert loses patience: Fast and robust inference with early exit
Wangchunshu Zhou, Canwen Xu, Tao Ge, Julian McAuley, Ke Xu, and Furu Wei. 2020 · 2020
Later among the works it cites.
Efficientbert: Progressively searching multilayer perceptron via warm-up knowledge distillation
Chenhe Dong, Guangrun Wang, Hang Xu, Jiefeng Peng, Xiaozhe Ren, and Xiaodan Liang. 2021 · 2021
Later among the works it cites.
On the transformer growth for progressive BERT training
Xiaotao Gu, Liyuan Liu, Hongkun Yu, Jing Li, Chen Chen, and Jiawei Han. 2021 · 2021
Later among the works it cites.
Block-skim: Efficient question answering for transformer
Yue Guan, Zhengyi Li, Jingwen Leng, Zhouhan Lin, Minyi Guo, and Yuhao Zhu. 2021 · 2021
Later among the works it cites.
Magic pyramid: Accelerating inference with early exiting and token pruning
Xuanli He, Iman Keivanloo, Yi Xu, Xiang He, Belinda Zeng, Santosh Rajagopalan, and Trishul Chilimbi. 2021 · 2021
Later among the works it cites.
Length-adaptive transformer: Train once with length drop, use anytime with search
Gyuwan Kim and Kyunghyun Cho. 2021 · 2021
Later among the works it cites.
Learned token pruning for transformers
Sehoon Kim, Sheng Shen, David Thorsley, Amir Gholami, Woosuk Kwon, Joseph Hassoun, and Kurt Keutzer. 2021 · 2021
Later among the works it cites.
Spatten: Efficient sparse attention architecture with cascade token and head pruning
Hanrui Wang, Zhekai Zhang, and Song Han. 2021 · 2021
Later among the works it cites.
TR-BERT: Dynamic token reduction for accelerating BERT inference
Deming Ye, Yankai Lin, Yufei Huang, and Maosong Sun. 2021 · 2021
Later among the works it cites.
SQuant: On-the-fly data-free quantization via diagonal hessian approximation
Cong Guo, Yuxian Qiu, Jingwen Leng, Xiaotian Gao, Chen Zhang, Yunxin Liu, Fan Yang, Yuhao Zhu, and Minyi Guo. 2022 · 2022
Closest in time.