Fetching the paper…
Reading the bibliography…
We introduce EdgeFormer -- a parameter-efficient Transformer for on-device seq2seq generation under the strict computation and memory constraints.
Pay less attention with lightweight and dynamic convolutions
Felix Wu, Angela Fan, Alexei Baevski, Yann Dauphin, and Michael Auli · 1901
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Mining revision log of language learning sns for automated japanese error correction of second language learners
Tomoya Mizumoto, Mamoru Komachi, Masaaki Nagata, and Yuji Matsumoto · 2011
Earlier work this paper cites.
A new dataset and method for automatically grading esol texts
Helen Yannakoudakis, Ted Briscoe, and Ben Medlock · 2011
Earlier work this paper cites.
Better evaluation for grammatical error correction
Daniel Dahlmeier and Hwee Tou Ng · 2012
Earlier work this paper cites.
The CoNLL-2013 shared task on grammatical error correction
Hwee Tou Ng, Siew Mei Wu, Yuanbin Wu, Christian Hadiwinoto, and Joel Tetreault · 2013
Earlier work this paper cites.
The conll-2014 shared task on grammatical error correction
Hwee Tou Ng, Siew Mei Wu, Ted Briscoe, Christian Hadiwinoto, Raymond Hendy Susanto, and Christopher Bryant · 2014
Earlier work this paper cites.
Sequence-level knowledge distillation
Yoon Kim and Alexander M Rush · 2016
Earlier work this paper cites.
Learning to ask: Neural question generation for reading comprehension
Xinya Du, Junru Shao, and Claire Cardie · 2017
Earlier work this paper cites.
Convolutional sequence to sequence learning
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N Dauphin · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Łukasz Kaiser · 2018
Earlier work this paper cites.
Taku Kudo and John Richardson · 2018
Earlier work this paper cites.
Shashi Narayan, Shay B Cohen, and Mirella Lapata · 2018
Earlier work this paper cites.
Scaling neural machine translation
Myle Ott, Sergey Edunov, David Grangier, and Michael Auli · 2018
Earlier work this paper cites.
A call for clarity in reporting bleu scores
Matt Post · 2018
Cited alongside, same era.
Improving sequence-to-sequence pre-training via sequence span rewriting
Wangchunshu Zhou, Tao Ge, Canwen Xu, Ke Xu, and Furu Wei · 2018
Cited alongside, same era.
The bea-2019 shared task on grammatical error correction
Christopher Bryant, Mariano Felice, Øistein E Andersen, and Ted Briscoe · 2019
Cited alongside, same era.
Reinforcement learning based graph-to-sequence model for natural question generation
Yu Chen, Lingfei Wu, and Mohammed J Zaki · 2019
Cited alongside, same era.
On-device machine learning: An algorithms and learning theory perspective
Sauptik Dhar, Junyao Guo, Jiayi Liu, Samarth Tripathi, Unmesh Kurup, and Mohak Shah · 2019
Cited alongside, same era.
Bert-of-theseus: Compressing bert by progressive module replacing
Canwen Xu, Wangchunshu Zhou, Tao Ge, Furu Wei, and Ming Zhou · 2020
Later among the works it cites.
Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models
Elad Ben Zaken, Shauli Ravfogel, and Yoav Goldberg · 2021
Later among the works it cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity, 2021
William Fedus, Barret Zoph, and Noam Shazeer · 2021
Later among the works it cites.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Later among the works it cites.
The power of scale for parameter-efficient prompt tuning
Brian Lester, Rami Al-Rfou, and Noah Constant · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Unified language model pre-training for natural language understanding and generation
Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon · 2019
Cited alongside, same era.
Reducing transformer depth on demand with structured dropout
Angela Fan, Edouard Grave, and Armand Joulin · 2019
Cited alongside, same era.
Parameter-efficient transfer learning for nlp
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly · 2019
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2019
Cited alongside, same era.
Improving the efficiency of grammatical error correction with erroneous span detection and correction
Mengyun Chen, Tao Ge, Xingxing Zhang, Furu Wei, and Ming Zhou · 2020
Cited alongside, same era.
Deep encoder, shallow decoder: Reevaluating non-autoregressive machine translation
Jungo Kasai, Nikolaos Pappas, Hao Peng, James Cross, and Noah A Smith · 2020
Cited alongside, same era.
Prefix-tuning: Optimizing continuous prompts for generation
Xiang Lisa Li and Percy Liang · 2021
Later among the works it cites.
Shuming Ma, Li Dong, Shaohan Huang, Dongdong Zhang, Alexandre Muzio, Saksham Singhal, Hany Hassan Awadalla, Xia Song, and Furu Wei · 2021
Later among the works it cites.
Shapeshifter: a parameter-efficient transformer using factorized reshaped matrices
Aliakbar Panahi, Seyran Saeedi, and Tom Arodz · 2021
Later among the works it cites.
Subformer: Exploring weight sharing for parameter efficiency in generative transformers
Machel Reid, Edison Marrese-Taylor, and Yutaka Matsuo · 2021
Later among the works it cites.
Instantaneous grammatical error correction with shallow aggressive decoding
Xin Sun, Tao Ge, Furu Wei, and Houfeng Wang · 2021
Later among the works it cites.
Lessons on parameter sharing across layers in transformers
Sho Takase and Shun Kiyono · 2021
Later among the works it cites.
Edgebert: Sentence-level energy optimizations for latency-aware multi-task nlp inference
Thierry Tambe, Coleman Hooper, Lillian Pentecost, Tianyu Jia, En-Yu Yang, Marco Donato, Victor Sanh, Paul Whatmough, Alexander M Rush, David Brooks, et al · 2021
Later among the works it cites.
Aston Zhang, Yi Tay, Shuai Zhang, Alvin Chan, Anh Tuan Luu, Siu Cheung Hui, and Jie Fu · 2021
Later among the works it cites.
Dq-bart: Efficient sequence-to-sequence model via joint distillation and quantization
Zheng Li, Zijian Wang, Ming Tan, Ramesh Nallapati, Parminder Bhatia, Andrew Arnold, Bing Xiang, and Dan Roth · 2022
Closest in time.