Fetching the paper…
Reading the bibliography…
Encoder layer fusion (EncoderFusion) is a technique to fuse all the encoder layers (instead of the uppermost layer) for sequence-to-sequence (Seq2Seq) models, which has proven effective on various NLP tasks.
Bleu: A method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Better evaluation for grammatical error correction
Daniel Dahlmeier and Hwee Tou Ng · 2012
Earlier work this paper cites.
A simple, fast, and effective reparameterization of IBM model 2
Chris Dyer, Victor Chahuneau, and Noah A. Smith · 2013
Earlier work this paper cites.
A deep architecture for matching short texts
Zhengdong Lu and Hang Li · 2013
Earlier work this paper cites.
Regularization of neural networks using dropconnect
Li Wan, Matthew Zeiler, Sixin Zhang, Yann Le Cun, and Rob Fergus · 2013
Earlier work this paper cites.
Convolutional neural network architectures for matching natural language sentences
Baotian Hu, Zhengdong Lu, Hang Li, and Qingcai Chen · 2014
Earlier work this paper cites.
Fully convolutional networks for semantic segmentation
Jonathan Long, Evan Shelhamer, and Trevor Darrell · 2014
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
On using monolingual corpora in neural machine translation
Caglar Gulcehre, Orhan Firat, Kelvin Xu, Kyunghyun Cho, Loic Barrault, Huei-Chi Lin, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2015
Earlier work this paper cites.
Layer normalization
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Learning to refine object segments
Pedro H. O. Pinheiro, Tsung-Yi Lin, Ronan Collobert, and Piotr Dollár · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al · 2016
Earlier work this paper cites.
Exfuse: Enhancing feature fusion for semantic segmentation
Zhenli Zhang, Xiangyu Zhang, Chao Peng, Dazhi Cheng, and Jian Sun · 2016
Earlier work this paper cites.
What do neural machine translation models learn about morphology?
Yonatan Belinkov, Nadir Durrani, Fahim Dalvi, Hassan Sajjad, and James Glass · 2017
Cited alongside, same era.
Convolutional sequence to sequence learning
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N. Dauphin · 2017
Cited alongside, same era.
Feature pyramid networks for object detection
Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie · 2017
Cited alongside, same era.
Using the Output Embedding to Improve Language Models
Ofir Press and Lior Wolf · 2017
Cited alongside, same era.
Cold fusion: Training seq2seq models together with language models
Anuroop Sriram, Heewoo Jun, Sanjeev Satheesh, and Adam Coates · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Dynamic layer aggregation for neural machine translation with routing-by-agreement
Zi-Yi Dou, Zhaopeng Tu, Xing Wang, Longyue Wang, Shuming Shi, and Tong Zhang · 2019
Later among the works it cites.
Representation degeneration problem in training natural language generation models
Jun Gao, Di He, Xu Tan, Tao Qin, Liwei Wang, and Tieyan Liu · 2019
Later among the works it cites.
Mask-predict: Parallel decoding of conditional masked language models
Marjan Ghazvininejad, Omer Levy, Yinhan Liu, and Luke Zettlemoyer · 2019
Later among the works it cites.
An empirical study of incorporating pseudo data into grammatical error correction
Shun Kiyono, Jun Suzuki, Masato Mita, Tomoya Mizumoto, and Kentaro Inui · 2019
Later among the works it cites.
Hallucinations in neural machine translation
Katherine Lee, Orhan Firat, Ashish Agarwal, Clara Fannjiang, and David Sussillo · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Training deeper neural machine translation models with transparent attention
Ankur Bapna, Mia Xu Chen, Orhan Firat, Yuan Cao, and Yonghui Wu · 2018
Cited alongside, same era.
Fine-grained attention mechanism for neural machine translation
Heeyoul Choi, Kyunghyun Cho, and Yoshua Bengio · 2018
Cited alongside, same era.
A multilayer convolutional encoder-decoder neural network for grammatical error correction
Shamil Chollampatt and Hwee Tou Ng · 2018
Cited alongside, same era.
How much attention do you need? a granular analysis of neural machine translation architectures
Tobias Domhan · 2018
Cited alongside, same era.
Exploiting deep representations for neural machine translation
Zi-Yi Dou, Zhaopeng Tu, Xing Wang, Shuming Shi, and Tong Zhang · 2018
Cited alongside, same era.
Layer-wise coordination between encoder and decoder for neural machine translation
Tianyu He, Xu Tan, Yingce Xia, Di He, Tao Qin, Zhibo Chen, and Tie-Yan Liu · 2018
Cited alongside, same era.
Xuebo Liu, Derek F. Wong, Yang Liu, Lidia S. Chao, Tong Xiao, and Jingbo Zhu · 2019
Later among the works it cites.
Improving lexical choice in neural machine translation
Toan Nguyen and David Chiang · 2019
Later among the works it cites.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli · 2019
Later among the works it cites.
Leveraging pre-trained checkpoints for sequence generation tasks
Sascha Rothe, Shashi Narayan, and Aliaksei Severyn · 2019
Later among the works it cites.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov · 2019
Later among the works it cites.
Pay less attention with lightweight and dynamic convolutions
Felix Wu, Angela Fan, Alexei Baevski, Yann Dauphin, and Michael Auli · 2019
Later among the works it cites.
Encoder-decoder models can benefit from pre-trained masked language models in grammatical error correction
Masahiro Kaneko, Masato Mita, Shun Kiyono, Jun Suzuki, and Kentaro Inui · 2020
Closest in time.
Neuron interaction based representation composition for neural machine translation
Jian Li, Xing Wang, Baosong Yang, Shuming Shi, Michael R Lyu, and Zhaopeng Tu · 2020
Closest in time.
Layer-wise cross-view decoding for sequence-to-sequence learning
Fenglin Liu, Xuancheng Ren, Guangxiang Zhao, and Xu Sun · 2020
Closest in time.
Rethinking batch normalization in transformers
Sheng Shen, Zhewei Yao, Amir Gholami, Michael Mahoney, and Kurt Keutzer · 2020
Closest in time.
On exposure bias, hallucination and domain shift in neural machine translation
Chaojun Wang and Rico Sennrich · 2020
Closest in time.
Rethinking the value of transformer components
Wenxuan Wang and Zhaopeng Tu · 2020
Closest in time.
Pegasus: Pre-training with extracted gap-sentences for abstractive summarization
Jingqing Zhang, Yao Zhao, Mohammad Saleh, and Peter J Liu · 2020
Closest in time.