Fetching the paper…
Reading the bibliography…
Checkpoint averaging is a simple and effective method to boost the performance of converged neural machine translation models.
Muse: Parallel multi-scale attention for sequence to sequence learning
Guangxiang Zhao, Xu Sun, Jingjing Xu, Zhiyuan Zhang, and Liangchen Luo. 2019 · 1911
Earlier work this paper cites.
Ensemble methods in machine learning
Thomas G Dietterich. 2000 · 2000
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Pushing the limits of semi-supervised learning for automatic speech recognition
Yu Zhang, James Qin, Daniel S. Park, Wei Han, Chung-Cheng Chiu, Ruoming Pang, Quoc V. Le, and Yonghui Wu. 2020 · 2010
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Is neural machine translation ready for deployment? a case study on 30 translation directions
Marcin Junczys-Dowmunt, Tomasz Dwojak, and Hieu Hoang. 2016 · 2016
Earlier work this paper cites.
Checkpoint ensembles: Ensemble methods from a single training process
Hugh Chen, Scott Lundberg, and Su-In Lee. 2017 · 2017
Earlier work this paper cites.
Snapshot ensembles: Train 1, get M for free
Gao Huang, Yixuan Li, Geoff Pleiss, Zhuang Liu, John E. Hopcroft, and Kilian Q. Weinberger. 2017 · 2017
Earlier work this paper cites.
Cyclical learning rates for training neural networks
Leslie N Smith. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Speech-transformer: a no-recurrence sequence-to-sequence model for speech recognition
Linhao Dong, Shuang Xu, and Bo Xu. 2018 · 2018
Earlier work this paper cites.
Averaging weights leads to wider optima and better generalization
Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry P. Vetrov, and Andrew Gordon Wilson. 2018 · 2018
Earlier work this paper cites.
A comparable study on model averaging, ensembling and reranking in nmt
Yuchen Liu, Long Zhou, Yining Wang, Yang Zhao, Jiajun Zhang, and Chengqing Zong. 2018 · 2018
Cited alongside, same era.
Training tips for the transformer model
Martin Popel and Ondřej Bojar. 2018 · 2018
Cited alongside, same era.
Self-attention with relative position representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. 2018 · 2018
Cited alongside, same era.
Data augmentation for end-to-end speech translation: FBK@IWSLT ‘19
Mattia A. Di Gangi, Matteo Negri, Viet Nhat Nguyen, Amirhossein Tebbifakhr, and Marco Turchi. 2019 · 2019
Cited alongside, same era.
A comparative study on transformer vs rnn in speech applications
Shigeki Karita, Nanxin Chen, Tomoki Hayashi, Takaaki Hori, Hirofumi Inaguma, Ziyan Jiang, Masao Someki, Nelson Yalta, Ryuichi Yamamoto, Xiao fei Wang, Shinji Watanabe, Takenori Yoshimura, and Wangyou Zhang. 2019 · 2019
Cited alongside, same era.
Transformer-based acoustic modeling for hybrid speech recognition
Yongqiang Wang, Abdelrahman Mohamed, Duc Le, Chunxi Liu, Alex Xiao, Jay Mahadeokar, Hongzhao Huang, Andros Tjandra, Xiaohui Zhang, Frank Zhang, Christian Fuegen, Geoffrey Zweig, and Michael L. Seltzer. 2020 · 2020
Later among the works it cites.
Proceedings of the Sixth Conference on Machine Translation
Loic Barrault, Ondrej Bojar, Fethi Bougares, Rajen Chatterjee, Marta R. Costa-jussa, Christian Federmann, Mark Fishel, Alexander Fraser, Markus Freitag, Yvette Graham, Roman Grundkiewicz, Paco Guzman, Barry Haddow, Matthias Huck, Antonio Jimeno Yepes, Philipp Koehn, Tom Kocmi, Andre Martins, Makoto Morishita, and Christof Monz, editors. 2021 · 2021
Later among the works it cites.
Tune in: The afrl wmt21 news-translation systems
Grant Erdmann, Jeremy Gwinnup, and Tim Anderson. 2021 · 2021
Later among the works it cites.
A comparative study on neural architectures and training methods for japanese speech recognition
Shigeki Karita, Yotaro Kubo, Michiel Bacchiani, and Llion Jones. 2021 · 2021
Later among the works it cites.
Scalable and efficient moe training for multitask multilingual models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019 · 2019
Cited alongside, same era.
The rwth aachen university machine translation systems for wmt 2019
Jan Rosendahl, Christian Herold, Yunsu Kim, Miguel Graça, Weiyue Wang, Parnia Bahar, Yingbo Gao, and Hermann Ney. 2019 · 2019
Cited alongside, same era.
CUED@WMT19:EWC&LMs
Felix Stahlberg, Danielle Saunders, Adrià de Gispert, and Bill Byrne. 2019 · 2019
Cited alongside, same era.
DFKI-NMT submission to the WMT19 news translation task
Jingyi Zhang and Josef van Genabith. 2019 · 2019
Cited alongside, same era.
A survey on ensemble learning
Xibin Dong, Zhiwen Yu, Wenming Cao, Yifan Shi, and Qianli Ma. 2020 · 2020
Cited alongside, same era.
Mask ctc: Non-autoregressive end-to-end asr with ctc and mask predict
Yosuke Higuchi, Shinji Watanabe, Nanxin Chen, Tetsuji Ogawa, and Tetsunori Kobayashi. 2020 · 2020
Cited alongside, same era.
Spike-triggered non-autoregressive transformer for end-to-end speech recognition
Zhengkun Tian, Jiangyan Yi, Jianhua Tao, Ye Bai, Shuai Zhang, and Zhengqi Wen. 2020 · 2020
Cited alongside, same era.
Young Jin Kim, Ammar Ahmad Awan, Alexandre Muzio, Andres Felipe Cruz Salinas, Liyang Lu, Amr Hendy, Samyam Rajbhandari, Yuxiong He, and Hany Hassan Awadalla. 2021 · 2021
Later among the works it cites.
Miss@wmt21: Contrastive learning-reinforced domain adaptation in neural machine translation
Zuchao Li, Masao Utiyama, Eiichiro Sumita, and Hai Zhao. 2021 · 2021
Later among the works it cites.
Merging models with fisher-weighted averaging
Michael Matena and Colin Raffel. 2021 · 2021
Later among the works it cites.
Nvidia nemo’s neural machine translation systems for english-german and english-russian news and biomedical tasks at wmt21
Sandeep Subramanian, Oleksii Hrinchuk, Virginia Adams, and Oleksii Kuchaiev. 2021 · 2021
Later among the works it cites.
Facebook ai’s wmt21 news translation task submission
Chau Tran, Shruti Bhosale, James Cross, Philipp Koehn, Sergey Edunov, and Angela Fan. 2021 · 2021
Later among the works it cites.
Hw-tsc’s participation in the wmt 2021 news translation shared task
Daimeng Wei, Zongyao Li, Zhanglin Wu, Zhengzhe Yu, Xiaoyu Chen, Hengchao Shang, Jiaxin Guo, Minghan Wang, Lizhi Lei, Min Zhang, Hao Yang, and Ying Qin. 2021 · 2021
Later among the works it cites.
Findings of the IWSLT 2022 evaluation campaign
Antonios Anastasopoulos, Loïc Barrault, Luisa Bentivogli, Marcely Zanon Boito, Ondřej Bojar, Roldano Cattoni, Anna Currey, Georgiana Dinu, Kevin Duh, Maha Elbayad, Clara Emmanuel, Yannick Estève, Marcello Federico, Christian Federmann, Souhir Gahbiche, Hongyu Gong, Roman Grundkiewicz, Barry Haddow, Benjamin Hsu, Dávid Javorský, Vĕra Kloudová, Surafel Lakew, Xutai Ma, Prashant Mathur, Paul McNamee, Kenton Murray, Maria Nǎdejde, Satoshi Nakamura, Matteo Negri, Jan Niehues, Xing Niu, John Ortega, Juan Pino, Elizabeth Salesky, Jiatong Shi, Matthias Sperber, Sebastian Stüker, Katsuhito Sudoh, Marco Turchi, Yogesh Virkar, Alexander Waibel, Changhan Wang, and Shinji Watanabe. 2022 · 2022
Closest in time.
HW-TSC’s participation in the IWSLT 2022 isometric spoken language translation
Zongyao Li, Jiaxin Guo, Daimeng Wei, Hengchao Shang, Minghan Wang, Ting Zhu, Zhanglin Wu, Zhengzhe Yu, Xiaoyu Chen, Lizhi Lei, Hao Yang, and Ying Qin. 2022 · 2022
Closest in time.