Fetching the paper…
Reading the bibliography…
We present an empirical study of scaling properties of encoder-decoder Transformer models used in neural machine translation (NMT).
Scaling and generalization in neural networks: a case study
Subutai Ahmad and Gerald Tesauro · 1988
Earlier work this paper cites.
Four types of learning curves
S. Amari, Naotake Fujita, and S. Shinomoto · 1992
Earlier work this paper cites.
BLEU: a Method for Automatic Evaluation of Machine Translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Translationese and Its Dialects
Moshe Koppel and Noam Ordan · 2011
Earlier work this paper cites.
Sequence-level knowledge distillation, 2016
Yoon Kim and Alexander M. Rush · 2016
Earlier work this paper cites.
Improving neural machine translation models with monolingual data
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al · 2016
Earlier work this paper cites.
High-dimensional dynamics of generalization error in neural networks, 2017
Madhu S. Advani and Andrew M. Saxe · 2017
Earlier work this paper cites.
Deep learning scaling is predictable, empirically
Joel Hestness, Sharan Narang, Newsha Ardalani, Gregory Diamos, Heewoo Jun, Hassan Kianinejad, Md Patwary, Mostofa Ali, Yang Yang, and Yanqi Zhou · 2017
Earlier work this paper cites.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
The best of both worlds: Combining recent advances in neural machine translation, 2018
Mia Xu Chen, Orhan Firat, Ankur Bapna, Melvin Johnson, Wolfgang Macherey, George Foster, Llion Jones, Niki Parmar, Mike Schuster, Zhifeng Chen, Yonghui Wu, and Macduff Hughes · 2018
Earlier work this paper cites.
Adafactor: Adaptive learning rates with sublinear memory cost
Noam Shazeer and Mitchell Stern · 2018
Earlier work this paper cites.
Massively multilingual neural machine translation in the wild: Findings and challenges
Naveen Arivazhagan, Ankur Bapna, Orhan Firat, Dmitry Lepikhin, Melvin Johnson, Maxim Krikun, Mia Xu Chen, Yuan Cao, George F. Foster, Colin Cherry, Wolfgang Macherey, Zhifeng Chen, and Yonghui Wu · 2019
Earlier work this paper cites.
Findings of the 2019 conference on machine translation (wmt19)
Loïc Barrault, Ondřej Bojar, Marta R Costa-Jussa, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Philipp Koehn, Shervin Malmasi, et al · 2019
Earlier work this paper cites.
APE at Scale and Its Implications on MT Evaluation Biases
Markus Freitag, Isaac Caswell, and Scott Roy · 2019
Cited alongside, same era.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Xu Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V Le, Yonghui Wu, et al · 2019
Cited alongside, same era.
A constructive prediction of the generalization error across scales
Jonathan S Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, and Nir Shavit · 2019
Cited alongside, same era.
Scaling laws of recovering bernoulli
Kyunghyun Cho · 2020
Cited alongside, same era.
BLEU might be guilty but references are not innocent
Markus Freitag, David Grangier, and Isaac Caswell · 2020
Cited alongside, same era.
Scaling description of generalization with number of parameters in deep learning
Mario Geiger, Arthur Jacot, Stefano Spigler, Franck Gabriel, Levent Sagun, Stéphane d’ Ascoli, Giulio Biroli, Clément Hongler, and Matthieu Wyart · 2020
Translationese as a language in “multilingual” nmt
Parker Riley, Isaac Caswell, Markus Freitag, and David Grangier · 2020
Later among the works it cites.
Bleurt: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur P Parikh · 2020
Later among the works it cites.
Thibault Sellam, Amy Pu, Hyung Won Chung, Sebastian Gehrmann, Qijun Tan, Markus Freitag, Dipanjan Das, and Ankur P Parikh · 2020
Later among the works it cites.
On layer normalization in the transformer architecture
Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tieyan Liu · 2020
Later among the works it cites.
Explaining neural scaling laws, 2021
Yasaman Bahri, Ethan Dyer, Jared Kaplan, Jaehoon Lee, and Utkarsh Sharma · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Linearized two-layers neural networks in high dimension, 2020
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2020
Cited alongside, same era.
Statistical power and translationese in machine translation evaluation
Yvette Graham, Barry Haddow, and Philipp Koehn · 2020
Cited alongside, same era.
Revisiting self-training for neural sequence generation, 2020
Junxian He, Jiatao Gu, Jiajun Shen, and Marc’Aurelio Ranzato · 2020
Cited alongside, same era.
Scaling laws for autoregressive generative modeling
Tom Henighan, Jared Kaplan, Mor Katz, Mark Chen, Christopher Hesse, Jacob Jackson, Heewoo Jun, Tom B Brown, Prafulla Dhariwal, Scott Gray, et al · 2020
Cited alongside, same era.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Cited alongside, same era.
Limits to depth efficiencies of self-attention
Yoav Levine, Noam Wies, Or Sharir, Hofit Bata, and Amnon Shashua · 2020
Cited alongside, same era.
Data and parameter scaling laws for neural machine translation
Mitchell A Gordon, Kevin Duh, and Jared Kaplan · 2021
Closest in time.
Scaling laws for transfer, 2021
Danny Hernandez, Jared Kaplan, Tom Henighan, and Sam McCandlish · 2021
Closest in time.
Marcus Hutter · 2021
Closest in time.
Deep encoder, shallow decoder: Reevaluating non-autoregressive machine translation, 2021
Jungo Kasai, Nikolaos Pappas, Hao Peng, James Cross, and Noah A. Smith · 2021
Closest in time.
{GS}hard: Scaling giant models with conditional computation and automatic sharding
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen · 2021
Closest in time.
Capturing the learning curves of generic features maps for realistic data sets with a teacher-student model, 2021
Bruno Loureiro, Cédric Gerbelot, Hugo Cui, Sebastian Goldt, Florent Krzakala, Marc Mézard, and Lenka Zdeborová · 2021
Closest in time.
Which transformer architecture fits my data? a vocabulary bottleneck in self-attention, 2021
Noam Wies, Yoav Levine, Daniel Jannai, and Amnon Shashua · 2021
Closest in time.
Understanding knowledge distillation in non-autoregressive machine translation, 2021
Chunting Zhou, Graham Neubig, and Jiatao Gu · 2021
Closest in time.