Fetching the paper…
Reading the bibliography…
In this work, we provide a large-scale empirical study of the scaling properties of multilingual neural machine translation models.
Massively multilingual neural machine translation in the wild: Findings and challenges
Arivazhagan, N., Bapna, A., Firat, O., Lepikhin, D., Johnson, M., Krikun, M., Chen, M. X., Cao, Y., Foster, G. F., Cherry, C., Macherey, W., Chen, Z., and Wu, Y · 1907
Earlier work this paper cites.
Massively multilingual neural machine translation in the wild: Findings and challenges
Arivazhagan, N., Bapna, A., Firat, O., Lepikhin, D., Johnson, M., Krikun, M., Chen, M. X., Cao, Y., Foster, G. F., Cherry, C., Macherey, W., Chen, Z., and Wu, Y · 1907
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2001
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J · 2002
Earlier work this paper cites.
Convex Optimization
Boyd, S. and Vandenberghe, L · 2004
Earlier work this paper cites.
Sequence transduction with recurrent neural networks, 2012
Graves, A · 2012
Earlier work this paper cites.
Multi-task learning for multiple language translation
Dong, D., Wu, H., He, W., Yu, D., and Wang, H · 2015
Earlier work this paper cites.
Multi-task sequence to sequence learning, 2015
Luong, M.-T., Le, Q. V., Sutskever, I., Vinyals, O., and Kaiser, L · 2015
Earlier work this paper cites.
chrF: character n-gram F-score for automatic MT evaluation
Popović, M · 2015
Earlier work this paper cites.
Enabling multi-source neural machine translation by concatenating source sentences in multiple languages
Dabre, R., Cromieres, F., and Kurohashi, S · 2017
Earlier work this paper cites.
Deep learning scaling is predictable, empirically
Hestness, J., Narang, S., Ardalani, N., Diamos, G., Jun, H., Kianinejad, H., Patwary, M., Ali, M., Yang, Y., and Zhou, Y · 2017
Earlier work this paper cites.
Six challenges for neural machine translation
Koehn, P. and Knowles, R · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Subword regularization: Improving neural network translation models with multiple subword candidates
Kudo, T · 2018
Earlier work this paper cites.
Adafactor: Adaptive learning rates with sublinear memory cost
Shazeer, N. M. and Stern, M · 2018
Cited alongside, same era.
Denoising neural machine translation training with trusted data and online data selection
Wang, W., Watanabe, T., Hughes, M., Nakagawa, T., and Chelba, C · 2018
Cited alongside, same era.
Findings of the 2019 conference on machine translation (WMT19)
Barrault, L., Bojar, O., Costa-jussà, M. R., Federmann, C., Fishel, M., Graham, Y., Haddow, B., Huck, M., Koehn, P., Malmasi, S., Monz, C., Müller, M., Pal, S., Post, M., and Zampieri, M · 2019
Cited alongside, same era.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Huang, Y., Cheng, Y., Bapna, A., Firat, O., Chen, D., Chen, M., Lee, H., Ngiam, J., Le, Q. V., Wu, Y., et al · 2019
Cited alongside, same era.
A constructive prediction of the generalization error across scales
Rosenfeld, J. S., Rosenfeld, A., Belinkov, Y., and Shavit, N · 2019
Cited alongside, same era.
Gradient vaccine: Investigating and improving multi-task optimization in massively multilingual models
Wang, Z., Tsvetkov, Y., Firat, O., and Cao, Y · 2021
Later among the works it cites.
Data scaling laws in nmt: The effect of noise and architecture
Bansal, Y., Ghorbani, B., Garg, A., Zhang, B., Krikun, M., Cherry, C., Neyshabur, B., and Firat, O · 2022
Later among the works it cites.
Building machine translation systems for the next thousand languages, 2022
Bapna, A., Caswell, I., Kreutzer, J., Firat, O., van Esch, D., Siddhant, A., Niu, M., Baljekar, P., Garcia, X., Macherey, W., Breiner, T., Axelrod, V., Riesa, J., Cao, Y., Chen, M. X., Macherey, K., Krikun, M., Wang, P., Gutkin, A., Shah, A., Huang, Y., Chen, Z., Wu, Y., and Hughes, M · 2022
Later among the works it cites.
Palm: Scaling language modeling with pathways, 2022
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., Schuh, P., Shi, K., Tsvyashchenko, S., Maynez, J., Rao, A., Barnes, P., Tay, Y., Shazeer, N., Prabhakaran, V., Reif, E., Du, N., Hutchinson, B., Pope, R., Bradbury, J., Austin, J., Isard, M., Gur-Ari, G., Yin, P., Duke, T., Levskaya, A., Ghemawat, S., Dev, S., Michalewski, H., Garcia, X., Misra, V., Robinson, K., Fedus, L., Zhou, D., Ippolito, D., Luan, D., Lim, H., Zoph, B., Spiridonov, A., Sepassi, R., Dohan, D., Agrawal, S., Omernick, M., Dai, A. M., Pillai, T. S., Pellat, M., Lewkowycz, A., Moreira, E., Child, R., Polozov, O., Lee, K., Zhou, Z., Wang, X., Saeta, B., Diaz, M., Firat, O., Catasta, M., Wei, J., Meier-Hellstern, K., Eck, D., Dean, J., Petrov, S., and Fiedel, N · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Cited alongside, same era.
Scaling laws for autoregressive generative modeling
Henighan, T., Kaplan, J., Katz, M., Chen, M., Hesse, C., Jackson, J., Jun, H., Brown, T. B., Dhariwal, P., Gray, S., et al · 2020
Cited alongside, same era.
Gshard: Scaling giant models with conditional computation and automatic sharding
Lepikhin, D., Lee, H., Xu, Y., Chen, D., Firat, O., Huang, Y., Krikun, M., Shazeer, N., and Chen, Z · 2020
Cited alongside, same era.
BLEURT: Learning robust metrics for text generation
Sellam, T., Das, D., and Parikh, A · 2020
Cited alongside, same era.
Balancing training for multilingual neural machine translation
Wang, X., Tsvetkov, Y., and Neubig, G · 2020
Cited alongside, same era.
On layer normalization in the transformer architecture
Xiong, R., Yang, Y., He, D., Zheng, K., Zheng, S., Xing, C., Zhang, H., Lan, Y., Wang, L., and Liu, T · 2020
Cited alongside, same era.
Data and parameter scaling laws for neural machine translation
Gordon, M. A., Duh, K., and Kaplan, J · 2021
Cited alongside, same era.
Later among the works it cites.
No language left behind: Scaling human-centered machine translation
Costa-jussà, M. R., Cross, J., Çelebi, O., Elbayad, M., Heafield, K., Heffernan, K., Kalbassi, E., Lam, J., Licht, D., Maillard, J., et al · 2022
Later among the works it cites.
Scaling laws for neural machine translation
Ghorbani, B., Firat, O., Freitag, M., Bapna, A., Krikun, M., García, X., Chelba, C., and Cherry, C · 2022
Later among the works it cites.
Training compute-optimal large language models, 2022
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., Driessche, G. v. d., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., Rae, J. W., Vinyals, O., and Sifre, L · 2022
Later among the works it cites.
In defense of the unitary scalarization for deep multi-task learning
Kurin, V., De Palma, A., Kostrikov, I., Whiteson, S., and Kumar, M. P · 2022
Later among the works it cites.
Reed, S., Zolna, K., Parisotto, E., Colmenarejo, S. G., Novikov, A., Barth-Maron, G., Gimenez, M., Sulsky, Y., Kay, J., Springenberg, J. T., et al · 2022
Later among the works it cites.
Multitask prompted training enables zero-shot task generalization
Sanh, V., Webson, A., Raffel, C., Bach, S., Sutawika, L., Alyafeai, Z., Chaffin, A., Stiegler, A., Raja, A., Dey, M., Bari, M. S., Xu, C., Thakker, U., Sharma, S. S., Szczechla, E., Kim, T., Chhablani, G., Nayak, N., Datta, D., Chang, J., Jiang, M. T.-J., Wang, H., Manica, M., Shen, S., Yong, Z. X., Pandey, H., Bawden, R., Wang, T., Neeraj, T., Rozen, J., Sharma, A., Santilli, A., Fevry, T., Fries, J. A., Teehan, R., Scao, T. L., Biderman, S., Gao, L., Wolf, T., and Rush, A. M · 2022
Later among the works it cites.
Do current multi-task optimization methods in deep learning even help?
Xin, D., Ghorbani, B., Garg, A., Firat, O., and Gilmer, J · 2022
Later among the works it cites.
Scaling vision transformers
Zhai, X., Kolesnikov, A., Houlsby, N., and Beyer, L · 2022
Later among the works it cites.
Examining scaling and transfer of language model architectures for machine translation
Zhang, B., Ghorbani, B., Bapna, A., Cheng, Y., Garcia, X., Shen, J., and Firat, O · 2022
Later among the works it cites.