A neural scaling law from the dimension of the data manifold, 2020
Sharma, U. and Kaplan, J · 2020
Later among the works it cites.
Leveraging monolingual data with self-supervision for multilingual neural machine translation, 2020
Siddhant, A., Bapna, A., Cao, Y., Firat, O., Chen, M., Kudugunta, S., Arivazhagan, N., and Wu, Y · 2020
Later among the works it cites.
On layer normalization in the transformer architecture, 2020
Xiong, R., Yang, Y., He, D., Zheng, K., Zheng, S., Xing, C., Zhang, H., Lan, Y., Wang, L., and Liu, T.-Y · 2020
Later among the works it cites.
Recipes for safety in open-domain chatbots
Original
Xu, J., Ju, D., Li, M., Boureau, Y.-L., Weston, J., and Dinan, E · 2020
Later among the works it cites.
Scaling laws vs model architectures: How does inductive bias influence scaling? an extensive empirical study on language tasks, 2021
Anonymous · 2021
Later among the works it cites.
Explaining neural scaling laws
Original
Bahri, Y., Dyer, E., Kaplan, J., Lee, J., and Sharma, U · 2021
Later among the works it cites.
Experts, errors, and context: A large-scale study of human evaluation for machine translation, 2021
Freitag, M., Foster, G., Grangier, D., Ratnakar, V., Tan, Q., and Macherey, W · 2021
Later among the works it cites.
An empirical exploration in quality filtering of text data, 2021
Gao, L · 2021
Later among the works it cites.
Scaling laws for neural machine translation, 2021
Ghorbani, B., Firat, O., Freitag, M., Bapna, A., Krikun, M., Garcia, X., Chelba, C., and Cherry, C · 2021
Later among the works it cites.
Data and parameter scaling laws for neural machine translation
Gordon, M., Duh, K., and Kaplan, J · 2021
Later among the works it cites.
Learning curves for analysis of deep networks
Hoiem, D., Gupta, T., Li, Z., and Shlapentokh-Rothman, M · 2021
Later among the works it cites.
To ship or not to ship: An extensive evaluation of automatic metrics for machine translation, 2021
Kocmi, T., Federmann, C., Grundkiewicz, R., Junczys-Dowmunt, M., Matsushita, H., and Menezes, A · 2021
Later among the works it cites.
Accuracy on the line: on the strong correlation between out-of-distribution and in-distribution generalization
Miller, J. P., Taori, R., Raghunathan, A., Sagawa, S., Koh, P. W., Shankar, V., Liang, P., Carmon, Y., and Schmidt, L · 2021
Later among the works it cites.
Language models are good translators, 2021
Wang, S., Tu, Z., Tan, Z., Wang, W., Sun, M., and Liu, Y · 2021
Later among the works it cites.
Lamda: Language models for dialog applications, 2022
Thoppilan, R., Freitas, D. D., Hall, J., Shazeer, N., Kulshreshtha, A., Cheng, H.-T., Jin, A., Bos, T., Baker, L., Du, Y., Li, Y., Lee, H., Zheng, H. S., Ghafouri, A., Menegali, M., Huang, Y., Krikun, M., Lepikhin, D., Qin, J., Chen, D., Xu, Y., Chen, Z., Roberts, A., Bosma, M., Zhou, Y., Chang, C.-C., Krivokon, I., Rusch, W., Pickett, M., Meier-Hellstern, K., Morris, M. R., Doshi, T., Santos, R. D., Duke, T., Soraker, J., Zevenbergen, B., Prabhakaran, V., Diaz, M., Hutchinson, B., Olson, K., Molina, A., Hoffman-John, E., Lee, J., Aroyo, L., Rajakumar, R., Butryna, A., Lamm, M., Kuzmina, V., Fenton, J., Cohen, A., Bernstein, R., Kurzweil, R., Aguera-Arcas, B., Cui, C., Croak, M., Chi, E., and Le, Q · 2022
Closest in time.
Intelligent selection of language model training data
Moore, R. C. and Lewis, W · 2041
Closest in time.