Scaling laws for transfer
Original
Hernandez, D., Kaplan, J., Henighan, T., and McCandlish, S · 2021
Later among the works it cites.
{GS}hard: Scaling giant models with conditional computation and automatic sharding
Lepikhin, D., Lee, H., Xu, Y., Chen, D., Firat, O., Huang, Y., Krikun, M., Shazeer, N., and Chen, Z · 2021
Later among the works it cites.
Scaling language models: Methods, analysis & insights from training gopher
Original
Rae, J. W., Borgeaud, S., Cai, T., Millican, K., Hoffmann, J., Song, H. F., Aslanides, J., Henderson, S., Ring, R., Young, S., Rutherford, E., Hennigan, T., Menick, J., Cassirer, A., Powell, R., van den Driessche, G., Hendricks, L. A., Rauh, M., Huang, P., Glaese, A., Welbl, J., Dathathri, S., Huang, S., Uesato, J., Mellor, J., Higgins, I., Creswell, A., McAleese, N., Wu, A., Elsen, E., Jayakumar, S. M., Buchatskaya, E., Budden, D., Sutherland, E., Simonyan, K., Paganini, M., Sifre, L., Martens, L., Li, X. L., Kuncoro, A., Nematzadeh, A., Gribovskaya, E., Donato, D., Lazaridou, A., Mensch, A., Lespiau, J., Tsimpoukelli, M., Grigorev, N., Fritz, D., Sottiaux, T., Pajarskas, M., Pohlen, T., Gong, Z., Toyama, D., de Masson d’Autume, C., Li, Y., Terzi, T., Mikulik, V., Babuschkin, I., Clark, A., de Las Casas, D., Guy, A., Jones, C., Bradbury, J., Johnson, M., Hechtman, B. A., Weidinger, L., Gabriel, I., Isaac, W. S., Lockhart, E., Osindero, S., Rimell, L., Dyer, C., Vinyals, O., Ayoub, K., Stanway, J., Bennett, L., Hassabis, D., Kavukcuoglu, K., and Irving, G · 2021
Later among the works it cites.
Template-guided clarifying question generation for web search clarification
Wang, J. and Li, W · 2021
Later among the works it cites.
mT5: A massively multilingual pre-trained text-to-text transformer
Xue, L., Constant, N., Roberts, A., Kale, M., Al-Rfou, R., Siddhant, A., Barua, A., and Raffel, C · 2021
Later among the works it cites.
Palm: Scaling language modeling with pathways, 2022
Original
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., Schuh, P., Shi, K., Tsvyashchenko, S., Maynez, J., Rao, A., Barnes, P., Tay, Y., Shazeer, N., Prabhakaran, V., Reif, E., Du, N., Hutchinson, B., Pope, R., Bradbury, J., Austin, J., Isard, M., Gur-Ari, G., Yin, P., Duke, T., Levskaya, A., Ghemawat, S., Dev, S., Michalewski, H., Garcia, X., Misra, V., Robinson, K., Fedus, L., Zhou, D., Ippolito, D., Luan, D., Lim, H., Zoph, B., Spiridonov, A., Sepassi, R., Dohan, D., Agrawal, S., Omernick, M., Dai, A. M., Pillai, T. S., Pellat, M., Lewkowycz, A., Moreira, E., Child, R., Polozov, O., Lee, K., Zhou, Z., Wang, X., Saeta, B., Diaz, M., Firat, O., Catasta, M., Wei, J., Meier-Hellstern, K., Eck, D., Dean, J., Petrov, S., and Fiedel, N · 2022
Later among the works it cites.
Using natural language prompts for machine translation, 2022
Original
Garcia, X. and Firat, O · 2022
Later among the works it cites.
Scaling laws for neural machine translation
Ghorbani, B., Firat, O., Freitag, M., Bapna, A., Krikun, M., Garcia, X., Chelba, C., and Cherry, C · 2022
Later among the works it cites.
Assistance with large language models
Krasheninnikov, D., Krasheninnikov, E., and Krueger, D · 2022
Later among the works it cites.
Controlling translation formality using pre-trained multilingual language models
Rippeth, E., Agrawal, S., and Carpuat, M · 2022
Later among the works it cites.
Language models are multilingual chain-of-thought reasoners, 2022
Original
Shi, F., Suzgun, M., Freitag, M., Wang, X., Srivats, S., Vosoughi, S., Chung, H. W., Tay, Y., Ruder, S., Zhou, D., Das, D., and Wei, J · 2022
Later among the works it cites.
Prompting palm for translation: Assessing strategies and performance, 2022
Original
Vilar, D., Freitag, M., Cherry, C., Luo, J., Ratnakar, V., and Foster, G · 2022
Later among the works it cites.
Measuring and mitigating name biases in neural machine translation
Wang, J., Rubinstein, B., and Cohn, T · 2022
Later among the works it cites.
Ai chains: Transparent and controllable human-ai interaction by chaining large language model prompts
Wu, T., Terry, M., and Cai, C. J · 2022
Later among the works it cites.
Least-to-most prompting enables complex reasoning in large language models, 2022
Original
Zhou, D., Schärli, N., Hou, L., Wei, J., Scales, N., Wang, X., Schuurmans, D., Cui, C., Bousquet, O., Le, Q., and Chi, E · 2022
Later among the works it cites.
Prompting large language model for machine translation: A case study, 2023
Original
Zhang, B., Haddow, B., and Birch, A · 2023
Closest in time.