Fetching the paper…
Reading the bibliography…
Small Language Models (SLMs) offer efficient alternatives to LLMs for specific domains.
Bertscore: Evaluating text generation with bert, 2020
Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., and Artzi, Y · 1904
Earlier work this paper cites.
Blimp: The benchmark of linguistic minimal pairs for english, 2023b
Warstadt, A., Parrish, A., Liu, H., Mohananey, A., Peng, W., Wang, S.-F., and Bowman, S. R · 1912
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J · 2002
Earlier work this paper cites.
Toward a Psychology of Language Acquisition , pp. 323–328
Tomasello, M · 2003
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Lin, C.-Y · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Banerjee, S. and Lavie, A · 2005
Earlier work this paper cites.
Language models are few-shot learners, 2020
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2005
Earlier work this paper cites.
Comparative study of indexing and search strategies for the hindi, marathi, and bengali languages
Dolamic, L. and Savoy, J · 2010
Earlier work this paper cites.
Mtld, vocd-d, and hd-d: A validation study of sophisticated approaches to lexical diversity assessment
McCarthy, P. M. and Jarvis, S · 2010
Earlier work this paper cites.
Mapping the early language environment using all-day recordings and automated analysis
Gilkerson, J., Richards, J. A., Warren, S. F., Montgomery, J. K., Greenwood, C. R., Oller, D. K., Hansen, J. H. L., and Paul, T. D · 2017
Earlier work this paper cites.
The effect of translationese in machine translation test sets
Zhang, M. and Toral, A · 2019
Earlier work this paper cites.
Multilingual denoising pre-training for neural machine translation
Liu, Y., Gu, J., Goyal, N., Li, X., Edunov, S., Ghazvininejad, M., Lewis, M., and Zettlemoyer, L · 2020
Earlier work this paper cites.
mt5: A massively multilingual pre-trained text-to-text transformer
Xue, L., Constant, N., Roberts, A., Kale, M., Al-Rfou, R., Siddhant, A., Barua, A., and Raffel, C · 2020
Earlier work this paper cites.
Gpt-neo: Large scale autoregressive language modeling with mesh-tensorflow
Black, S., Gao, L., Wang, P., Leahy, C., and Biderman, S · 2021
Earlier work this paper cites.
Indicbart: A pre-trained model for indic natural language generation
Dabre, R., Shrotriya, H., Kunchukuttan, A., Puduppully, R., Khapra, M. M., and Kumar, P · 2021
Cited alongside, same era.
Beyond english-centric multilingual machine translation
Fan, A., Bhosale, S., Schwenk, H., Ma, Z., El-Kishky, A., Goyal, S., Baines, M., Celebi, O., Wenzek, G., Chaudhary, V., et al · 2021
Cited alongside, same era.
Subword-based cross-lingual transfer of embeddings from Hindi to Marathi and Nepali
Bafna, N. and Žabokrtský, Z · 2022
Cited alongside, same era.
Palm: Scaling language modeling with pathways, 2022
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., Schuh, P., Shi, K., Tsvyashchenko, S., Maynez, J., Rao, A., Barnes, P., Tay, Y., Shazeer, N., Prabhakaran, V., Reif, E., Du, N., Hutchinson, B., Pope, R., Bradbury, J., Austin, J., Isard, M., Gur-Ari, G., Yin, P., Duke, T., Levskaya, A., Ghemawat, S., Dev, S., Michalewski, H., Garcia, X., Misra, V., Robinson, K., Fedus, L., Zhou, D., Ippolito, D., Luan, D., Lim, H., Zoph, B., Spiridonov, A., Sepassi, R., Dohan, D., Agrawal, S., Omernick, M., Dai, A. M., Pillai, T. S., Pellat, M., Lewkowycz, A., Moreira, E., Child, R., Polozov, O., Lee, K., Zhou, Z., Wang, X., Saeta, B., Diaz, M., Firat, O., Catasta, M., Wei, J., Meier-Hellstern, K., Eck, D., Dean, J., Petrov, S., and Fiedel, N · 2022
Trained on 100 million words and still in shape: BERT meets British National Corpus
Samuel, D., Kutuzov, A., Øvrelid, L., and Velldal, E · 2023
Later among the works it cites.
Democratizing neural machine translation with OPUS-MT
Tiedemann, J., Aulamo, M., Bakshandaeva, D., Boggia, M., Grönroos, S.-A., Nieminen, T., Raganato A., Scherrer, Y., Vazquez, R., and Virpioja, S · 2023
Later among the works it cites.
Tokenization and the noiseless channel
Zouhar, V., Meister, C., Gastaldi, J., Du, L., Sachan, M., and Cotterell, R · 2023
Later among the works it cites.
Why do language models perform worse for morphologically complex languages?, 2024
Arnett, C. and Bergen, B. K · 2024
Later among the works it cites.
Sutra: Scalable multilingual language model architecture, 2024
Bendale, A., Sapienza, M., Ripplinger, S., Gibbs, S., Lee, J., and Mistry, P · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Training compute-optimal large language models, 2022
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., de Las Casas, D., Hendricks, L. A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., van den Driessche, G., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., Rae, J. W., Vinyals, O., and Sifre, L · 2022
Cited alongside, same era.
nanogpt, 2022
Karpathy, A · 2022
Cited alongside, same era.
Guiding neural story generation with reader models, 2022
Peng, X., Xie, K., Alabdulkarim, A., Kayam, H., Dani, S., and Riedl, M. O · 2022
Cited alongside, same era.
No language left behind: Scaling human-centered machine translation, 2022
Team, N., Costa-jussà, M. R., Cross, J., Çelebi, O., Elbayad, M., Heafield, K., Heffernan, K., Kalbassi, E., Lam, J., Licht, D., Maillard, J., Sun, A., Wang, S., Wenzek, G., Youngblood, A., Akula, B., Barrault, L., Gonzalez, G. M., Hansanti, P., Hoffman, J., Jarrett, S., Sadagopan, K. R., Rowe, D., Spruit, S., Tran, C., Andrews, P., Ayan, N. F., Bhosale, S., Edunov, S., Fan, A., Gao, C., Goswami, V., Guzmán, F., Koehn, P., Mourachko, A., Ropers, C., Saleem, S., Schwenk, H., and Wang, J · 2022
Cited alongside, same era.
Gala, J., Chitale, P. A., AK, R., Gumma, V., Doddapaneni, S., Kumar, A., Nawale, J., Sujatha, A., Puduppully, R., Raghavan, V., et al · 2023
Cited alongside, same era.
Bridging the resource gap: Exploring the efficacy of English and multilingual LLMs for Swedish
Holmström, O., Kunz, J., and Kuhlmann, M · 2023
Cited alongside, same era.
Is chatgpt a good translator? yes with gpt-4 as the engine, 2023
Jiao, W., Wang, W., tse Huang, J., Wang, X., Shi, S., and Tu, Z · 2023
Cited alongside, same era.
Large language models are state-of-the-art evaluators of translation quality, 2023
Kocmi, T. and Federmann, C · 2023
Cited alongside, same era.
Boughorbel, S., Parvez, M. R., and Hawasly, M · 2024
Later among the works it cites.
Humans or LLMs as the judge? a study on judgement bias
Chen, G. H., Chen, S., Liu, Z., Jiang, F., and Wang, B · 2024
Later among the works it cites.
Pretraining language models using translationese
Doshi, M., Dabre, R., and Bhattacharyya, P · 2024
Later among the works it cites.
Emergent abilities in reduced-scale generative language models, 2024
Muckatira, S., Deshpande, V., Lialin, V., and Rumshisky, A · 2024
Later among the works it cites.
Tiktoken, 2024
OpenAI · 2024
Later among the works it cites.
Sarvam 1 : The first indian language llm, 2024
Sarvam · 2024
Later among the works it cites.
Berttime stories: Investigating the role of synthetic story data in language pre-training, 2024
Theodoropoulos, N., Filandrianos, G., Lyberatos, V., Lymperaiou, M., and Stamou, G · 2024
Later among the works it cites.
Wang, F., Zhang, Z., Zhang, X., Wu, Z., Mo, T., Lu, Q., Wang, W., Li, R., Xu, J., Tang, X., He, Q., Ma, Y., Huang, M., and Wang, S · 2024
Later among the works it cites.
Predicting trends in the quality of state-of-the-art neural networks without access to training or testing data
Martin, C. H., Peng, T. S., and Mahoney, M. W · 2041
Closest in time.