Fetching the paper…
Reading the bibliography…
The high compute cost associated with pretraining large language models limits their research.
2001
Earlier work this paper cites.
Sennrich, R., Haddow, B., Birch, A.: Neural machine translation of rare words with subword units. In: Erk, K., Smith, N.A. (eds.) Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). pp. 1715–1725. Association for Computational Linguistics (2016). https://doi.org/10.18653/v1/P16-1162
2016
Earlier work this paper cites.
Cataneo Silveira, I., Deratani Mauá, D.: Advances in automatically solving the ENEM. In: 2018 7th Brazilian Conference on Intelligent Systems (BRACIS). pp. 43–48 (2018). https://doi.org/10.1109/BRACIS.2018.00016
2018
Earlier work this paper cites.
Shazeer, N., Stern, M.: Adafactor: adaptive learning rates with sublinear memory cost. In: Dy, J., Krause, A. (eds.) Proceedings of the 35th International Conference on Machine Learning. Proceedings of Machine Learning Research, vol. 80, pp. 4596–4604. PMLR (2018), https://proceedings.mlr.press/v80/shazeer18a.html
2018
Earlier work this paper cites.
Sayama, H.F., Araujo, A.V., Fernandes, E.R.: FaQuAD: reading comprehension dataset in the domain of Brazilian higher education. In: 2019 8th Brazilian Conference on Intelligent Systems (BRACIS). pp. 443–448 (2019). https://doi.org/10.1109/BRACIS.2019.00084
2019
Earlier work this paper cites.
Brown, T., et al.: Language models are few-shot learners. In: Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., Lin, H. (eds.) Advances in Neural Information Processing Systems. vol. 33, pp. 1877–1901. Curran Associates, Inc. (2020), https://proceedings.neurips.cc/paper_files/paper/2020/file/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf
2020
Earlier work this paper cites.
Carlini, N., et al.: Extracting training data from large language models. In: 30th USENIX Security Symposium (USENIX Security 21). pp. 2633–2650. USENIX Association (2021), https://www.usenix.org/conference/usenixsecurity21/presentation/carlini-extracting
2021
Earlier work this paper cites.
Hendrycks, D., et al.: Measuring massive multitask language understanding. In: International Conference on Learning Representations (2021), https://openreview.net/forum?id=d7KBjmI3GmQ
2021
Earlier work this paper cites.
Polo, F., et al.: LegalNLP - natural language processing methods for the Brazilian legal language. In: Anais do XVIII Encontro Nacional de Inteligência Artificial e Computacional. pp. 763–774. SBC (2021). https://doi.org/10.5753/eniac.2021.18301
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
Xue, L., et al.: mT5: a massively multilingual pre-trained text-to-text transformer. In: Toutanova, K., et al. (eds.) Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. pp. 483–498. Association for Computational Linguistics (2021). https://doi.org/10.18653/v1/2021.naacl-main.41
2021
Earlier work this paper cites.
2022
Earlier work this paper cites.
Wei, J., et al.: Emergent abilities of large language models. Transactions on Machine Learning Research (2022), https://openreview.net/forum?id=yzkSU5zdwD
2022
Cited alongside, same era.
Almeida, T.S., Laitz, T., Bonás, G.K., Nogueira, R.: BLUEX: a benchmark based on Brazilian leading universities entrance exams. In: Naldi, M.C., Bianchi, R.A.C. (eds.) Intelligent Systems. pp. 337–347. Springer Nature Switzerland (2023). https://doi.org/10.1007/978-3-031-45368-7_22
2023
Cited alongside, same era.
Chowdhery, A., et al.: PaLM: scaling language modeling with pathways. Journal of Machine Learning Research 24
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Roberts, A., et al.: Scaling up models and data with t5x and seqio. Journal of Machine Learning Research 24
2023
Later among the works it cites.
Sakiyama, K., Montanari, R., Malaquias Junior, R., Nogueira, R., Romero, R.A.F.: Exploring text decoding methods for Portuguese legal text generation. In: Naldi, M.C., Bianchi, R.A.C. (eds.) Intelligent Systems. pp. 63–77. Springer Nature Switzerland (2023). https://doi.org/10.1007/978-3-031-45368-7_5
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Pires, R., Abonizio, H., Almeida, T.S., Nogueira, R.: Sabiá: Portuguese large language models. In: Naldi, M.C., Bianchi, R.A.C. (eds.) Intelligent Systems. pp. 226–240. Springer Nature Switzerland (2023). https://doi.org/10.1007/978-3-031-45392-2_15
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Zheng, L., et al.: Judging LLM-as-a-judge with MT-Bench and chatbot arena. In: Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., Levine, S. (eds.) Advances in Neural Information Processing Systems. vol. 36, pp. 46595–46623. Curran Associates, Inc. (2023), https://proceedings.neurips.cc/paper_files/paper/2023/file/91f18a1287b398d378ef22505bf41832-Paper-Datasets_and_Benchmarks.pdf
2023
Later among the works it cites.
2024
Closest in time.
Colombo, P., et al.: SaulLM-54B & SaulLM-141B: scaling up domain adaptation for the legal domain. In: Globerson, A., et al. (eds.) Advances in Neural Information Processing Systems. vol. 37, pp. 129672–129695. Curran Associates, Inc. (2024), https://proceedings.neurips.cc/paper_files/paper/2024/file/ea3f85a33f9ba072058e3df233cf6cca-Paper-Conference.pdf
2024
Closest in time.
2024
Closest in time.
Garcia, E., et al.: RoBERTaLexPT: a legal RoBERTa model pretrained with deduplication for Portuguese. In: Proceedings of the 16th International Conference on Computational Processing of Portuguese - Vol. 1. pp. 374–383. Association for Computational Lingustics (2024), https://aclanthology.org/2024.propor-1.38/
2024
Closest in time.
2024
Closest in time.
Groeneveld, D., et al.: OLMo: accelerating the science of language models. In: Ku, L.W., Martins, A., Srikumar, V. (eds.) Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). pp. 15789–15809. Association for Computational Linguistics (2024). https://doi.org/10.18653/v1/2024.acl-long.841
2024
Closest in time.