Fetching the paper…
Reading the bibliography…
Moderate-sized large language models (LLMs) -- those with 7B or 13B parameters -- exhibit promising machine translation (MT) performance.
Bleu: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J · 2002
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Oord, A. v. d., Li, Y., and Vinyals, O · 2018
Earlier work this paper cites.
A call for clarity in reporting BLEU scores
Post, M · 2018
Earlier work this paper cites.
Fine-tuning language models from human preferences
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
A simple framework for contrastive learning of visual representations
Chen, T., Kornblith, S., Norouzi, M., and Hinton, G · 2020
Earlier work this paper cites.
Momentum contrast for unsupervised visual representation learning
He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R · 2020
Earlier work this paper cites.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
Rasley, J., Rajbhandari, S., Ruwase, O., and He, Y · 2020
Earlier work this paper cites.
BLEURT: Learning robust metrics for text generation
Sellam, T., Das, D., and Parikh, A · 2020
Earlier work this paper cites.
Beyond english-centric multilingual machine translation
Fan, A., Bhosale, S., Schwenk, H., Ma, Z., El-Kishky, A., Goyal, S., Baines, M., Celebi, O., Wenzek, G., Chaudhary, V., et al · 2021
Earlier work this paper cites.
Contrastive learning with hard negative samples
Robinson, J. D., Chuang, C.-Y., Sra, S., and Jegelka, S · 2021
Earlier work this paper cites.
BERT, mBERT, or BiBERT? a study on contextualized embeddings for neural machine translation
Xu, H., Van Durme, B., and Murray, K · 2021
Earlier work this paper cites.
mT5: A massively multilingual pre-trained text-to-text transformer
Xue, L., Constant, N., Roberts, A., Kale, M., Al-Rfou, R., Siddhant, A., Barua, A., and Raffel, C · 2021
Cited alongside, same era.
LoRA: Low-rank adaptation of large language models
Hu, E. J., yelong shen, Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2022
Cited alongside, same era.
Findings of the 2022 conference on machine translation (WMT22)
Kocmi, T., Bawden, R., Bojar, O., Dvorkovich, A., Federmann, C., Fishel, M., Gowda, T., Graham, Y., Grundkiewicz, R., Haddow, B., Knowles, R., Koehn, P., Monz, C., Morishita, M., Nagata, M., Nakazawa, T., Novák, M., Popel, M., and Popović, M · 2022
Cited alongside, same era.
No language left behind: Scaling human-centered machine translation
NLLB TEAM, Costa-jussà, M. R., Cross, J., Çelebi, O., Elbayad, M., Heafield, K., Heffernan, K., Kalbassi, E., Lam, J., Licht, D., Maillard, J., et al · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Findings of the 2023 conference on machine translation (WMT23): LLMs are here but not quite there yet
Kocmi, T., Avramidis, E., Bawden, R., Bojar, O., Dvorkovich, A., Federmann, C., Fishel, M., Freitag, M., Gowda, T., Grundkiewicz, R., Haddow, B., Koehn, P., Marie, B., Monz, C., Morishita, M., Murray, K., Nagata, M., Nakazawa, T., Popel, M., Popović, M., and Shmatova, M · 2023
Later among the works it cites.
Madlad-400: A multilingual and document-level large audited dataset, 2023
Kudugunta, S., Caswell, I., Zhang, B., Garcia, X., Choquette-Choo, C. A., Lee, K., Xin, D., Kusupati, A., Stella, R., Bapna, A., and Firat, O · 2023
Later among the works it cites.
Li, J., Zhou, H., Huang, S., Chen, S., and Chen, J · 2023
Later among the works it cites.
Small data, big impact: Leveraging minimal data for effective machine translation
Maillard, J., Gao, C., Kalbassi, E., Sadagopan, K. R., Goswami, V., Koehn, P., Fan, A., and Guzman, F · 2023
Later among the works it cites.
Gpt-4 technical report, 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
COMET-22: Unbabel-IST 2022 submission for the metrics shared task
Rei, R., C. de Souza, J. G., Alves, D., Zerva, C., Farinha, A. C., Glushkova, T., Lavie, A., Coheur, L., and Martins, A. F. T · 2022
Cited alongside, same era.
Falcon-40B: an open large language model with state-of-the-art performance
Almazrouei, E., Alobeidli, H., Alshamsi, A., Cappelli, A., Cojocaru, R., Debbah, M., Goffinet, E., Heslow, D., Launay, J., Malartic, Q., Noune, B., Pannier, B., and Penedo, G · 2023
Cited alongside, same era.
Improving translation faithfulness of large language models via augmenting instructions
Chen, Y., Liu, Y., Meng, F., Chen, Y., Xu, J., and Zhou, J · 2023
Cited alongside, same era.
Results of WMT23 metrics shared task: Metrics might be guilty but references are not innocent
Freitag, M., Mathur, N., Lo, C.-k., Avramidis, E., Rei, R., Thompson, B., Kocmi, T., Blain, F., Deutsch, D., Stewart, C., Zerva, C., Castilho, S., Lavie, A., and Foster, G · 2023
Cited alongside, same era.
xcomet: Transparent machine translation evaluation through fine-grained error detection
Guerreiro, N. M., Rei, R., van Stigt, D., Coheur, L., Colombo, P., and Martins, A. F · 2023
Cited alongside, same era.
Contrastive prefence learning: Learning from human feedback without rl
Hejna, J., Rafailov, R., Sikchi, H., Finn, C., Niekum, S., Knox, W. B., and Sadigh, D · 2023
Cited alongside, same era.
How good are gpt models at machine translation? a comprehensive evaluation
Hendy, A., Abdelrehim, M., Sharaf, A., Raunak, V., Gabr, M., Matsushita, H., Kim, Y. J., Afify, M., and Awadalla, H. H · 2023
Cited alongside, same era.
OpenAI · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C · 2023
Later among the works it cites.
Scaling up cometkiwi: Unbabel-ist 2023 submission for the quality estimation shared task
Rei, R., Guerreiro, N. M., Pombal, J., van Stigt, D., Treviso, M., Coheur, L., de Souza, J. G., and Martins, A. F · 2023
Later among the works it cites.
Multilingual representation distillation with contrastive learning
Tan, W., Heffernan, K., Schwenk, H., and Koehn, P · 2023
Later among the works it cites.
Exploring prompt engineering with GPT language models for document-level machine translation: Insights and findings
Wu, Y. and Hu, G · 2023
Later among the works it cites.
A paradigm shift in machine translation: Boosting translation performance of large language models, 2023
Xu, H., Kim, Y. J., Sharaf, A., and Awadalla, H. H · 2023
Later among the works it cites.
Yang, W., Li, C., Zhang, J., and Zong, C · 2023
Later among the works it cites.
Tim: Teaching large language models to translate with comparison
Zeng, J., Meng, F., Yin, Y., and Zhou, J · 2023
Later among the works it cites.
Zhang, S., Fang, Q., Zhang, Z., Ma, Z., Zhou, Y., Huang, L., Bu, M., Gui, S., Chen, Y., Chen, X., et al · 2023
Later among the works it cites.
Navigating the metrics maze: Reconciling score magnitudes and accuracies
Kocmi, T., Zouhar, V., Federmann, C., and Post, M · 2024
Closest in time.