Fetching the paper…
Reading the bibliography…
Automatic evaluation of sequence generation, traditionally reliant on metrics like BLEU and ROUGE, often fails to capture the semantic accuracy of generated text sequences due to their emphasis on n-gram overlap.
R. J. Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” Reinforcement learning , 1992
1992
Earlier work this paper cites.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: a method for automatic evaluation of machine translation,” in Proc. of ACL , 2002, pp. 311–318
2002
Earlier work this paper cites.
I. J. Myung, “Tutorial on maximum likelihood estimation,” Journal of mathematical Psychology , 2003
2003
Earlier work this paper cites.
C.-Y. Lin, “ROUGE: A package for automatic evaluation of summaries,” in Text Summarization Branches Out , 2004, pp. 74–81
2004
Earlier work this paper cites.
G. Neubig, “Travatar: A forest-to-string machine translation engine based on tree transducers,” in Proc. of ACL , 2013, pp. 91–96
2013
Earlier work this paper cites.
M. Shen, D. Kawahara, and S. Kurohashi, “Dependency parse reranking with rich subtree features,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2014
2014
Earlier work this paper cites.
K. M. Hermann, T. Kociský, E. Grefenstette, L. Espeholt, W. Kay, M. Suleyman, and P. Blunsom, “Teaching machines to read and comprehend,” in Proc. of NeurIPS , 2015, pp. 1693–1701
2015
Earlier work this paper cites.
Y. Kim and A. M. Rush, “Sequence-level knowledge distillation,” in Proc. of EMNLP , 2016, pp. 1317–1327
2016
Earlier work this paper cites.
S. Shen, Y. Cheng, Z. He, W. He, H. Wu, M. Sun, and Y. Liu, “Minimum risk training for neural machine translation,” in Proc. of ACL , 2016, pp. 1683–1692
2016
Earlier work this paper cites.
J. Zhang, X. Wu, and V. S. Sheng, “Learning from crowdsourced labeled data: a survey,” Artificial Intelligence Review , 2016
2016
Earlier work this paper cites.
P. F. Christiano, J. Leike, T. B. Brown, M. Martic, S. Legg, and D. Amodei, “Deep reinforcement learning from human preferences,” in Proc. of NeurIPS , 2017, pp. 4299–4307
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proc. of NeurIPS , 2017, pp. 5998–6008
2017
Earlier work this paper cites.
M. Freitag and Y. Al-Onaizan, “Beam search strategies for neural machine translation,” in Proceedings of the First Workshop on Neural Machine Translation , 2017, pp. 56–60
2017
Earlier work this paper cites.
H. Shimanaka, T. Kajiwara, and M. Komachi, “RUSE: Regressor using sentence embeddings for automatic machine translation evaluation,” in Proceedings of the Third Conference on Machine Translation: Shared Task Papers , 2018, pp. 751–758
2018
Earlier work this paper cites.
S. Rao and J. Tetreault, “Dear sir or madam, may I introduce the GYAFC dataset: Corpus, benchmarks and metrics for formality style transfer,” in Proc. of NAACL , 2018, pp. 129–140
2018
Earlier work this paper cites.
J. Wieting, T. Berg-Kirkpatrick, K. Gimpel, and G. Neubig, “Beyond BLEU: Training neural machine translation with semantic similarity,” in Proc. of ACL , 2019, pp. 4344–4355
2019
Earlier work this paper cites.
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov, “Roberta: A robustly optimized bert pretraining approach,” ArXiv preprint , 2019
2019
Earlier work this paper cites.
W. Zhao, M. Peyrard, F. Liu, Y. Gao, C. M. Meyer, and S. Eger, “MoverScore: Text generation evaluating with contextualized embeddings and earth mover distance,” in Proc. of EMNLP , 2019, pp. 563–578
2019
Earlier work this paper cites.
T. Zhang, V. Kishore, F. Wu, K. Q. Weinberger, and Y. Artzi, “Bertscore: Evaluating text generation with BERT,” in Proc. of ICLR , 2020
2020
Earlier work this paper cites.
T. Sellam, D. Das, and A. Parikh, “BLEURT: Learning robust metrics for text generation,” in Proc. of ACL , 2020, pp. 7881–7892
2020
Earlier work this paper cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” J. Mach. Learn. Res. , pp. 140:1–140:67, 2020
2020
Earlier work this paper cites.
B. Thompson and M. Post, “Automatic machine translation evaluation in many languages via zero-shot paraphrasing,” in Proc. of EMNLP , 2020, pp. 90–121
2020
Earlier work this paper cites.
M. Fomicheva, S. Sun, L. Yankovskaya, F. Blain, F. Guzmán, M. Fishel, N. Aletras, V. Chaudhary, and L. Specia, “Unsupervised quality estimation for neural machine translation,” TACL , pp. 539–555, 2020
2020
Earlier work this paper cites.
R. Rei, C. Stewart, A. C. Farinha, and A. Lavie, “COMET: A neural framework for MT evaluation,” in Proc. of EMNLP , 2020, pp. 2685–2702
2020
Earlier work this paper cites.
A. Conneau, K. Khandelwal, N. Goyal, V. Chaudhary, G. Wenzek, F. Guzmán, E. Grave, M. Ott, L. Zettlemoyer, and V. Stoyanov, “Unsupervised cross-lingual representation learning at scale,” in Proc. of ACL , 2020, pp. 8440–8451
2020
Earlier work this paper cites.
M. Lewis, Y. Liu, N. Goyal, M. Ghazvininejad, A. Mohamed, O. Levy, V. Stoyanov, and L. Zettlemoyer, “BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension,” in Proc. of ACL , 2020, pp. 7871–7880
2020
Earlier work this paper cites.
R. Rei, C. Stewart, A. C. Farinha, and A. Lavie, “Unbabel’s participation in the WMT20 metrics shared task,” in Proceedings of the Fifth Conference on Machine Translation , 2020, pp. 911–920
2020
Earlier work this paper cites.
N. Roberts, D. Liang, G. Neubig, and Z. C. Lipton, “Decoding and diversity in machine translation,” ArXiv preprint , 2020
2020
Cited alongside, same era.
A. Holtzman, J. Buys, L. Du, M. Forbes, and Y. Choi, “The curious case of neural text degeneration,” in Proc. of ICLR , 2020
2020
Cited alongside, same era.
W. Yuan, G. Neubig, and P. Liu, “Bartscore: Evaluating generated text as text generation,” in Proc. of NeurIPS , 2021, pp. 27 263–27 277
2021
Cited alongside, same era.
T. Schick and H. Schütze, “Generating datasets with pretrained language models,” in Proc. of EMNLP , 2021, pp. 6943–6951
2021
Cited alongside, same era.
J. Austin, A. Odena, M. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. Cai, M. Terry, Q. Le et al. , “Program synthesis with large language models,” ArXiv preprint , 2021
2021
Cited alongside, same era.
Y. Liu, D. Iter, Y. Xu, S. Wang, R. Xu, and C. Zhu, “Gpteval: Nlg evaluation using gpt-4 with better human alignment,” ArXiv preprint , 2023
2023
Closest in time.
Z. Luo, Q. Xie, and S. Ananiadou, “Chatgpt as a factual inconsistency evaluator for abstractive text summarization,” ArXiv preprint , 2023
2023
Closest in time.
N. Wu, M. Gong, L. Shou, S. Liang, and D. Jiang, “Large language models are diverse role-players for summarization evaluation,” ArXiv preprint , 2023
2023
Closest in time.
A. M. Tripathi and O. J. Pandey, “Divide and distill: new outlooks on knowledge distillation for environmental sound classification,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , pp. 1100–1113, 2023
2023
Closest in time.
N. Ho, L. Schmid, and S.-Y. Yun, “Large language models are reasoning teachers,” in Proc. of ACL , 2023, pp. 14 852–14 882
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Kiegeland and J. Kreutzer, “Revisiting the weaknesses of reinforcement learning for neural machine translation,” in Proc. of NAACL , 2021, pp. 1673–1681
2021
Cited alongside, same era.
A. Lee, M. Auli, and M. Ranzato, “Discriminative reranking for neural machine translation,” in Proc. of ACL , 2021, pp. 7250–7264
2021
Cited alongside, same era.
R. Liu, Z. Lin, and W. Wang, “Addressing extraction and generation separately: Keyphrase prediction with pre-trained language models,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , pp. 3180–3191, 2021
2021
Cited alongside, same era.
R. Shu, K. M. Yoo, and J.-W. Ha, “Reward optimization for neural machine translation with learned metrics,” ArXiv preprint , 2021
2021
Cited alongside, same era.
C. Hu, C. Wang, X. Ma, X. Meng, Y. Li, T. Xiao, J. Zhu, and C. Li, “RankNAS: Efficient neural architecture search by pairwise ranking,” in Proc. of EMNLP , 2021, pp. 2469–2480
2021
Cited alongside, same era.
H. Lai, A. Toral, and M. Nissim, “Thank you BART! rewarding pre-trained models improves formality style transfer,” in Proc. of ACL , 2021, pp. 484–494
2021
Cited alongside, same era.
A. R. Fabbri, W. Kryściński, B. McCann, C. Xiong, R. Socher, and D. Radev, “SummEval: Re-evaluating summarization evaluation,” TACL , pp. 391–409, 2021
2021
Cited alongside, same era.
2023
Closest in time.
C.-Y. Hsieh, C.-L. Li, C.-k. Yeh, H. Nakhost, Y. Fujii, A. Ratner, R. Krishna, C.-Y. Lee, and T. Pfister, “Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes,” in Proc. of ACL Findings , 2023, pp. 8003–8017
2023
Closest in time.
H. Lee, S. Phatale, H. Mansoor, K. Lu, T. Mesnard, C. Bishop, V. Carbune, and A. Rastogi, “Rlaif: Scaling reinforcement learning from human feedback with ai feedback,” ArXiv preprint , 2023
2023
Closest in time.
G. Cui, L. Yuan, N. Ding, G. Yao, W. Zhu, Y. Ni, G. Xie, Z. Liu, and M. Sun, “Ultrafeedback: Boosting language models with high-quality feedback,” ArXiv preprint , 2023
2023
Closest in time.
IDEA-CCNL, “Ziya-llama-7b-reward,” 2023
2023
Closest in time.
A. Mohtashami, M. Verzetti, and P. K. Rubenstein, “Learning translation quality evaluation on low resource languages from large language models,” ArXiv preprint , 2023
2023
Closest in time.
B. Ding, C. Qin, L. Liu, Y. K. Chia, B. Li, S. Joty, and L. Bing, “Is GPT-3 a good data annotator?” in Proc. of ACL , 2023, pp. 11 173–11 195
2023
Closest in time.
J. Kang, W. Xu, and A. Ritter, “Distill or annotate? cost-efficient fine-tuning of compact models,” in Proc. of ACL , 2023, pp. 11 100–11 119
2023
Closest in time.
C.-H. Chiang and H.-y. Lee, “Can large language models be an alternative to human evaluations?” in Proc. of ACL , 2023, pp. 15 607–15 631
2023
Closest in time.
T. Xiao and J. Zhu, “Introduction to transformers: an nlp perspective,” ArXiv preprint , 2023
2023
Closest in time.
R. Rei, N. M. Guerreiro, J. Pombal, D. van Stigt, M. Treviso, L. Coheur, J. G. C. de Souza, and A. Martins, “Scaling up CometKiwi: Unbabel-IST 2023 submission for the quality estimation shared task,” in Proceedings of the Eighth Conference on Machine Translation , 2023, pp. 841–848
2023
Closest in time.
H. Dong, W. Xiong, D. Goyal, R. Pan, S. Diao, J. Zhang, K. Shum, and T. Zhang, “Raft: Reward ranked finetuning for generative foundation model alignment,” ArXiv preprint , 2023
2023
Closest in time.
K. Wang, J. Zhu, M. Ren, Z. Liu, S. Li, Z. Zhang, C. Zhang, X. Wu, Q. Zhan, Q. Liu et al. , “A survey on data synthesis and augmentation for large language models,” ArXiv preprint , 2024
2024
Closest in time.
J. Fu, S.-K. Ng, Z. Jiang, and P. Liu, “GPTScore: Evaluate as you desire,” in Proc. of NAACL , 2024, pp. 6556–6576
2024
Closest in time.
C. Chan, W. Chen, Y. Su, J. Yu, W. Xue, S. Zhang, J. Fu, and Z. Liu, “Chateval: Towards better llm-based evaluators through multi-agent debate,” in Proc. of ICLR , 2024
2024
Closest in time.
Z. He, X. Wang, W. Jiao, Z. Zhang, R. Wang, S. Shi, and Z. Tu, “Improving machine translation with human feedback: An exploration of quality estimation as a reward model,” in Proc. of NAACL , 2024, pp. 8164–8180
2024
Closest in time.
C. Wang, H. Zhou, Y. Hu, Y. Huo, B. Li, T. Liu, T. Xiao, and J. Zhu, “ESRL: efficient sampling-based reinforcement learning for sequence generation,” in Proc. of AAAI , 2024, pp. 19 107–19 115
2024
Closest in time.
C. Burns, P. Izmailov, J. H. Kirchner, B. Baker, L. Gao, L. Aschenbrenner, Y. Chen, A. Ecoffet, M. Joglekar, J. Leike, I. Sutskever, and J. Wu, “Weak-to-strong generalization: Eliciting strong capabilities with weak supervision,” in Proc. of ICML , 2024
2024
Closest in time.
A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan et al. , “The llama 3 herd of models,” ArXiv preprint , 2024
2024
Closest in time.
A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei et al. , “Qwen2. 5 technical report,” ArXiv preprint , 2024
2024
Closest in time.
Z. Wang, Y. Dong, O. Delalleau, J. Zeng, G. Shen, D. Egert, J. Zhang, M. N. Sreedhar, and O. Kuchaiev, “Helpsteer 2: Open-source dataset for training top-performing reward models,” in Proc. of NeurIPS , 2024
2024
Closest in time.
T. Xiao and J. Zhu, “Foundations of large language models,” 2025
2025
Closest in time.
Y. Lin, Y. Li, Z. Wang, B. Li, Q. Du, T. Xiao, and J. Zhu, “Weight distillation: Transferring the knowledge in neural network parameters,” in Proc. of ACL , 2021, pp. 2076–2088
2088
Closest in time.