Fetching the paper…
Reading the bibliography…
Generative Artificial Intelligence (AI) has enabled the development of sophisticated models that are capable of producing high-caliber text, images, and other outputs through the utilization of large pre-trained models.
Roberta: A robustly optimized BERT pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 1907
Earlier work this paper cites.
Towards a human-like open-domain chatbot
Adiwardana, D., Luong, M., So, D. R., Hall, J., Fiedel, N., Thoppilan, R., Yang, Z., Kulshreshtha, A., Nemade, G., Lu, Y., and Le, Q. V · 2001
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Lin, C.-Y · 2004
Earlier work this paper cites.
Evaluating content selection in summarization: The pyramid method
Nenkova, A. and Passonneau, R · 2004
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2005
Earlier work this paper cites.
Spearman rank correlation
Zar, J. H · 2005
Earlier work this paper cites.
COMET: A neural framework for MT evaluation
Rei, R., Stewart, C., Farinha, A. C., and Lavie, A · 2009
Earlier work this paper cites.
Phrase-based statistical language generation using graphical models and active learning
Mairesse, F., Gasic, M., Jurcícek, F., Keizer, S., Thomson, B., Yu, K., and Young, S. J · 2010
Earlier work this paper cites.
A guide to appropriate use of correlation coefficient in medical research
Mukaka, M. M · 2012
Earlier work this paper cites.
Teaching machines to read and comprehend
Hermann, K. M., Kociský, T., Grefenstette, E., Espeholt, L., Kay, W., Suleyman, M., and Blunsom, P · 2015
Earlier work this paper cites.
From word embeddings to document distances
Kusner, M. J., Sun, Y., Kolkin, N. I., and Weinberger, K. Q · 2015
Earlier work this paper cites.
chrf: character n-gram f-score for automatic MT evaluation
Popovic, M · 2015
Earlier work this paper cites.
Semantically conditioned lstm-based natural language generation for spoken dialogue systems
Wen, T., Gasic, M., Mrksic, N., Su, P., Vandyke, D., and Young, S. J · 2015
Earlier work this paper cites.
Newsroom: A dataset of 1.3 million summaries with diverse extractive strategies
Grusky, M., Naaman, M., and Artzi, Y · 2018
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Large-scale, diverse, paraphrastic bitexts via sampling and clustering
Hu, J. E., Singh, A., Holzenberger, N., Post, M., and Durme, B. V · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Cited alongside, same era.
Moverscore: Text generation evaluating with contextualized embeddings and earth mover distance
Zhao, W., Peyrard, M., Liu, F., Gao, Y., Meyer, C. M., and Eger, S · 2019
Cited alongside, same era.
Re-evaluating evaluation in text summarization
Bhandari, M., Gour, P. N., Ashfaq, A., Liu, P., and Neubig, G · 2020
Cited alongside, same era.
BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., Stoyanov, V., and Zettlemoyer, L · 2020
Cited alongside, same era.
Unsupervised evaluation of interactive dialog with dialogpt
Mehri, S. and Eskénazi, M · 2020
Cited alongside, same era.
Towards holistic and automatic evaluation of open-domain dialogue generation
Questeval: Summarization asks for fact-based evaluation
Scialom, T., Dray, P., Lamprier, S., Piwowarski, B., Staiano, J., Wang, A., and Gallinari, P · 2021
Later among the works it cites.
Bartscore: Evaluating generated text as text generation
Yuan, W., Neubig, G., and Liu, P · 2021
Later among the works it cites.
Dynaeval: Unifying turn and dialogue level evaluation
Zhang, C., Chen, Y., D’Haro, L. F., Zhang, Y., Friedrichs, T., Lee, G., and Li, H · 2021
Later among the works it cites.
Robust preference learning for storytelling via contrastive reinforcement learning
Castricato, L., Havrilla, A., Matiana, S., Pieler, M., Ye, A., Yang, I., Frazier, S., and Riedl, M. O · 2022
Later among the works it cites.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., Schuh, P., Shi, K., Tsvyashchenko, S., Maynez, J., Rao, A., Barnes, P., Tay, Y., Shazeer, N., Prabhakaran, V., Reif, E., Du, N., Hutchinson, B., Pope, R., Bradbury, J., Austin, J., Isard, M., Gur-Ari, G., Yin, P., Duke, T., Levskaya, A., Ghemawat, S., Dev, S., Michalewski, H., Garcia, X., Misra, V., Robinson, K., Fedus, L., Zhou, D., Ippolito, D., Luan, D., Lim, H., Zoph, B., Spiridonov, A., Sepassi, R., Dohan, D., Agrawal, S., Omernick, M., Dai, A. M., Pillai, T. S., Pellat, M., Lewkowycz, A., Moreira, E., Child, R., Polozov, O., Lee, K., Zhou, Z., Wang, X., Saeta, B., Diaz, M., Firat, O., Catasta, M., Wei, J., Meier-Hellstern, K., Eck, D., Dean, J., Petrov, S., and Fiedel, N · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pang, B., Nijkamp, E., Han, W., Zhou, L., Liu, Y., and Tu, K · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Cited alongside, same era.
BLEURT: learning robust metrics for text generation
Sellam, T., Das, D., and Parikh, A. P · 2020
Cited alongside, same era.
Automatic machine translation evaluation in many languages via zero-shot paraphrasing
Thompson, B. and Post, M · 2020
Cited alongside, same era.
Asking and answering questions to evaluate the factual consistency of summaries
Wang, A., Cho, K., and Lewis, M · 2020
Cited alongside, same era.
Bertscore: Evaluating text generation with BERT
Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., and Artzi, Y · 2020
Cited alongside, same era.
Summeval: Re-evaluating summarization evaluation
Fabbri, A. R., Kryscinski, W., McCann, B., Xiong, C., Socher, R., and Radev, D. R · 2021
Cited alongside, same era.
Later among the works it cites.
Scaling instruction-finetuned language models
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, E., Wang, X., Dehghani, M., Brahma, S., et al · 2022
Later among the works it cites.
Polyglot prompt: Multilingual multitask promptraining
Fu, J., Ng, S.-K., and Liu, P · 2022
Later among the works it cites.
Revisiting the gold standard: Grounding summarization evaluation with robust human evaluation
Liu, Y., Fabbri, A. R., Liu, P., Zhao, Y., Nan, L., Han, R., Han, S., Joty, S., Wu, C.-S., Xiong, C., et al · 2022
Later among the works it cites.
Rethinking the role of demonstrations: What makes in-context learning work?
Min, S., Lyu, X., Holtzman, A., Artetxe, M., Lewis, M., Hajishirzi, H., and Zettlemoyer, L · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Later among the works it cites.
BLOOM: A 176b-parameter open-access multilingual language model
Scao, T. L., Fan, A., Akiki, C., Pavlick, E., Ilic, S., Hesslow, D., Castagné, R., Luccioni, A. S., Yvon, F., Gallé, M., Tow, J., Rush, A. M., Biderman, S., Webson, A., Ammanamanchi, P. S., Wang, T., Sagot, B., Muennighoff, N., del Moral, A. V., Ruwase, O., Bawden, R., Bekman, S., McMillan-Major, A., Beltagy, I., Nguyen, H., Saulnier, L., Tan, S., Suarez, P. O., Sanh, V., Laurençon, H., Jernite, Y., Launay, J., Mitchell, M., Raffel, C., Gokaslan, A., Simhi, A., Soroa, A., Aji, A. F., Alfassy, A., Rogers, A., Nitzav, A. K., Xu, C., Mou, C., Emezue, C., Klamm, C., Leong, C., van Strien, D., Adelani, D. I., and et al · 2022
Later among the works it cites.
Generative ai: A creative new world
Sequoia, T · 2022
Later among the works it cites.
Super-naturalinstructions: Generalization via declarative instructions on 1600+ nlp tasks
Wang, Y., Mishra, S., Alipoormolabashi, P., Kordi, Y., Mirzaei, A., Arunkumar, A., Ashok, A., Dhanasekaran, A. S., Naik, A., Stap, D., et al · 2022
Later among the works it cites.
Chain of thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Chi, E. H., Le, Q., and Zhou, D · 2022
Later among the works it cites.
Not all errors are equal: Learning text generation metrics using stratified error synthesis
Xu, W., Tuan, Y., Lu, Y., Saxon, M., Li, L., and Wang, W. Y · 2022
Later among the works it cites.
Towards a unified multi-dimensional evaluator for text generation
Zhong, M., Liu, Y., Yin, D., Mao, Y., Jiao, Y., Liu, P., Zhu, C., Ji, H., and Han, J · 2022
Later among the works it cites.