Fetching the paper…
Reading the bibliography…
Recent large language models (LLMs) have shown remarkable performance in aligning generated text with user intentions across various tasks.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 1910
Earlier work this paper cites.
Better summarization evaluation with word embeddings for ROUGE
Jun-Ping Ng and Viktoria Abrecht. 2015 · 1930
Earlier work this paper cites.
Rhetorical structure theory: Toward a functional theory of text organization
William C Mann and Sandra A Thompson. 1988 · 1988
Earlier work this paper cites.
News analysis
Teun A Van Dijk. 1988 · 1988
Earlier work this paper cites.
The discourse-level structure of empirical abstracts: An exploratory study
Elizabeth DuRoss Liddy. 1991 · 1991
Earlier work this paper cites.
Centering: A framework for modeling the local coherence of discourse
Barbara J. Grosz, Aravind K. Joshi, and Scott Weinstein. 1995 · 1995
Earlier work this paper cites.
Manual and automatic evaluation of summaries
Chin-Yew Lin and Eduard Hovy. 2002 · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Zone analysis in biology articles as a basis for information extraction
Yoko Mizuta, Anna Korhonen, Tony Mullen, and Nigel Collier. 2006 · 2006
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with high levels of correlation with human judgments
Alon Lavie and Abhaya Agarwal. 2007 · 2007
Earlier work this paper cites.
The penn discourse treebank 2.0
Rashmi Prasad, Nikhil Dinesh, Alan Lee, Eleni Miltsakaki, Livio Robaldo, Aravind K Joshi, and Bonnie L Webber. 2008 · 2008
Earlier work this paper cites.
News as discourse
Teun A Van Dijk. 2013 · 2013
Earlier work this paper cites.
Towards coherent and cohesive long-form text generation
Woon Sang Cho, Pengchuan Zhang, Yizhe Zhang, Xiujun Li, Michel Galley, Chris Brockett, Mengdi Wang, and Jianfeng Gao. 2019 · 2019
Cited alongside, same era.
ELI5: Long form question answering
Angela Fan, Yacine Jernite, Ethan Perez, David Grangier, Jason Weston, and Michael Auli. 2019 · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Natural Questions: A Benchmark for Question Answering Research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019 · 2019
Cited alongside, same era.
Step-by-step: Separating planning from realization in neural data-to-text generation
Amit Moryossef, Yoav Goldberg, and Ido Dagan. 2019 · 2019
Cited alongside, same era.
Can transformer models measure coherence in text: Re-thinking the shuffle test
Philippe Laban, Luke Dai, Lucas Bandarkar, and Marti A. Hearst. 2021 · 2021
Later among the works it cites.
Recipe1m+: A dataset for learning cross-modal embeddings for cooking recipes and food images
Javier Marín, Aritro Biswas, Ferda Ofli, Nicholas Hynes, Amaia Salvador, Yusuf Aytar, Ingmar Weber, and Antonio Torralba. 2021 · 2021
Later among the works it cites.
Bartscore: Evaluating generated text as text generation
Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021 · 2021
Later among the works it cites.
Model criticism for long-form text generation
Yuntian Deng, Volodymyr Kuleshov, and Alexander Rush. 2022 · 2022
Later among the works it cites.
Plug-and-play recipe generation with content planning
Yinhong Liu, Yixuan Su, Ehsan Shareghi, and Nigel Collier. 2022 · 2022
Later among the works it cites.
Webgpt: Browser-assisted question-answering with human feedback
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019 · 2019
Cited alongside, same era.
MoverScore: Text generation evaluating with contextualized embeddings and earth mover distance
Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M. Meyer, and Steffen Eger. 2019 · 2019
Cited alongside, same era.
Discourse as a function of event: Profiling discourse structure in news articles around the main event
Prafulla Kumar Choubey, Aaron Lee, Ruihong Huang, and Lu Wang. 2020 · 2020
Cited alongside, same era.
GRUEN for evaluating linguistic quality of generated text
Wanzheng Zhu and Suma Bhat. 2020 · 2020
Cited alongside, same era.
Is incoherence surprising? targeted evaluation of coherence prediction from language models
Anne Beyer, Sharid Loáiciga, and David Schlangen. 2021 · 2021
Cited alongside, same era.
Long text generation by modeling sentence-level and discourse-level coherence
Jian Guan, Xiaoxi Mao, Changjie Fan, Zitao Liu, Wenbiao Ding, and Minlie Huang. 2021 · 2021
Cited alongside, same era.
Discourse probing of pretrained language models
Fajri Koto, Jey Han Lau, and Timothy Baldwin. 2021 · 2021
Cited alongside, same era.
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, Xu Jiang, Karl Cobbe, Tyna Eloundou, Gretchen Krueger, Kevin Button, Matthew Knight, Benjamin Chess, and John Schulman. 2022 · 2022
Later among the works it cites.
Sequentially controlled text generation
Alexander Spangher, Yao Ming, Xinyu Hua, and Nanyun Peng. 2022 · 2022
Later among the works it cites.
Instruct-sctg: Guiding sequential controlled text generation through instructions
Yinhong Liu, Yixuan Su, Ehsan Shareghi, and Nigel Collier. 2023 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
Later among the works it cites.
DiscoScore: Evaluating text generation with BERT and discourse coherence
Wei Zhao, Michael Strube, and Steffen Eger. 2023 · 2023
Later among the works it cites.
Aligning with human judgement: The role of pairwise preference in large language model evaluators
Yinhong Liu, Han Zhou, Zhijiang Guo, Ehsan Shareghi, Ivan Vulic, Anna Korhonen, and Nigel Collier. 2024 · 2024
Closest in time.