Fetching the paper…
Reading the bibliography…
Multi-agent collaboration among models has shown promise in reasoning tasks but is underexplored in long-form generation tasks like summarization and question-answering.
Eli5: Long form question answering
Angela Fan, Yacine Jernite, Ethan Perez, David Grangier, Jason Weston, and Michael Auli. 2019 · 1907
Earlier work this paper cites.
Make the most of prior data: A solution for interactive text summarization with preference feedback
Duy-Hung Nguyen, Nguyen Viet Dung Nghiem, Bao-Sinh Nguyen, Dung Tien Tien Le, Shahab Sabahi, Minh-Tien Nguyen, and Hung Le. 2022 · 1930
Earlier work this paper cites.
Statistical significance tests for machine translation evaluation
Philipp Koehn. 2004 · 2004
Earlier work this paper cites.
Natural language processing with Python: analyzing text with the natural language toolkit
Steven Bird, Ewan Klein, and Edward Loper. 2009 · 2009
Earlier work this paper cites.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano. 2020 · 2020
Earlier work this paper cites.
Summac: Re-visiting nli-based models for inconsistency detection in summarization
Philippe Laban, Tobias Schnabel, Paul N. Bennett, and Marti A. Hearst. 2021 · 2021
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, Xu Jiang, Karl Cobbe, Tyna Eloundou, Gretchen Krueger, Kevin Button, Matthew Knight, Benjamin Chess, and John Schulman. 2021 · 2021
Earlier work this paper cites.
Evidence-based factual error correction
James Thorne and Andreas Vlachos. 2021 · 2021
Earlier work this paper cites.
Recursively summarizing books with human feedback
Jeff Wu, Long Ouyang, Daniel M. Ziegler, Nisan Stiennon, Ryan Lowe, Jan Leike, and Paul Christiano. 2021 · 2021
Earlier work this paper cites.
MediaSum: A large-scale media interview dataset for dialogue summarization
Chenguang Zhu, Yang Liu, Jie Mei, and Michael Zeng. 2021 · 2021
Earlier work this paper cites.
Correcting diverse factual errors in abstractive summarization via post-editing and language model infilling
Vidhisha Balachandran, Hannaneh Hajishirzi, William Cohen, and Yulia Tsvetkov. 2022 · 2022
Earlier work this paper cites.
Hallucinated but factual! inspecting the factuality of hallucinations in abstractive summarization
Meng Cao, Yue Dong, and Jackie Cheung. 2022 · 2022
Earlier work this paper cites.
Improving factual consistency in summarization with compression-based post-editing
Alex Fabbri, Prafulla Kumar Choubey, Jesse Vig, Chien-Sheng Wu, and Caiming Xiong. 2022 · 2022
Earlier work this paper cites.
Mutual information alleviates hallucinations in abstractive summarization
Liam van der Poel, Ryan Cotterell, and Clara Meister. 2022 · 2022
Earlier work this paper cites.
RL4F: Generating natural language feedback with reinforcement learning for repairing model outputs
Afra Feyza Akyurek, Ekin Akyurek, Ashwin Kalyan, Peter Clark, Derry Tanti Wijaya, and Niket Tandon. 2023 · 2023
Earlier work this paper cites.
Chateval: Towards better llm-based evaluators through multi-agent debate
Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu. 2023 · 2023
Earlier work this paper cites.
Improving factuality of abstractive summarization via contrastive reward learning
I-chun Chern, Zhiruo Wang, Sanjan Das, Bhavuk Sharma, Pengfei Liu, and Graham Neubig. 2023b · 2023
Earlier work this paper cites.
Can large language models be an alternative to human evaluations?
Cheng-Han Chiang and Hung-yi Lee. 2023 · 2023
Earlier work this paper cites.
Enhancing chat language models by scaling high-quality instructional conversations
Ning Ding, Yulin Chen, Bokai Xu, Yujia Qin, Shengding Hu, Zhiyuan Liu, Maosong Sun, and Bowen Zhou. 2023 · 2023
Earlier work this paper cites.
Improving factuality and reasoning in language models through multiagent debate
Yilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum, and Igor Mordatch. 2023 · 2023
Earlier work this paper cites.
Human-like summarization evaluation with chatgpt
Mingqi Gao, Jie Ruan, Renliang Sun, Xunjian Yin, Shiping Yang, and Xiaojun Wan. 2023 · 2023
Cited alongside, same era.
Self-verification improves few-shot clinical information extraction
Zelalem Gero, Chandan Singh, Hao Cheng, Tristan Naumann, Michel Galley, Jianfeng Gao, and Hoifung Poon. 2023 · 2023
Cited alongside, same era.
Hallucinations in large multilingual translation models
Nuno M. Guerreiro, Duarte M. Alves, Jonas Waldendorf, Barry Haddow, Alexandra Birch, Pierre Colombo, and André F. T. Martins. 2023 · 2023
Cited alongside, same era.
MeetingBank: A benchmark dataset for meeting summarization
Yebowen Hu, Timothy Ganter, Hanieh Deilamsalehy, Franck Dernoncourt, Hassan Foroosh, and Fei Liu. 2023 · 2023
Cited alongside, same era.
Selfevolve: A code evolution framework via large language models
Shuyang Jiang, Yuhao Wang, and Yu Wang. 2023 · 2023
Cited alongside, same era.
CRITIC: Large language models can self-correct with tool-interactive critiquing
Zhibin Gou, Zhihong Shao, Yeyun Gong, yelong shen, Yujiu Yang, Nan Duan, and Weizhu Chen. 2024 · 2024
Later among the works it cites.
Improving llm reasoning with multi-agent tree-of-thought validator agent
Fatemeh Haji, Mazal Bethany, Maryam Tabar, Jason Chiang, Anthony Rios, and Peyman Najafirad. 2024 · 2024
Later among the works it cites.
Embrace divergence for richer insights: A multi-document summarization benchmark and a case study on summarizing diverse information from news articles
Kung-Hsiang Huang, Philippe Laban, Alexander Fabbri, Prafulla Kumar Choubey, Shafiq Joty, Caiming Xiong, and Chien-Sheng Wu. 2024b · 2024
Later among the works it cites.
Kyungha Kim, Sangyun Lee, Kung-Hsiang Huang, Hou Pong Chan, Manling Li, and Heng Ji. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Encouraging divergent thinking in large language models through multi-agent debate
Tian Liang, Zhiwei He, Wenxiang Jiao, Xing Wang, Yan Wang, Rui Wang, Yujiu Yang, Zhaopeng Tu, and Shuming Shi. 2023 · 2023
Cited alongside, same era.
G-eval: NLG evaluation using gpt-4 with better human alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023a · 2023
Cited alongside, same era.
Self-refine: Iterative refinement with self-feedback
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark. 2023 · 2023
Cited alongside, same era.
Leveraging GPT-4 for automatic translation post-editing
Vikas Raunak, Amr Sharaf, Yiren Wang, Hany Awadalla, and Arul Menezes. 2023 · 2023
Cited alongside, same era.
The troubling emergence of hallucination in large language models - an extensive definition, quantification, and prescriptive remediations
Vipula Rawte, Swagata Chakraborty, Agnibh Pathak, Anubhav Sarkar, S.M Towhidul Islam Tonmoy, Aman Chadha, Amit Sheth, and Amitava Das. 2023 · 2023
Cited alongside, same era.
Improving automated prediction of English lexical blends through the use of observable linguistic features
Jarem Saunders. 2023 · 2023
Cited alongside, same era.
Understanding factual errors in summarization: Errors, summarizers, datasets, error detectors
Liyan Tang, Tanya Goyal, Alex Fabbri, Philippe Laban, Jiacheng Xu, Semih Yavuz, Wojciech Kryscinski, Justin Rousseau, and Greg Durrett. 2023 · 2023
Cited alongside, same era.
Zhen Li, Xiaohan Xu, Tao Shen, Can Xu, Jia-Chen Gu, Yuxuan Lai, Chongyang Tao, and Shuai Ma. 2024 · 2024
Later among the works it cites.
Benchmarking generation and evaluation capabilities of large language models for instruction controllable summarization
Yixin Liu, Alexander Fabbri, Jiawen Chen, Yilun Zhao, Simeng Han, Shafiq Joty, Pengfei Liu, Dragomir Radev, Chien-Sheng Wu, and Arman Cohan. 2024 · 2024
Later among the works it cites.
LLM discussion: Enhancing the creativity of large language models via discussion framework and role-play
Li-Chun Lu, Shou-Jen Chen, Tsung-Min Pai, Chan-Hung Yu, Hung yi Lee, and Shao-Hua Sun. 2024 · 2024
Later among the works it cites.
Is self-repair a silver bullet for code generation?
Theo X. Olausson, Jeevana Priya Inala, Chenglong Wang, Jianfeng Gao, and Armando Solar-Lezama. 2024 · 2024
Later among the works it cites.
Hello gpt-4o
OpenAI. 2024 · 2024
Later among the works it cites.
REFINER: Reasoning feedback on intermediate representations
Debjit Paul, Mete Ismayilzada, Maxime Peyrard, Beatriz Borges, Antoine Bosselut, Robert West, and Boi Faltings. 2024 · 2024
Later among the works it cites.
Training language models with language feedback at scale
Jérémy Scheurer, Jon Ander Campos, Tomasz Korbak, Jun Shern Chan, Angelica Chen, Kyunghyun Cho, and Ethan Perez. 2024 · 2024
Later among the works it cites.
Veriscore: Evaluating the factuality of verifiable claims in long-form text generation
Yixiao Song, Yekyung Kim, and Mohit Iyyer. 2024 · 2024
Later among the works it cites.
Towards detecting llms hallucination via markov chain-based multi-agent debate framework
Xiaoxi Sun, Jinpeng Li, Yan Zhong, Dongyan Zhao, and Rui Yan. 2024 · 2024
Later among the works it cites.
Minicheck: Efficient fact-checking of llms on grounding documents
Liyan Tang, Philippe Laban, and Greg Durrett. 2024a · 2024
Later among the works it cites.
TofuEval: Evaluating hallucinations of LLMs on topic-focused dialogue summarization
Liyan Tang, Igor Shalyminov, Amy Wong, Jon Burnsky, Jake Vincent, Yu’an Yang, Siffi Singh, Song Feng, Hwanjun Song, Hang Su, Lijia Sun, Yi Zhang, Saab Mansour, and Kathleen McKeown. 2024b · 2024
Later among the works it cites.
Tofueval: Evaluating hallucinations of llms on topic-focused dialogue summarization
Liyan Tang, Igor Shalyminov, Amy Wong, Jon Burnsky, Jake Vincent, Yu’an Yang, Siffi Singh, Song Feng, Hwanjun Song, Hang Su, Justin Sun, Yi Zhang, Saab Mansour, and Kathleen McKeown. 2024c · 2024
Later among the works it cites.
Learning to refine with fine-grained natural language feedback
Manya Wadhwa, Xinyu Zhao, Junyi Jessy Li, and Greg Durrett. 2024 · 2024
Later among the works it cites.
ACUEval: Fine-grained hallucination evaluation and correction for abstractive summarization
David Wan, Koustuv Sinha, Srini Iyer, Asli Celikyilmaz, Mohit Bansal, and Ramakanth Pasunuru. 2024 · 2024
Later among the works it cites.
Unleashing the emergent cognitive synergy in large language models: A task-solving agent through multi-persona self-collaboration
Zhenhailong Wang, Shaoguang Mao, Wenshan Wu, Tao Ge, Furu Wei, and Heng Ji. 2024 · 2024
Later among the works it cites.
LLMRefine: Pinpointing and refining large language models via fine-grained actionable feedback
Wenda Xu, Daniel Deutsch, Mara Finkelstein, Juraj Juraska, Biao Zhang, Zhongtao Liu, William Yang Wang, Lei Li, and Markus Freitag. 2024 · 2024
Later among the works it cites.