Fetching the paper…
Reading the bibliography…
Recent work has explored the capability of large language models (LLMs) to identify and correct errors in LLM-generated responses.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Binary codes capable of correcting deletions, insertions, and reversals
Vladimir I Levenshtein et al. 1966 · 1966
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. 2018 · 2018
Earlier work this paper cites.
Inherent disagreements in human textual inferences
Ellie Pavlick and Tom Kwiatkowski. 2019 · 2019
Earlier work this paper cites.
Uncertain natural language inference
Tongfei Chen, Zhengping Jiang, Adam Poliak, Keisuke Sakaguchi, and Benjamin Van Durme. 2020 · 2020
Earlier work this paper cites.
TRL: Transformer Reinforcement Learning
Leandro von Werra, Younes Belkada, Lewis Tunstall, Edward Beeching, Tristan Thrush, Nathan Lambert, and Shengyi Huang. 2020 · 2020
Earlier work this paper cites.
What can we learn from collective human opinions on natural language inference data?
Yixin Nie, Xiang Zhou, and Mohit Bansal. 2020 · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021 · 2021
Earlier work this paper cites.
Investigating memorization of conspiracy theories in text generation
Sharon Levy, Michael Saxon, and William Yang Wang. 2021 · 2021
Earlier work this paper cites.
TruthfulQA: Measuring How Models Mimic Human Falsehoods
Stephanie C. Lin, Jacob Hilton, and Owain Evans. 2021 · 2021
Earlier work this paper cites.
Copy that! editing sequences by copying spans
Sheena Panthaplackel, Miltiadis Allamanis, and Marc Brockschmidt. 2021 · 2021
Earlier work this paper cites.
Evidence-based factual error correction
James Thorne and Andreas Vlachos. 2021 · 2021
Earlier work this paper cites.
MediaSum: A large-scale media interview dataset for dialogue summarization
Chenguang Zhu, Yang Liu, Jie Mei, and Michael Zeng. 2021 · 2021
Earlier work this paper cites.
Correcting diverse factual errors in abstractive summarization via post-editing and language model infilling
Vidhisha Balachandran, Hannaneh Hajishirzi, William Cohen, and Yulia Tsvetkov. 2022 · 2022
Earlier work this paper cites.
CodeT: Code Generation with Generated Tests
Bei Chen, Fengji Zhang, Anh Nguyen, Daoguang Zan, Zeqi Lin, Jian-Guang Lou, and Weizhu Chen. 2022 · 2022
Earlier work this paper cites.
Improving factual consistency in summarization with compression-based post-editing
Alex Fabbri, Prafulla Kumar Choubey, Jesse Vig, Chien-Sheng Wu, and Caiming Xiong. 2022 · 2022
Earlier work this paper cites.
LoRA: Low-rank adaptation of large language models
Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022 · 2022
Earlier work this paper cites.
Self-critiquing models for assisting human evaluators
William Saunders, Catherine Yeh, Jeff Wu, Steven Bills, Long Ouyang, Jonathan Ward, and Jan Leike. 2022 · 2022
Earlier work this paper cites.
The Unreliability of Explanations in Few-shot Prompting for Textual Reasoning
Xi Ye and Greg Durrett. 2022 · 2022
Earlier work this paper cites.
RL4F: Generating natural language feedback with reinforcement learning for repairing model outputs
Afra Feyza Akyurek, Ekin Akyurek, Ashwin Kalyan, Peter Clark, Derry Tanti Wijaya, and Niket Tandon. 2023 · 2023
Earlier work this paper cites.
UltraFeedback: Boosting Language Models with High-quality Feedback
Ganqu Cui, Lifan Yuan, Ning Ding, Guanming Yao, Wei Zhu, Yuan Ni, Guotong Xie, Zhiyuan Liu, and Maosong Sun. 2023 · 2023
Earlier work this paper cites.
Enhancing chat language models by scaling high-quality instructional conversations
Ning Ding, Yulin Chen, Bokai Xu, Yujia Qin, Shengding Hu, Zhiyuan Liu, Maosong Sun, and Bowen Zhou. 2023 · 2023
Cited alongside, same era.
FLEEK: Factual error detection and correction with evidence retrieved from external knowledge
Farima Fatahi Bayat, Kun Qian, Benjamin Han, Yisi Sang, Anton Belyy, Samira Khorshidi, Fei Wu, Ihab Ilyas, and Yunyao Li. 2023 · 2023
Cited alongside, same era.
Self-verification improves few-shot clinical information extraction
Zelalem Gero, Chandan Singh, Hao Cheng, Tristan Naumann, Michel Galley, Jianfeng Gao, and Hoifung Poon. 2023 · 2023
Cited alongside, same era.
SelfEvolve: A Code Evolution Framework via Large Language Models
Shuyang Jiang, Yuhao Wang, and Yu Wang. 2023 · 2023
Cited alongside, same era.
Prometheus: Inducing evaluation capability in language models
Seungone Kim, Jamin Shin, Yejin Cho, Joel Jang, Shayne Longpre, Hwaran Lee, Sangdoo Yun, Seongjin Shin, Sungdong Kim, James Thorne, et al. 2023 · 2023
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael Jordan, Joseph E Gonzalez, et al. 2024 · 2024
Closest in time.
Alpacafarm: A simulation framework for methods that learn from human feedback
Yann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy S Liang, and Tatsunori B Hashimoto. 2024 · 2024
Closest in time.
CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing
Zhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen, Yujiu Yang, Nan Duan, and Weizhu Chen. 2024 · 2024
Closest in time.
TIGERScore: Building Explainable Metric for All Text Generation Task
Dongfu Jiang, Yishan Li, Ge Zhang, Wenhao Huang, Bill Yuchen Lin, and Wenhu Chen. 2024 · 2024
Closest in time.
Leveraging Large Language Models for NLG Evaluation: A Survey
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On improving summarization factual consistency from natural language feedback
Yixin Liu, Budhaditya Deb, Milagro Teruel, Aaron Halfaker, Dragomir Radev, and Ahmed Hassan Awadallah. 2023 · 2023
Cited alongside, same era.
Self-refine: Iterative refinement with self-feedback
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al. 2023 · 2023
Cited alongside, same era.
Leveraging GPT-4 for automatic translation post-editing
Vikas Raunak, Amr Sharaf, Yiren Wang, Hany Awadalla, and Arul Menezes. 2023 · 2023
Cited alongside, same era.
Improving automated prediction of English lexical blends through the use of observable linguistic features
Jarem Saunders. 2023 · 2023
Cited alongside, same era.
On second thought, let’s not think step by step! bias and toxicity in zero-shot reasoning
Omar Shaikh, Hongxin Zhang, William Held, Michael Bernstein, and Diyi Yang. 2023 · 2023
Cited alongside, same era.
Dynamic routing transformer network for multimodal sarcasm detection
Yuan Tian, Nan Xu, Ruike Zhang, and Wenji Mao. 2023 · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
Cited alongside, same era.
Zhen Li, Xiaohan Xu, Tao Shen, Can Xu, Jia-Chen Gu, and Chongyang Tao. 2024 · 2024
Closest in time.
Introducing Meta Llama 3: The most capable openly available LLM to date
Meta. 2024 · 2024
Closest in time.
Fine-grained hallucination detection and editing for language models
Abhika Mishra, Akari Asai, Vidhisha Balachandran, Yizhong Wang, Graham Neubig, Yulia Tsvetkov, and Hannaneh Hajishirzi. 2024 · 2024
Closest in time.
Is self-repair a silver bullet for code generation?
Theo X. Olausson, Jeevana Priya Inala, Chenglong Wang, Jianfeng Gao, and Armando Solar-Lezama. 2024 · 2024
Closest in time.
Automatically correcting large language models: Surveying the landscape of diverse automated correction strategies
Liangming Pan, Michael Saxon, Wenda Xu, Deepak Nathani, Xinyi Wang, and William Yang Wang. 2024 · 2024
Closest in time.
REFINER: Reasoning feedback on intermediate representations
Debjit Paul, Mete Ismayilzada, Maxime Peyrard, Beatriz Borges, Antoine Bosselut, Robert West, and Boi Faltings. 2024 · 2024
Closest in time.
Reflexion: Language agents with verbal reinforcement learning
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2024 · 2024
Closest in time.
Learning performance-improving code edits
Alexander G Shypula, Aman Madaan, Yimeng Zeng, Uri Alon, Jacob R. Gardner, Yiming Yang, Milad Hashemi, Graham Neubig, Parthasarathy Ranganathan, Osbert Bastani, and Amir Yazdanbakhsh. 2024 · 2024
Closest in time.
ReGAL: Refactoring Programs to Discover Generalizable Abstractions
Elias Stengel-Eskin, Archiki Prasad, and Mohit Bansal. 2024 · 2024
Closest in time.
TofuEval: Evaluating hallucinations of LLMs on topic-focused dialogue summarization
Liyan Tang, Igor Shalyminov, Amy Wong, Jon Burnsky, Jake Vincent, Yu’an Yang, Siffi Singh, Song Feng, Hwanjun Song, Hang Su, Justin Sun, Yi Zhang, Saab Mansour, and Kathleen McKeown. 2024b · 2024
Closest in time.
InfoLossQA: Characterizing and recovering information loss in text simplification
Jan Trienes, Sebastian Joseph, Jörg Schlötterer, Christin Seifert, Kyle Lo, Wei Xu, Byron C. Wallace, and Junyi Jessy Li. 2024 · 2024
Closest in time.
Language models don’t always say what they think: unfaithful explanations in chain-of-thought prompting
Miles Turpin, Julian Michael, Ethan Perez, and Samuel Bowman. 2024 · 2024
Closest in time.
Using natural language explanations to rescale human judgments
Manya Wadhwa, Jifan Chen, Junyi Jessy Li, and Greg Durrett. 2024 · 2024
Closest in time.
Large language models are not fair evaluators
Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Lingpeng Kong, Qi Liu, Tianyu Liu, and Zhifang Sui. 2024 · 2024
Closest in time.
LLMRefine: Pinpointing and refining large language models via fine-grained actionable feedback
Wenda Xu, Daniel Deutsch, Mara Finkelstein, Juraj Juraska, Biao Zhang, Zhongtao Liu, William Yang Wang, Lei Li, and Markus Freitag. 2024 · 2024
Closest in time.
How language model hallucinations can snowball
Muru Zhang, Ofir Press, William Merrill, Alisa Liu, and Noah A. Smith. 2024 · 2024
Closest in time.