Fetching the paper…
Reading the bibliography…
Intrinsic self-correction was proposed to improve LLMs' responses via feedback prompts solely based on their inherent capability.
Boolq: Exploring the surprising difficulty of natural yes/no questions
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019 · 1905
Earlier work this paper cites.
The framing of decisions and the psychology of choice
Amos Tversky and Daniel Kahneman. 1981 · 1981
Earlier work this paper cites.
Divergence measures based on the shannon entropy
Jianhua Lin. 1991 · 1991
Earlier work this paper cites.
The role of rumination in depressive disorders and mixed anxiety/depressive symptoms
Susan Nolen-Hoeksema. 2000 · 2000
Earlier work this paper cites.
Clinical perfectionism: A cognitive–behavioural analysis
Roz Shafran, Zafra Cooper, and Christopher G Fairburn. 2002 · 2002
Earlier work this paper cites.
Alfworld: Aligning text and embodied environments for interactive learning
Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk, Adam Trischler, and Matthew Hausknecht. 2020 · 2010
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton. 2015 · 2015
Earlier work this paper cites.
The paradox of choice
Barry Schwartz. 2015 · 2015
Earlier work this paper cites.
The reflective practitioner: How professionals think in action
Donald A Schön. 2017 · 2017
Earlier work this paper cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W Cohen, Ruslan Salakhutdinov, and Christopher D Manning. 2018 · 2018
Earlier work this paper cites.
Shallow-deep networks: Understanding and mitigating network overthinking
Yigitcan Kaya, Sanghyun Hong, and Tudor Dumitras. 2019 · 2019
Earlier work this paper cites.
Interpreting gpt: The logit lens
Nostalgebraist. 2020 · 2020
Earlier work this paper cites.
Curve circuits
Nick Cammarata, Gabriel Goh, Shan Carter, Chelsea Voss, Ludwig Schubert, and Chris Olah. 2021 · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde De Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021 · 2021
Earlier work this paper cites.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, et al. 2021 · 2021
Earlier work this paper cites.
Causal abstractions of neural networks
Atticus Geiger, Hanson Lu, Thomas Icard, and Christopher Potts. 2021 · 2021
Earlier work this paper cites.
From theory to practice: the application of cognitive load theory to the practice of medicine
Adam Szulewski, Daniel Howes, Jeroen JG van Merriënboer, and John Sweller. 2021 · 2021
Earlier work this paper cites.
Calibrate before use: Improving few-shot performance of language models
Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021 · 2021
Cited alongside, same era.
Large language models are few-shot clinical information extractors
Monica Agrawal, Stefan Hegselmann, Hunter Lang, Yoon Kim, and David Sontag. 2022 · 2022
Cited alongside, same era.
The gifts of imperfection: Let go of who you think you’re supposed to be and embrace who you are
Brené Brown. 2022 · 2022
Cited alongside, same era.
Capturing failures of large language models via human cognitive biases
Erik Jones and Jacob Steinhardt. 2022 · 2022
Cited alongside, same era.
Large language models with controllable working memory
Daliang Li, Ankit Singh Rawat, Manzil Zaheer, Xin Wang, Michal Lukasik, Andreas Veit, Felix Yu, and Sanjiv Kumar. 2022 · 2022
Cited alongside, same era.
Ask again, then fail: Large language models’ vacillations in judgement
Qiming Xie, Zengzhi Wang, Yi Feng, and Rui Xia. 2023 · 2023
Later among the works it cites.
Cognitive overload: Jailbreaking large language models with overloaded logical thinking
Nan Xu, Fei Wang, Ben Zhou, Bang Zheng Li, Chaowei Xiao, and Muhao Chen. 2023 · 2023
Later among the works it cites.
Exploring collaboration mechanisms for llm agents: A social psychology view
Jintian Zhang, Xin Xu, Ningyu Zhang, Ruibo Liu, Bryan Hooi, and Shumin Deng. 2023 · 2023
Later among the works it cites.
Working memory capacity of chatgpt: An empirical study
Dongyu Gong, Xingchen Wan, and Dingmin Wang. 2024 · 2024
Closest in time.
Large language models cannot self-correct reasoning yet
Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, and Denny Zhou. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Generating sequences by learning to self-correct
Sean Welleck, Ximing Lu, Peter West, Faeze Brahman, Tianxiao Shen, Daniel Khashabi, and Yejin Choi. 2022 · 2022
Cited alongside, same era.
Rl4f: Generating natural language feedback with reinforcement learning for repairing model outputs
Afra Feyza Akyürek, Ekin Akyürek, Aman Madaan, Ashwin Kalyan, Peter Clark, Derry Wijaya, and Niket Tandon. 2023 · 2023
Cited alongside, same era.
Eliciting latent predictions from transformers with the tuned lens
Nora Belrose, Zach Furman, Logan Smith, Danny Halawi, Igor Ostrovsky, Lev McKinney, Stella Biderman, and Jacob Steinhardt. 2023 · 2023
Cited alongside, same era.
Can llm-generated misinformation be detected?
Canyu Chen and Kai Shu. 2023 · 2023
Cited alongside, same era.
Critic: Large language models can self-correct with tool-interactive critiquing
Zhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen, Yujiu Yang, Nan Duan, and Weizhu Chen. 2023 · 2023
Cited alongside, same era.
Human-like intuitive behavior and reasoning biases emerged in large language models but disappeared in chatgpt
Thilo Hagendorff, Sarah Fabi, and Michal Kosinski. 2023 · 2023
Cited alongside, same era.
Overthinking the truth: Understanding how language models process false demonstrations
Danny Halawi, Jean-Stanislas Denain, and Jacob Steinhardt. 2023 · 2023
Cited alongside, same era.
Closest in time.
When can llms actually correct their own mistakes? a critical survey of self-correction of llms
Ryo Kamoi, Yusen Zhang, Nan Zhang, Jiawei Han, and Rui Zhang. 2024 · 2024
Closest in time.
Why and when llm-based assistants can go wrong: Investigating the effectiveness of prompt-based interactions for software help-seeking
Anjali Khurana, Hariharan Subramonyam, and Parmit K Chilana. 2024 · 2024
Closest in time.
Large language models have intrinsic self-correction ability
Dancheng Liu, Amir Nassereldine, Ziming Yang, Chenhui Xu, Yuting Hu, Jiajie Li, Utkarsh Kumar, Changjae Lee, and Jinjun Xiong. 2024 · 2024
Closest in time.
Self-refine: Iterative refinement with self-feedback
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al. 2024 · 2024
Closest in time.
Countering reward over-optimization in llm with demonstration-guided reinforcement learning
Mathieu Rita, Florian Strub, Rahma Chaabouni, Paul Michel, Emmanuel Dupoux, and Olivier Pietquin. 2024 · 2024
Closest in time.
Reflexion: Language agents with verbal reinforcement learning
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2024 · 2024
Closest in time.
Large language models can self-correct with key condition verification
Zhenyu Wu, Qingkai Zeng, Zhihan Zhang, Zhaoxuan Tan, Chao Shen, and Meng Jiang. 2024 · 2024
Closest in time.
Course-correction: Safety alignment using synthetic preferences
Rongwu Xu, Yishuo Cai, Zhenhong Zhou, Renjie Gu, Haiqin Weng, Liu Yan, Tianwei Zhang, Wei Xu, and Han Qiu. 2024a · 2024
Closest in time.
The earth is flat because…: Investigating llms’ belief towards misinformation via persuasive conversation
Rongwu Xu, Brian S Lin, Shujian Yang, Tianqi Zhang, Weiyan Shi, Tianwei Zhang, Zhixuan Fang, Wei Xu, and Han Qiu. 2024b · 2024
Closest in time.
Walking in others’ shoes: How perspective-taking guides large language models in reducing toxicity and bias
Rongwu Xu, Zi’an Zhou, Tianwei Zhang, Zehan Qi, Su Yao, Ke Xu, Wei Xu, and Han Qiu. 2024c · 2024
Closest in time.
Quiet-star: Language models can teach themselves to think before speaking
Eric Zelikman, Georges Harik, Yijia Shao, Varuna Jayasiri, Nick Haber, and Noah D Goodman. 2024 · 2024
Closest in time.
Promptbench: A unified library for evaluation of large language models
Kaijie Zhu, Qinlin Zhao, Hao Chen, Jindong Wang, and Xing Xie. 2024 · 2024
Closest in time.