Fetching the paper…
Reading the bibliography…
Large language models (LLMs) often exhibit deficient reasoning or generate hallucinations.
A. Tarski, Introduction to logic: And to the methodology of deductive sciences . Oxford University Press, 1941
1941
Earlier work this paper cites.
M. Jin, Q. Yu et al. , “The impact of reasoning step length on large language models,” in Proc. of ACL Findings , 2024, pp. 1830–1842
1999
Earlier work this paper cites.
I. Harvey, “The microbial genetic algorithm,” in Advances in Artificial Life. Darwin Meets von Neumann , 2011, pp. 126–133
2011
Earlier work this paper cites.
Y. Gal and Z. Ghahramani, “Dropout as a bayesian approximation: Representing model uncertainty in deep learning,” in Proc. of ICML , 2016, pp. 1050–1059
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
A. Ignatiev, A. Morgado, and J. Marques-Silva, “RC2: An Efficient MaxSAT Solver,” Journal on Satisfiability, Boolean Modeling and Computation , pp. 53–64, 2019
2019
Earlier work this paper cites.
M. T. Pilehvar and J. Camacho-Collados, “Wic: the word-in-context dataset for evaluating context-sensitive meaning representations,” Proceedings of NAACL 2019 (short) , 2019
2019
Earlier work this paper cites.
2020
Earlier work this paper cites.
D. Hendrycks, C. Burns et al. , “Measuring massive multitask language understanding,” in Proc. of ICLR , 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
E. M. Bender, T. Gebru et al. , “On the dangers of stochastic parrots: Can language models be too big?” in Proceedings of the 2021 ACM conference on fairness, accountability, and transparency , 2021, pp. 610–623
2021
Earlier work this paper cites.
Y. Xiao and W. Y. Wang, “On hallucination and predictive uncertainty in conditional language generation,” in Proc. of EACL , 2021, pp. 2734–2744
2021
Earlier work this paper cites.
W. Yuan, G. Neubig, and P. Liu, “BARTScore: Evaluating generated text as text generation,” in Proc. of NeurIPS , 2021
2021
Earlier work this paper cites.
Y. Elazar, N. Kassner et al. , “Measuring and improving consistency in pretrained language models,” TACL , pp. 1012–1031, 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
D. Hendrycks, C. Burns et al. , “Measuring mathematical problem solving with the math dataset,” NeurIPS , 2021
2021
Earlier work this paper cites.
J. Wei, X. Wang et al. , “Chain of thought prompting elicits reasoning in large language models,” in Proc. of NeurIPS , 2022
2022
Earlier work this paper cites.
S. Lin, J. Hilton, and O. Evans, “TruthfulQA: Measuring how models mimic human falsehoods,” in Proc. of ACL , 2022, pp. 3214–3252
2022
Earlier work this paper cites.
Y. Ma, D. Tsao, and H.-Y. Shum, “On the principles of parsimony and self-consistency for the emergence of intelligence,” Frontiers of Information Technology & Electronic Engineering , pp. 1298–1323, 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
E. Mitchell, J. Noh et al. , “Enhancing self-consistency and performance of pre-trained language models through natural language inference,” in Proc. of EMNLP , 2022, pp. 1754–1768
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
J. Jung, L. Qin et al. , “Maieutic prompting: Logically consistent reasoning with recursive explanations,” in Proc. of EMNLP , 2022, pp. 1266–1279
2022
Earlier work this paper cites.
A. Agarwal, A. Tzen, and C. Tew, “Improving logical consistency in pre-trained language models using natural language inference,” 2022
2022
Earlier work this paper cites.
K. Yang, Y. Tian et al. , “Re3: Generating longer stories with recursive reprompting and revision,” in Proc. of EMNLP , 2022, pp. 4393–4479
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
K. Ethayarajh, Y. Choi, and S. Swayamdipta, “Understanding dataset difficulty with 𝒱 \mathcal{V} -usable information,” in Proc. of ICML , 2022, pp. 5988–6008
2022
Earlier work this paper cites.
L. Ouyang, J. Wu et al. , “Training language models to follow instructions with human feedback,” Proc. of NeurIPS , pp. 27 730–27 744, 2022
2022
Earlier work this paper cites.
M. Jang, D. S. Kwon, and T. Lukasiewicz, “BECEL: Benchmark for consistency evaluation of language models,” in Proc. of COLING , 2022, pp. 3680–3696
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
X. Wang, J. Wei et al. , “Self-consistency improves chain of thought reasoning in language models,” in Proc. of ICLR , 2023
2023
Earlier work this paper cites.
Z. Yin, Q. Sun et al. , “Do large language models know what they don’t know?” in Proc. of ACL Findings , 2023, pp. 8653–8665
2023
Earlier work this paper cites.
K. Li, O. Patel et al. , “Inference-time intervention: Eliciting truthful answers from a language model,” in Proc. of NeurIPS , 2023
2023
Earlier work this paper cites.
A. Madaan, N. Tandon et al. , “Self-refine: Iterative refinement with self-feedback,” in Proc. of NeurIPS , 2023
2023
Earlier work this paper cites.
S. Welleck, X. Lu et al. , “Generating sequences by learning to self-correct,” in Proc. of ICLR , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
B. Liu, J. T. Ash et al. , “Exposing attention glitches with flip-flop language modeling,” in Proc. of NeurIPS , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
S. Yao, D. Yu et al. , “Tree of thoughts: Deliberate problem solving with large language models,” in Proc. of NeurIPS , 2023
2023
Earlier work this paper cites.
J. Huang, S. Gu et al. , “Large language models can self-improve,” in Proc. of EMNLP , 2023, pp. 1051–1068
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
K. Xiong, X. Ding et al. , “Examining inter-consistency of large language models collaboration: An in-depth analysis via debate,” in Proc. of EMNLP Findings , 2023, pp. 7572–7590
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
N. Shinn, F. Cassano et al. , “Reflexion: language agents with verbal reinforcement learning,” in Proc. of NeurIPS , 2023, pp. 8634–8652
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
D. Deng, G. Chen et al. , “Uncertainty estimation by fisher information-based evidential deep learning,” in Proc. of ICML , 2023, pp. 7596–7616
2023
Cited alongside, same era.
M. Besta, N. Blach et al. , “Graph of thoughts: Solving elaborate problems with large language models,” Proc. of AAAI , pp. 17 682–17 690, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
D. Paul, M. Ismayilzada et al. , “REFINER: Reasoning feedback on intermediate representations,” in Proc. of EACL , 2024, pp. 1100–1126
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
2023
Cited alongside, same era.
P. Manakul, A. Liusie, and M. Gales, “SelfCheckGPT: Zero-resource black-box hallucination detection for generative large language models,” in Proc. of EMNLP , 2023, pp. 9004–9017
2023
Cited alongside, same era.
R. Cohen, M. Hamri et al. , “LM vs LM: Detecting factual errors via cross examination,” in Proc. of EMNLP , 2023, pp. 12 621–12 640
2023
Cited alongside, same era.
J. Fu, S.-K. Ng et al. , “Gptscore: Evaluate as you desire,” arXiv preprint arXiv:2302.04166 , 2023
2023
Cited alongside, same era.
Y. Xie, K. Kawaguchi et al. , “Self-evaluation guided beam search for reasoning,” in Proc. of NeurIPS , 2023
2023
Cited alongside, same era.
H. Chen, A. Saha et al. , “Personalized distillation: Empowering open-sourced LLMs with adaptive learning for code generation,” in Proc. of EMNLP , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2024
Closest in time.
X. Chen, M. Lin et al. , “Teaching large language models to self-debug,” in Proc. of ICLR , 2024
2024
Closest in time.
2024
Closest in time.
Y.-S. Chuang, Y. Xie et al. , “Dola: Decoding by contrasting layers improves factuality in large language models,” in Proc. of ICLR , 2024
2024
Closest in time.
2024
Closest in time.
Z. Lin, S. Trivedi, and J. Sun, “Generating with confidence: Uncertainty quantification for black-box large language models,” TMLR , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
C. Chen, K. Liu et al. , “INSIDE: LLMs’ internal states retain the power of hallucination detection,” in Proc. of ICLR , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
O. Khattab, A. Singhvi et al. , “DSPy: Compiling declarative language model calls into state-of-the-art pipelines,” in Proc. of ICLR , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
L. Wang, C. Ma et al. , “A survey on large language model based autonomous agents,” Frontiers of Computer Science , p. 186345, 2024
2024
Closest in time.
J. Huang, X. Chen et al. , “Large language models cannot self-correct reasoning yet,” in Proc. of ICLR , 2024
2024
Closest in time.
A. P. Jacob, Y. Shen et al. , “The consensus game: Language model generation via equilibrium search,” in Proc. of ICLR , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Z. Chen, Y. Deng et al. , “Self-play fine-tuning converts weak language models to strong language models,” in Proc. of ICML , 2024, pp. 6621–6642
2024
Closest in time.
X. Pang, S. Tang et al. , “Self-alignment of large language models via monopolylogue-based social scene simulation,” in Proc. of ICML , 2024, pp. 39 416–39 447
2024
Closest in time.
Z. Sun, Y. Shen et al. , “SALMON: Self-alignment with instructable reward models,” in Proc. of ICLR , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Y. Gu, L. Dong et al. , “MiniLLM: Knowledge distillation of large language models,” in Proc. of ICLR , 2024
2024
Closest in time.
R. Agarwal, N. Vieillard et al. , “On-policy distillation of language models: Learning from self-generated mistakes,” in Proc. of ICLR , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Y. Chang, X. Wang et al. , “A survey on evaluation of large language models,” ACM Trans. Intell. Syst. Technol. , 2024
2024
Closest in time.
Y. Huang, Y. Bai et al. , “C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models,” Proc. of NeurIPS , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
R. Battiti, Maximum satisfiability problemMaximum Satisfiability Problem , 2009, pp. 2035–2041
2041
Closest in time.