Fetching the paper…
Reading the bibliography…
We explore how iterative revising a chain of thoughts with the help of information retrieval significantly improves large language models' reasoning and generation ability in long-horizon generation tasks, while hugely mitigating hallucination.
Trueskill™: a bayesian skill rating system
R. Herbrich, T. Minka, and T. Graepel · 2006
Earlier work this paper cites.
The Oxford handbook of thinking and reasoning
K. J. Holyoak and R. G. Morrison · 2012
Earlier work this paper cites.
Search engine guided neural machine translation
J. Gu, Y. Wang, K. Cho, and V. O. Li · 2018
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
N. Reimers and I. Gurevych · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
Program synthesis with large language models
J. Austin, A. Odena, M. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. Cai, M. Terry, Q. Le, et al · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. d. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman · 2021
Earlier work this paper cites.
Show your work: Scratchpads for intermediate computation with language models
M. Nye, A. J. Andreassen, G. Gur-Ari, H. Michalewski, J. Austin, D. Bieber, D. Dohan, A. Lewkowycz, M. Bosma, D. Luan, et al · 2021
Earlier work this paper cites.
Docprompting: Generating code by retrieving the docs
S. Zhou, U. Alon, F. F. Xu, Z. Jiang, and G. Neubig · 2021
Earlier work this paper cites.
Video pretraining (vpt): Learning to act by watching unlabeled online videos
B. Baker, I. Akkaya, P. Zhokhov, J. Huizinga, J. Tang, A. Ecoffet, B. Houghton, R. Sampedro, and J. Clune · 2022
Earlier work this paper cites.
Faithful reasoning using large language models
A. Creswell and M. Shanahan · 2022
Earlier work this paper cites.
Selection-inference: Exploiting large language models for interpretable logical reasoning
A. Creswell, M. Shanahan, and I. Higgins · 2022
Earlier work this paper cites.
Pal: Program-aided language models
L. Gao, A. Madaan, S. Zhou, U. Alon, P. Liu, Y. Yang, J. Callan, and G. Neubig · 2022
Earlier work this paper cites.
Language models as zero-shot planners: Extracting actionable knowledge for embodied agents
W. Huang, P. Abbeel, D. Pathak, and I. Mordatch · 2022
Earlier work this paper cites.
Adapting a language model while preserving its general knowledge
Z. Ke, Y. Shao, H. Lin, H. Xu, L. Shu, and B. Liu · 2022
Earlier work this paper cites.
Large language models are zero-shot reasoners
T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa · 2022
Earlier work this paper cites.
Reacc: A retrieval-augmented code completion framework
S. Lu, N. Duan, H. Han, D. Guo, S.-w. Hwang, and A. Svyatkovskiy · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Cited alongside, same era.
Entailment tree explanations via iterative retrieval-generation reasoner
D. Ribeiro, S. Wang, X. Ma, R. Dong, X. Wei, H. Zhu, X. Chen, Z. Huang, P. Xu, A. Arnold, et al · 2022
Cited alongside, same era.
Palm: Scaling language modeling with pathways
G. P. Team · 2022
Cited alongside, same era.
Self-consistency improves chain of thought reasoning in language models
X. Wang, J. Wei, D. Schuurmans, Q. Le, E. Chi, S. Narang, A. Chowdhery, and D. Zhou · 2022
Cited alongside, same era.
Chain of thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, E. Chi, Q. Le, and D. Zhou · 2022
Gpt-4 technical report, 2023
OpenAI · 2023
Later among the works it cites.
A survey of hallucination in large foundation models
V. Rawte, A. Sheth, and A. Das · 2023
Later among the works it cites.
Code llama: Open foundation models for code
B. Rozière, J. Gehring, F. Gloeckle, S. Sootla, I. Gat, X. Tan, Y. Adi, J. Liu, T. Remez, J. Rapin, A. Kozhevnikov, I. Evtimov, J. Bitton, M. P. Bhatt, C. C. Ferrer, A. Grattafiori, W. Xiong, A. D’efossez, J. Copet, F. Azhar, H. Touvron, L. Martin, N. Usunier, T. Scialom, and G. Synnaeve · 2023
Later among the works it cites.
Reflexion: an autonomous agent with dynamic memory and self-reflection
N. Shinn, B. Labash, and A. Gopinath · 2023
Later among the works it cites.
Improving the domain adaptation of retrieval augmented generation (rag) models for open domain question answering
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
React: Synergizing reasoning and acting in language models
S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao · 2022
Cited alongside, same era.
Star: Bootstrapping reasoning with reasoning
E. Zelikman, Y. Wu, J. Mu, and N. Goodman · 2022
Cited alongside, same era.
Self-rag: Learning to retrieve, generate, and critique through self-reflection
A. Asai, Z. Wu, Y. Wang, A. Sil, and H. Hajishirzi · 2023
Cited alongside, same era.
Knowledge-augmented language model prompting for zero-shot knowledge graph question answering
J. Baek, A. F. Aji, and A. Saffari · 2023
Cited alongside, same era.
Graph of thoughts: Solving elaborate problems with large language models
M. Besta, N. Blach, A. Kubicek, R. Gerstenberger, L. Gianinazzi, J. Gajda, T. Lehmann, M. Podstawski, H. Niewiadomski, P. Nyczyk, et al · 2023
Cited alongside, same era.
Open-world multi-task control through goal-aware representation learning and adaptive horizon prediction
S. Cai, Z. Wang, X. Ma, A. Liu, and Y. Liang · 2023
Cited alongside, same era.
Chain-of-verification reduces hallucination in large language models
S. Dhuliawala, M. Komeili, J. Xu, R. Raileanu, X. Li, A. Celikyilmaz, and J. Weston · 2023
Cited alongside, same era.
S. Siriwardhana, R. Weerasekera, E. Wen, T. Kaluarachchi, R. Rana, and S. Nanayakkara · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al · 2023
Later among the works it cites.
Self-consistency improves chain of thought reasoning in language models
X. Wang, J. Wei, D. Schuurmans, Q. V. Le, E. H. Chi, S. Narang, A. Chowdhery, and D. Zhou · 2023
Later among the works it cites.
Grove: a retrieval-augmented complex story generation framework with a forest of evidence
Z. Wen, Z. Tian, W. Wu, Y. Yang, Y. Shi, Z. Huang, and D. Li · 2023
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models, 2023
S. Yao, D. Yu, J. Zhao, I. Shafran, T. L. Griffiths, Y. Cao, and K. Narasimhan · 2023
Later among the works it cites.
Plan4mc: Skill reinforcement learning and planning for open-world minecraft tasks
H. Yuan, C. Zhang, H. Wang, F. Xie, P. Cai, H. Dong, and Z. Lu · 2023
Later among the works it cites.
Proagent: Building proactive cooperative ai with large language models
C. Zhang, K. Yang, S. Hu, Z. Wang, G. Li, Y. Sun, C. Zhang, Z. Zhang, A. Liu, S.-C. Zhu, et al · 2023
Later among the works it cites.
Retrieving multimodal information for augmented generation: A survey
R. Zhao, H. Chen, W. Wang, F. Jiao, X. L. Do, C. Qin, B. Ding, X. Guo, M. Li, X. Li, and S. R. Joty · 2023
Later among the works it cites.
Least-to-most prompting enables complex reasoning in large language models
D. Zhou, N. Schärli, L. Hou, J. Wei, N. Scales, X. Wang, D. Schuurmans, C. Cui, O. Bousquet, Q. V. Le, and E. H. Chi · 2023
Later among the works it cites.
Deepseek-coder: When the large language model meets programming – the rise of code intelligence
D. Guo, Q. Zhu, D. Yang, Z. Xie, K. Dong, W. Zhang, G. Chen, X. Bi, Y. Wu, Y. K. Li, F. Luo, Y. Xiong, and W. Liang · 2024
Closest in time.
Selecting large language model to fine-tune via rectified scaling law
H. Lin, B. Huang, H. Ye, Q. Chen, Z. Wang, S. Li, J. Ma, X. Wan, J. Zou, and Y. Liang · 2024
Closest in time.
Chain-of-thought reasoning without prompting
X. Wang and D. Zhou · 2024
Closest in time.
Pre-training goal-based models for sample-efficient reinforcement learning
H. Yuan, Z. Mu, F. Xie, and Z. Lu · 2024
Closest in time.