Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have demonstrated impressive performance in various NLP tasks, but they still suffer from challenges such as hallucination and weak numerical reasoning.
The probabilistic relevance framework: Bm25 and beyond
S. Robertson, H. Zaragoza, et al · 2009
Earlier work this paper cites.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
M. Joshi, E. Choi, D. Weld, and L. Zettlemoyer · 2017
Earlier work this paper cites.
FEVER: a large-scale dataset for fact extraction and VERification
J. Thorne, A. Vlachos, C. Christodoulopoulos, and A. Mittal · 2018
Earlier work this paper cites.
HotpotQA: A dataset for diverse, explainable multi-hop question answering
Z. Yang, P. Qi, S. Zhang, Y. Bengio, W. Cohen, R. Salakhutdinov, and C. D. Manning · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
SciREX: A challenge dataset for document-level information extraction
S. Jain, M. van Zuylen, H. Hajishirzi, and I. Beltagy · 2020
Earlier work this paper cites.
Dense passage retrieval for open-domain question answering
V. Karpukhin, B. Oguz, S. Min, P. Lewis, L. Wu, S. Edunov, D. Chen, and W.-t. Yih · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel, et al · 2020
Earlier work this paper cites.
A dataset for answering time-sensitive questions
W. Chen, X. Wang, and W. Y. Wang · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt · 2021
Earlier work this paper cites.
Towards unsupervised dense information retrieval with contrastive learning
G. Izacard, M. Caron, L. Hosseini, S. Riedel, P. Bojanowski, A. Joulin, and E. Grave · 2021
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
R. Nakano, J. Hilton, S. Balaji, J. Wu, L. Ouyang, C. Kim, C. Hesse, S. Jain, V. Kosaraju, W. Saunders, et al · 2021
Earlier work this paper cites.
Investigating the limitations of transformers with simple arithmetic tasks
R. Nogueira, Z. Jiang, and J. Lin · 2021
Earlier work this paper cites.
Ethical and social risks of harm from language models
L. Weidinger, J. Mellor, M. Rauh, C. Griffin, J. Uesato, P.-S. Huang, M. Cheng, M. Glaese, B. Balle, A. Kasirzadeh, et al · 2021
Earlier work this paper cites.
SituatedQA: Incorporating extra-linguistic contexts into QA
M. Zhang and E. Choi · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Y. Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, N. DasSarma, D. Drain, S. Fort, D. Ganguli, T. Henighan, et al · 2022
Earlier work this paper cites.
Improving language models by retrieving from trillions of tokens
S. Borgeaud, A. Mensch, J. Hoffmann, T. Cai, E. Rutherford, K. Millican, G. B. Van Den Driessche, J.-B. Lespiau, B. Damoc, A. Clark, et al · 2022
Earlier work this paper cites.
Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks, 2022
W. Chen, X. Ma, X. Wang, and W. W. Cohen · 2022
Earlier work this paper cites.
Palm: Scaling language modeling with pathways
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, et al · 2022
Earlier work this paper cites.
Scaling instruction-finetuned language models
H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y. Tay, W. Fedus, E. Li, X. Wang, M. Dehghani, S. Brahma, et al · 2022
Earlier work this paper cites.
Time-aware language models as temporal knowledge bases
B. Dhingra, J. R. Cole, J. M. Eisenschlos, D. Gillick, J. Eisenstein, and W. W. Cohen · 2022
Earlier work this paper cites.
Pal: Program-aided language models
L. Gao, A. Madaan, S. Zhou, U. Alon, P. Liu, Y. Yang, J. Callan, and G. Neubig · 2022
Earlier work this paper cites.
Few-shot learning with retrieval augmented language models
G. Izacard, P. Lewis, M. Lomeli, L. Hosseini, F. Petroni, T. Schick, J. Dwivedi-Yu, A. Joulin, S. Riedel, and E. Grave · 2022
Cited alongside, same era.
Realtime qa: What’s the answer right now?
J. Kasai, K. Sakaguchi, Y. Takahashi, R. L. Bras, A. Asai, X. Yu, D. Radev, N. A. Smith, Y. Choi, and K. Inui · 2022
Cited alongside, same era.
Large language models are zero-shot reasoners
T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa · 2022
Cited alongside, same era.
Solving quantitative reasoning problems with language models
A. Lewkowycz, A. Andreassen, D. Dohan, E. Dyer, H. Michalewski, V. Ramasesh, A. Slone, C. Anil, I. Schlag, T. Gutman-Solo, et al · 2022
Cited alongside, same era.
Unsupervised cross-task generalization via retrieval augmentation
B. Y. Lin, K. Tan, C. S. Miller, B. Tian, and X. Ren · 2022
Language models can solve computer tasks, 2023
G. Kim, P. Baldi, and S. McAleer · 2023
Closest in time.
Api-bank: A benchmark for tool-augmented llms, 2023
M. Li, F. Song, B. Yu, H. Yu, Z. Li, F. Huang, and Y. Li · 2023
Closest in time.
Chameleon: Plug-and-play compositional reasoning with large language models
P. Lu, B. Peng, H. Cheng, M. Galley, K.-W. Chang, Y. N. Wu, S.-C. Zhu, and J. Gao · 2023
Closest in time.
Gpt-4 technical report
OpenAI · 2023
Closest in time.
Introducing chatgpt, 2023
OpenAI · 2023
Closest in time.
Art: Automatic multi-step reasoning and tool-use for large language models
B. Paranjape, S. Lundberg, S. Singh, H. Hajishirzi, L. Zettlemoyer, and M. T. Ribeiro · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Dynamic prompt learning via policy gradient for semi-structured mathematical reasoning
P. Lu, L. Qiu, K.-W. Chang, Y. N. Wu, S.-C. Zhu, T. Rajpurohit, P. Clark, and A. Kalyan · 2022
Cited alongside, same era.
Reacc: A retrieval-augmented code completion framework
S. Lu, N. Duan, H. Han, D. Guo, S.-w. Hwang, and A. Svyatkovskiy · 2022
Cited alongside, same era.
Text and patterns: For effective chain of thought, it takes two to tango
A. Madaan and A. Yazdanbakhsh · 2022
Cited alongside, same era.
A. Mallen, A. Asai, V. Zhong, R. Das, H. Hajishirzi, and D. Khashabi · 2022
Cited alongside, same era.
Lila: A unified benchmark for mathematical reasoning
S. Mishra, M. Finlayson, P. Lu, L. Tang, S. Welleck, C. Baral, T. Rajpurohit, O. Tafjord, A. Sabharwal, P. Clark, et al · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Cited alongside, same era.
Talm: Tool augmented language models
A. Parisi, Y. Zhao, and N. Fiedel · 2022
Cited alongside, same era.
Closest in time.
Gorilla: Large language model connected with massive apis
S. G. Patil, T. Zhang, X. Wang, and J. E. Gonzalez · 2023
Closest in time.
B. Peng, M. Galley, P. He, H. Cheng, Y. Xie, Y. Hu, Q. Huang, L. Liden, Z. Yu, W. Chen, et al · 2023
Closest in time.
Tool learning with foundation models, 2023
Y. Qin, S. Hu, Y. Lin, W. Chen, N. Ding, G. Cui, Z. Zeng, Y. Huang, C. Xiao, C. Han, Y. R. Fung, Y. Su, H. Wang, C. Qian, R. Tian, K. Zhu, S. Liang, X. Shen, B. Xu, Z. Zhang, Y. Ye, B. Li, Z. Tang, J. Yi, Y. Zhu, Z. Dai, L. Yan, X. Cong, Y. Lu, W. Zhao, Y. Huang, J. Yan, X. Han, X. Sun, D. Li, J. Phang, C. Yang, T. Wu, H. Ji, Z. Liu, and M. Sun · 2023
Closest in time.
Toolformer: Language models can teach themselves to use tools
T. Schick, J. Dwivedi-Yu, R. Dessì, R. Raileanu, M. Lomeli, L. Zettlemoyer, N. Cancedda, and T. Scialom · 2023
Closest in time.
Hugginggpt: Solving ai tasks with chatgpt and its friends in huggingface
Y. Shen, K. Song, X. Tan, D. Li, W. Lu, and Y. Zhuang · 2023
Closest in time.
Replug: Retrieval-augmented black-box language models
W. Shi, S. Min, M. Yasunaga, M. Seo, R. James, M. Lewis, L. Zettlemoyer, and W.-t. Yih · 2023
Closest in time.
Reflexion: an autonomous agent with dynamic memory and self-reflection
N. Shinn, B. Labash, and A. Gopinath · 2023
Closest in time.
Adaplanner: Adaptive planning from feedback with language models, 2023
H. Sun, Y. Zhuang, L. Kong, B. Dai, and C. Zhang · 2023
Closest in time.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2023
Closest in time.
Describe, explain, plan and select: Interactive planning with large language models enables open-world multi-task agents, 2023
Z. Wang, S. Cai, A. Liu, X. Ma, and Y. Liang · 2023
Closest in time.
Wolfram|Alpha as the Way to Bring Computational Knowledge Superpowers to ChatGPT
S. Wolfram · 2023
Closest in time.
Visual chatgpt: Talking, drawing and editing with visual foundation models
C. Wu, S. Yin, W. Qi, X. Wang, Z. Tang, and N. Duan · 2023
Closest in time.
On the tool manipulation capability of open-source large language models
Q. Xu, F. Hong, B. Li, C. Hu, Z. Chen, and J. Zhang · 2023
Closest in time.
Weakly-supervised scientific document classification via retrieval-augmented multi-stage training
R. Xu, Y. Yu, J. C. Ho, and C. Yang · 2023
Closest in time.
Mm-react: Prompting chatgpt for multimodal reasoning and action
Z. Yang, L. Li, J. Wang, K. Lin, E. Azarnasab, F. Ahmed, Z. Liu, C. Liu, M. Zeng, and L. Wang · 2023
Closest in time.
React: Synergizing reasoning and acting in language models
S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. R. Narasimhan, and Y. Cao · 2023
Closest in time.
Graph-toolformer: To empower llms with graph reasoning ability via prompt augmented by chatgpt, 2023
J. Zhang · 2023
Closest in time.