Fetching the paper…
Reading the bibliography…
Although contemporary large language models (LMs) demonstrate impressive question-answering capabilities, their answers are typically the product of a single call to the model.
Logic for Mathematicians
A. Hamilton · 1988
Earlier work this paper cites.
Generating sequences with recurrent neural networks
A. Graves · 2013
Earlier work this paper cites.
Adaptive computation time for recurrent neural networks
A. Graves · 2016
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord · 2018
Earlier work this paper cites.
Learning to explain: Datasets and models for identifying valid reasoning chains in multihop question-answering
H. Jhamtani and P. Clark · 2020
Earlier work this paper cites.
Nile: Natural language inference with faithful natural language explanations
S. Kumar and P. Talukdar · 2020
Earlier work this paper cites.
Prover: Proof generation for interpretable reasoning over rules
S. Saha, S. Ghosh, S. Srivastava, and M. Bansal · 2020
Earlier work this paper cites.
Bleurt: Learning robust metrics for text generation
T. Sellam, D. Das, and A. Parikh · 2020
Earlier work this paper cites.
Worldtree v2: A corpus of science-domain structured explanations and inference patterns supporting multi-hop inference
Z. Xie, S. Thiem, J. Martin, E. Wainwright, S. Marmorstein, and P. Jansen · 2020
Earlier work this paper cites.
A. Banino, J. Balaguer, and C. Blundell · 2021
Earlier work this paper cites.
On the dangers of stochastic parrots: Can language models be too big?
E. M. Bender, T. Gebru, A. McMillan-Major, and S. Shmitchell · 2021
Earlier work this paper cites.
Critical thinking for language models
G. Betz, C. Voigt, and K. Richardson · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, J. Hilton, R. Nakano, C. Hesse, and J. Schulman · 2021
Cited alongside, same era.
Explaining answers with entailment trees
B. Dalvi, P. Jansen, O. Tafjord, Z. Xie, H. Smith, L. Pipatanangkura, and P. Clark · 2021
Cited alongside, same era.
Webgpt: Browser-assisted question-answering with human feedback
R. Nakano, J. Hilton, S. Balaji, J. Wu, L. Ouyang, C. Kim, C. Hesse, S. Jain, V. Kosaraju, W. Saunders, et al · 2021
Cited alongside, same era.
Scaling language models: Methods, analysis & insights from training gopher
J. W. Rae, S. Borgeaud, T. Cai, K. Millican, J. Hoffmann, F. Song, J. Aslanides, S. Henderson, R. Ring, S. Young, et al · 2021
Cited alongside, same era.
Proofwriter: Generating implications, proofs, and abductive statements over natural language
Right for the right reason: Evidence extraction for trustworthy tabular reasoning
V. Gupta, S. Zhang, A. Vempala, Y. He, T. Choji, and V. Srikumar · 2022
Closest in time.
Training compute-optimal large language models
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark, et al · 2022
Closest in time.
Language models (mostly) know what they know
S. Kadavath, T. Conerly, A. Askell, T. Henighan, D. Drain, E. Perez, N. Schiefer, Z. H. Dodds, N. DasSarma, E. Tran-Johnson, et al · 2022
Closest in time.
Large language models are zero-shot reasoners
T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa · 2022
Closest in time.
Show your work: Scratchpads for intermediate computation with language models
M. Nye, A. J. Andreassen, G. Gur-Ari, H. Michalewski, J. Austin, D. Bieber, D. Dohan, A. Lewkowycz, M. Bosma, D. Luan, et al · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
O. Tafjord, B. Dalvi, and P. Clark · 2021
Cited alongside, same era.
Ethical and social risks of harm from language models
L. Weidinger, J. Mellor, M. Rauh, C. Griffin, J. Uesato, P. Huang, M. Cheng, M. Glaese, B. Balle, A. Kasirzadeh, Z. Kenton, S. Brown, W. Hawkins, T. Stepleton, C. Biles, A. Birhane, J. Haas, L. Rimell, L. A. Hendricks, W. S. Isaac, S. Legassick, G. Irving, and I. Gabriel · 2021
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, et al · 2022
Cited alongside, same era.
Natural language deduction through search over statement compositions
K. Bostrom, Z. Sprague, S. Chaudhuri, and G. Durrett · 2022
Cited alongside, same era.
Selection-inference: Exploiting large language models for interpretable logical reasoning
A. Creswell, M. Shanahan, and I. Higgins · 2022
Cited alongside, same era.
Towards teachable reasoning systems
B. Dalvi, O. Tafjord, and P. Clark · 2022
Cited alongside, same era.
Language models show human-like content effects on reasoning
I. Dasgupta, A. K. Lampinen, S. C. Chan, A. Creswell, D. Kumaran, J. L. McClelland, and F. Hill · 2022
Cited alongside, same era.
Closest in time.
Entailment tree explanations via iterative retrieval-generation reasoner
D. Ribeiro, S. Wang, X. Ma, R. Dong, X. Wei, H. Zhu, X. Chen, Z. Huang, P. Xu, A. Arnold, et al · 2022
Closest in time.
Chain of thought prompting elicits reasoning in large language models, 2022
J. Wei, X. Wang, D. Schuurmans, M. Bosma, E. Chi, Q. Le, and D. Zhou · 2022
Closest in time.
Star: Bootstrapping reasoning with reasoning
E. Zelikman, Y. Wu, and N. D. Goodman · 2022
Closest in time.
Socratic models: Composing zero-shot multimodal reasoning with language
A. Zeng, A. Wong, S. Welker, K. Choromanski, F. Tombari, A. Purohit, M. Ryoo, V. Sindhwani, J. Lee, V. Vanhoucke, et al · 2022
Closest in time.
On the paradox of learning to reason from data
H. Zhang, L. H. Li, T. Meng, K.-W. Chang, and G. V. d. Broeck · 2022
Closest in time.
Least-to-most prompting enables complex reasoning in large language models
D. Zhou, N. Schärli, L. Hou, J. Wei, N. Scales, X. Wang, D. Schuurmans, O. Bousquet, Q. Le, and E. Chi · 2022
Closest in time.