Fetching the paper…
Reading the bibliography…
Recent work has shown that asking language models to generate reasoning steps improves performance on many reasoning tasks.
Minimum bayes-risk decoding for statistical machine translation
S. Kumar and W. Byrne · 2004
Earlier work this paper cites.
Generative language modeling for automated theorem proving, 2020
S. Polu and I. Sutskever · 2009
Earlier work this paper cites.
On the foundations of noise-free selective classification
R. El-Yaniv et al · 2010
Earlier work this paper cites.
A. Graves, G. Wayne, and I. Danihelka · 2014
Earlier work this paper cites.
Learning to automatically solve algebra word problems
N. Kushman, Y. Artzi, L. Zettlemoyer, and R. Barzilay · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Earlier work this paper cites.
Neural programmer-interpreters
S. Reed and N. De Freitas · 2015
Earlier work this paper cites.
Concrete problems in ai safety
D. Amodei, C. Olah, J. Steinhardt, P. Christiano, J. Schulman, and D. Mané · 2016
Earlier work this paper cites.
Adaptive computation time for recurrent neural networks
A. Graves · 2016
Earlier work this paper cites.
Neural program lattices
C. Li, D. Tarlow, A. L. Gaunt, M. Brockschmidt, and N. Kushman · 2016
Earlier work this paper cites.
Thinking fast and slow with deep learning and tree search
T. Anthony, Z. Tian, and D. Barber · 2017
Earlier work this paper cites.
Making neural programming architectures generalize via recursion
J. Cai, R. Shin, and D. Song · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
Reinforcement learning with a corrupted reward channel
T. Everitt, V. Krakovna, L. Orseau, M. Hutter, and S. Legg · 2017
Earlier work this paper cites.
Selective classification for deep neural networks
Y. Geifman and R. El-Yaniv · 2017
Earlier work this paper cites.
Weakly-supervised semantic parsing with abstract examples
O. Goldman, V. Latcinnik, U. Naveh, A. Globerson, and J. Berant · 2017
Earlier work this paper cites.
Program induction by rationale generation: Learning to solve and explain algebraic word problems
W. Ling, D. Yogatama, C. Dyer, and P. Blunsom · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2017
Earlier work this paper cites.
Mastering the game of go without human knowledge
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, et al · 2017
Earlier work this paper cites.
Maximum a posteriori policy optimisation
A. Abdolmaleki, J. T. Springenberg, Y. Tassa, R. Munos, N. Heess, and M. Riedmiller · 2018
Earlier work this paper cites.
Supervising strong learners by amplifying weak experts
P. Christiano, B. Shlegeris, and D. Amodei · 2018
Earlier work this paper cites.
M. Dehghani, S. Gouws, O. Vinyals, J. Uszkoreit, and Ł. Kaiser · 2018
Earlier work this paper cites.
G. Irving, P. Christiano, and D. Amodei · 2018
Cited alongside, same era.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Z. Yang, P. Qi, S. Zhang, Y. Bengio, W. W. Cohen, R. Salakhutdinov, and C. D. Manning · 2018
Cited alongside, same era.
Mathqa: Towards interpretable math word problem solving with operation-based formalisms
A. Amini, S. Gabriel, P. Lin, R. Koncel-Kedziorski, Y. Choi, and H. Hajishirzi · 2019
Cited alongside, same era.
An investigation of model-free planning
A. Guez, M. Mirza, K. Gregor, R. Kabra, S. Racanière, T. Weber, D. Raposo, A. Santoro, L. Orseau, T. Eccles, et al · 2019
Cited alongside, same era.
Z. Kenton, T. Everitt, L. Weidinger, I. Gabriel, V. Mikulik, and G. Irving · 2021
Later among the works it cites.
A graph placement methodology for fast chip design
A. Mirhoseini, A. Goldie, M. Yazgan, J. W. Jiang, E. Songhori, S. Wang, Y.-J. Lee, E. Johnson, O. Pathak, A. Nazi, et al · 2021
Later among the works it cites.
Webgpt: Browser-assisted question-answering with human feedback
R. Nakano, J. Hilton, S. Balaji, J. Wu, L. Ouyang, C. Kim, C. Hesse, S. Jain, V. Kosaraju, W. Saunders, et al · 2021
Later among the works it cites.
Show your work: Scratchpads for intermediate computation with language models
M. Nye, A. J. Andreassen, G. Gur-Ari, H. Michalewski, J. Austin, D. Bieber, D. Dohan, A. Lewkowycz, M. Bosma, D. Luan, et al · 2021
Later among the works it cites.
True few-shot learning with language models
E. Perez, D. Kiela, and K. Cho · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P. Christiano, and G. Irving · 2019
Cited alongside, same era.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Cited alongside, same era.
Measuring systematic generalization in neural proof generation with transformers
N. Gontier, K. Sinha, S. Reddy, and C. Pal · 2020
Cited alongside, same era.
Specification gaming: the flip side of ai ingenuity, 2020
V. Krakovna, J. Uesato, V. Mikulik, M. Rahtz, T. Everitt, R. Kumar, Z. Kenton, J. Leike, and S. Legg · 2020
Cited alongside, same era.
REALab: An embedded perspective on tampering
R. Kumar, J. Uesato, R. Ngo, T. Everitt, V. Krakovna, and S. Legg · 2020
Cited alongside, same era.
A diverse corpus for evaluating and developing english math word problem solvers
S.-y. Miao, C.-C. Liang, and K.-Y. Su · 2020
Cited alongside, same era.
Unsupervised question decomposition for question answering
E. Perez, P. Lewis, W.-t. Yih, K. Cho, and D. Kiela · 2020
Cited alongside, same era.
Mastering atari, go, chess and shogi by planning with a learned model
J. Schrittwieser, I. Antonoglou, T. Hubert, K. Simonyan, L. Sifre, S. Schmitt, A. Guez, E. Lockhart, D. Hassabis, T. Graepel, et al · 2020
Cited alongside, same era.
Later among the works it cites.
Recursively summarizing books with human feedback
J. Wu, L. Ouyang, D. M. Ziegler, N. Stiennon, R. Lowe, J. Leike, and P. Christiano · 2021
Later among the works it cites.
Palm: Scaling language modeling with pathways
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, et al · 2022
Closest in time.
Without specific countermeasures, the easiest path to transformative ai likely leads to ai takeover, 2022
A. Cotra · 2022
Closest in time.
Selection-inference: Exploiting large language models for interpretable logical reasoning
A. Creswell, M. Shanahan, and I. Higgins · 2022
Closest in time.
D. Dohan, W. Xu, A. Lewkowycz, J. Austin, D. Bieber, R. G. Lopes, Y. Wu, H. Michalewski, R. A. Saurous, J. Sohl-dickstein, et al · 2022
Closest in time.
Training compute-optimal large language models
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark, T. Hennigan, E. Noland, K. Millican, G. v. d. Driessche, B. Damoc, A. Guy, S. Osindero, K. Simonyan, E. Elsen, J. W. Rae, O. Vinyals, and L. Sifre · 2022
Closest in time.
Large language models are zero-shot reasoners
T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa · 2022
Closest in time.
Solving quantitative reasoning problems with language models, 2022
A. Lewkowycz, A. Andreassen, D. Dohan, E. Dyer, H. Michalewski, V. Ramasesh, A. Slone, C. Anil, I. Schlag, T. Gutman-Solo, Y. Wu, B. Neyshabur, G. Gur-Ari, and V. Misra · 2022
Closest in time.
On the advance of making language models better reasoners
Y. Li, Z. Lin, S. Zhang, Q. Fu, B. Chen, J.-G. Lou, and W. Chen · 2022
Closest in time.
Teaching language models to support answers with verified quotes
J. Menick, M. Trebacz, V. Mikulik, J. Aslanides, F. Song, M. Chadwick, M. Glaese, S. Young, L. Campbell-Gillingham, G. Irving, et al · 2022
Closest in time.
Training language models to follow instructions with human feedback, 2022
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. Christiano, J. Leike, and R. Lowe · 2022
Closest in time.
Characteristics of harmful text: Towards rigorous benchmarking of language models
M. Rauh, J. Mellor, J. Uesato, P.-S. Huang, J. Welbl, L. Weidinger, S. Dathathri, A. Glaese, G. Irving, I. Gabriel, et al · 2022
Closest in time.
Supervise process, not outcomes
A. Stuhlmüller and J. Byun · 2022
Closest in time.
Self-consistency improves chain of thought reasoning in language models
X. Wang, J. Wei, D. Schuurmans, Q. Le, E. Chi, and D. Zhou · 2022
Closest in time.
Chain of thought prompting elicits reasoning in large language models, 2022
J. Wei, X. Wang, D. Schuurmans, M. Bosma, E. Chi, Q. Le, and D. Zhou · 2022
Closest in time.
Star: Bootstrapping reasoning with reasoning
E. Zelikman, Y. Wu, and N. D. Goodman · 2022
Closest in time.