Fetching the paper…
Reading the bibliography…
Chain-of-thought (CoT) significantly enhances the reasoning performance of large language models (LLM).
Efficient noise-tolerant learning from statistical queries
Michael Kearns · 1998
Earlier work this paper cites.
Poly-time universality and limitations of deep learning, 2020
Emmanuel Abbe and Colin Sandon · 2001
Earlier work this paper cites.
Learning parities with neural networks, 2020
Amit Daniely and Eran Malach · 2002
Earlier work this paper cites.
Noise-tolerant learning, the parity problem, and the statistical query model
Avrim Blum, Adam Kalai, and Hal Wasserman · 2003
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma · 2014
Earlier work this paper cites.
Densely connected convolutional networks
Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger · 2017
Earlier work this paper cites.
Time-space hardness of learning sparse parities
Gillat Kol, Ran Raz, and Avishay Tal · 2017
Earlier work this paper cites.
Failures of gradient-based deep learning, 2017
Shai Shalev-Shwartz, Ohad Shamir, and Shaked Shammah · 2017
Earlier work this paper cites.
Provable limitations of deep learning
Emmanuel Abbe and Colin Sandon · 2018
Earlier work this paper cites.
High-Dimensional Probability: An Introduction with Applications in Data Science , volume 47 of Cambridge Series in Statistical and Probabilistic Mathematics
Roman Vershynin · 2018
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al · 2021
Earlier work this paper cites.
Data distributional properties drive emergent in-context learning in transformers
Stephanie Chan, Adam Santoro, Andrew Lampinen, Jane Wang, Aaditya Singh, Pierre Richemond, James McClelland, and Felix Hill · 2022
Earlier work this paper cites.
An empirical analysis of compute-optimal large language model training
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katherine Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Oriol Vinyals, Jack William Rae, and Laurent Sifre · 2022
Earlier work this paper cites.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa · 2022
Earlier work this paper cites.
Solving quantitative reasoning problems with language models
Aitor Lewkowycz, Anders Johan Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Venkatesh Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra · 2022
Earlier work this paper cites.
When hardness of approximation meets hardness of learning
Eran Malach and Shai Shalev-Shwartz · 2022
Earlier work this paper cites.
Show your work: Scratchpads for intermediate computation with language models
Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, et al · 2022
Earlier work this paper cites.
Chain of thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed H. Chi, Quoc V Le, and Denny Zhou · 2022
Earlier work this paper cites.
STar: Bootstrapping reasoning with reasoning
Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah Goodman · 2022
Cited alongside, same era.
Automatic chain of thought prompting in large language models, 2022
Zhuosheng Zhang, Aston Zhang, Mu Li, and Alex Smola · 2022
Cited alongside, same era.
Sgd learning on neural networks: leap complexity and saddle-to-saddle dynamics, 2023
Emmanuel Abbe, Enric Boix-Adsera, and Theodor Misiakiewicz · 2023
Cited alongside, same era.
Hidden progress in deep learning: Sgd learns parities near the computational limit, 2023
Boaz Barak, Benjamin L. Edelman, Surbhi Goel, Sham Kakade, Eran Malach, and Cyril Zhang · 2023
Cited alongside, same era.
Selection-inference: Exploiting large language models for interpretable logical reasoning
Antonia Creswell, Murray Shanahan, and Irina Higgins · 2023
Towards revealing the mystery behind chain of thought: a theoretical perspective
Guhao Feng, Bohang Zhang, Yuntian Gu, Haotian Ye, Di He, and Liwei Wang · 2024
Closest in time.
Why are sensitive functions hard for transformers?, 2024
Michael Hahn and Mark Rofin · 2024
Closest in time.
Unveiling the statistical foundations of chain-of-thought prompting methods, 2024
Xinyang Hu, Fengzhuo Zhang, Siyu Chen, and Zhuoran Yang · 2024
Closest in time.
Yiwen Kou, Zixiang Chen, Quanquan Gu, and Sham M. Kakade · 2024
Closest in time.
How do nonlinear transformers acquire generalization-guaranteed cot ability?
Hongkang Li, Meng Wang, Songtao Lu, Xiaodong Cui, and Pin-Yu Chen · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Pareto frontiers in neural feature learning: Data, compute, width, and luck, 2023
Benjamin L. Edelman, Surbhi Goel, Sham Kakade, Eran Malach, and Cyril Zhang · 2023
Cited alongside, same era.
Towards revealing the mystery behind chain of thought: A theoretical perspective
Guhao Feng, Bohang Zhang, Yuntian Gu, Haotian Ye, Di He, and Liwei Wang · 2023
Cited alongside, same era.
Complexity-based prompting for multi-step reasoning
Yao Fu, Hao Peng, Ashish Sabharwal, Peter Clark, and Tushar Khot · 2023
Cited alongside, same era.
Seungone Kim, Se June Joo, Doyoung Kim, Joel Jang, Seonghyeon Ye, Jamin Shin, and Minjoon Seo · 2023
Cited alongside, same era.
Tight time-space lower bounds for constant-pass learning, 2023
Xin Lyu, Avishay Tal, Hongxun Wu, and Junzhao Yang · 2023
Cited alongside, same era.
Why think step by step? reasoning emerges from the locality of experience, 2023
Ben Prystawski, Michael Y. Li, and Noah D. Goodman · 2023
Cited alongside, same era.
Language models are greedy reasoners: A systematic formal analysis of chain-of-thought
Abulhair Saparov and He He · 2023
Cited alongside, same era.
Closest in time.
Learning on transformers is provable low-rank and sparse: A one-layer analysis
Hongkang Li, Meng Wang, Shuai Zhang, Sijia Liu, and Pin-Yu Chen · 2024
Closest in time.
Let’s verify step by step
Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe · 2024
Closest in time.
The expressive power of transformers with chain of thought, 2024
William Merrill and Ashish Sabharwal · 2024
Closest in time.
How transformers learn causal structure with gradient descent, 2024
Eshaan Nichani, Alex Damian, and Jason D. Lee · 2024
Closest in time.
On the representational capacity of neural language models with chain-of-thought reasoning, 2024
Franz Nowak, Anej Svete, Alexandra Butoi, and Ryan Cotterell · 2024
Closest in time.
Why think step by step? reasoning emerges from the locality of experience
Ben Prystawski, Michael Li, and Noah Goodman · 2024
Closest in time.
Introducing qwen2-math
Team Qwen · 2024
Closest in time.
Implicit regularization of gradient flow on one-layer softmax attention, 2024
Heejune Sheen, Siyu Chen, Tianhao Wang, and Harrison H. Zhou · 2024
Closest in time.
Chain-of-thought reasoning without prompting, 2024
Xuezhi Wang and Denny Zhou · 2024
Closest in time.
Rnns are not transformers (yet): The key bottleneck on in-context retrieval, 2024
Kaiyue Wen, Xingyu Dang, and Kaifeng Lyu · 2024
Closest in time.
In-context learning from training on unstructured data: The role of co-occurrence, positional information, and training data structure
Kevin Christian Wibisono and Yixin Wang · 2024
Closest in time.
How in-context learning emerges from training on unstructured data: The role of co-occurrence, positional information, and noise structures
Kevin Christian Wibisono and Yixin Wang · 2024
Closest in time.
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, et al · 2024
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan · 2024
Closest in time.
Metamath: Bootstrap your own mathematical questions for large language models, 2024
Longhui Yu, Weisen Jiang, Han Shi, Jincheng Yu, Zhengying Liu, Yu Zhang, James T. Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu · 2024
Closest in time.