Fetching the paper…
Reading the bibliography…
Large language models have consistently struggled with complex reasoning tasks, such as mathematical problem-solving.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, and 1 others. 2020 · 1901
Earlier work this paper cites.
Frequency principle: Fourier analysis sheds light on deep neural networks
Zhi-Qin John Xu, Yaoyu Zhang, Tao Luo, Yanyang Xiao, and Zheng Ma. 2019 · 1901
Earlier work this paper cites.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019 · 1905
Earlier work this paper cites.
A multiscale visualization of attention in the transformer model
Jesse Vig. 2019 · 1906
Earlier work this paper cites.
Revealing the dark secrets of bert
Olga Kovaleva, Alexey Romanov, Anna Rogers, and Anna Rumshisky. 2019 · 1908
Earlier work this paper cites.
Clutrr: A diagnostic benchmark for inductive reasoning from text
Koustuv Sinha, Shagun Sodhani, Jin Dong, Joelle Pineau, and William L Hamilton. 2019 · 1908
Earlier work this paper cites.
Transformers as soft reasoners over language
Peter Clark, Oyvind Tafjord, and Kyle Richardson. 2020 · 2002
Earlier work this paper cites.
Attention is not only a weight: Analyzing transformers with vector norms
Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, and Kentaro Inui. 2020 · 2004
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
A convergence analysis of gradient descent for deep linear neural networks
Sanjeev Arora, Nadav Cohen, Noah Golowich, and Wei Hu. 2018 · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler. 2018 · 2018
Earlier work this paper cites.
Generating wikipedia by summarizing long sequences
Peter J Liu, Mohammad Saleh, Etienne Pot, Ben Goodrich, Ryan Sepassi, Lukasz Kaiser, and Noam Shazeer. 2018 · 2018
Earlier work this paper cites.
How sgd selects the global minima in over-parameterized learning: A dynamical stability perspective
Lei Wu, Chao Ma, and 1 others. 2018 · 2018
Earlier work this paper cites.
Zhanxing Zhu, Jingfeng Wu, Bing Yu, Lei Wu, and Jinwen Ma. 2018 · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, and 1 others. 2019 · 2019
Earlier work this paper cites.
Implicit regularization of random feature models
Arthur Jacot, Berfin Simsek, Francesco Spadaro, Clément Hongler, and Franck Gabriel. 2020 · 2020
Earlier work this paper cites.
Investigating gender bias in language models using causal mediation analysis
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart Shieber. 2020 · 2020
Earlier work this paper cites.
Advancing mathematics by guiding human intuition with ai
Alex Davies, Petar Veličković, Lars Buesing, Sam Blackwell, Daniel Zheng, Nenad Tomašev, Richard Tanburn, Peter Battaglia, Charles Blundell, András Juhász, and 1 others. 2021 · 2021
Earlier work this paper cites.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, and 6 others. 2021 · 2021
Earlier work this paper cites.
On the validity of modeling sgd with stochastic differential equations (sdes)
Zhiyuan Li, Sadhika Malladi, and Sanjeev Arora. 2021 · 2021
Earlier work this paper cites.
Phase diagram for two-layer relu neural networks at infinite-width limit
Tao Luo, Zhi-Qin John Xu, Zheng Ma, and Yaoyu Zhang. 2021 · 2021
Earlier work this paper cites.
Transformers can do bayesian inference
Samuel Müller, Noah Hollmann, Sebastian Pineda Arango, Josif Grabocka, and Frank Hutter. 2021 · 2021
Earlier work this paper cites.
Chiyuan Zhang, Maithra Raghu, Jon Kleinberg, and Samy Bengio. 2021 · 2021
Earlier work this paper cites.
Learning to reason with neural networks: Generalization, unseen data and boolean measures
Emmanuel Abbe, Samy Bengio, Elisabetta Cornacchia, Jon Kleinberg, Aryo Lotfi, Maithra Raghu, and Chiyuan Zhang. 2022 · 2022
Earlier work this paper cites.
Understanding gradient descent on the edge of stability in deep learning
Sanjeev Arora, Zhiyuan Li, and Abhishek Panigrahi. 2022 · 2022
Earlier work this paper cites.
Qiming Bao, Alex Yuxuan Peng, Tim Hartill, Neset Tan, Zhenyun Deng, Michael Witbrock, and Jiamou Liu. 2022 · 2022
Earlier work this paper cites.
Analyzing transformers in embedding space
Guy Dar, Mor Geva, Ankit Gupta, and Jonathan Berant. 2022 · 2022
Cited alongside, same era.
A survey on in-context learning
Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Zhiyong Wu, Baobao Chang, Xu Sun, Jingjing Xu, and Zhifang Sui. 2022 · 2022
Cited alongside, same era.
What can transformers learn in-context? a case study of simple function classes
Shivam Garg, Dimitris Tsipras, Percy S Liang, and Gregory Valiant. 2022 · 2022
Cited alongside, same era.
What changed? investigating debiasing methods using causal mediation analysis
Sullam Jeoung and Jana Diesner. 2022 · 2022
Cited alongside, same era.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022 · 2022
The implicit regularization of dynamical stability in stochastic gradient descent
Lei Wu and Weijie J Su. 2023 · 2023
Later among the works it cites.
Beam retrieval: General end-to-end retrieval for multi-hop question answering
Jiahao Zhang, Haiyang Zhang, Dongmei Zhang, Yong Liu, and Shen Huang. 2023 · 2023
Later among the works it cites.
How far can transformers reason? the locality barrier and inductive scratchpad
Emmanuel Abbe, Samy Bengio, Aryo Lotfi, Colin Sandon, and Omid Saremi. 2024 · 2024
Closest in time.
Phi-3 technical report: A highly capable language model locally on your phone
Marah Abdin, Sam Ade Jacobs, Ammar Ahmad Awan, Jyoti Aneja, Ahmed Awadallah, Hany Awadalla, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Harkirat Behl, and 1 others. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Locating and editing factual associations in gpt
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022 · 2022
Cited alongside, same era.
In-context learning and induction heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, and 1 others. 2022 · 2022
Cited alongside, same era.
Logicinference: A new dataset for teaching logical inference to seq2seq models
Santiago Ontanon, Joshua Ainslie, Vaclav Cvicek, and Zachary Fisher. 2022 · 2022
Cited alongside, same era.
Language models are greedy reasoners: A systematic formal analysis of chain-of-thought
Abulhair Saparov and He He. 2022 · 2022
Cited alongside, same era.
Stepgame: A new benchmark for robust multi-hop spatial reasoning in texts
Zhengxiang Shi, Qiang Zhang, and Aldo Lipani. 2022 · 2022
Cited alongside, same era.
Interpretability in the wild: a circuit for indirect object identification in gpt-2 small
Kevin Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt. 2022 · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, and 1 others. 2022 · 2022
Cited alongside, same era.
Noah Amsel, Gilad Yehudai, and Joan Bruna. 2024 · 2024
Closest in time.
Transformer block coupling and its correlation with generalization in llms
Murdock Aubry, Haoming Meng, Anton Sugolov, and Vardan Papyan. 2024 · 2024
Closest in time.
Birth of a transformer: A memory viewpoint
Alberto Bietti, Vivien Cabannes, Diane Bouchacourt, Herve Jegou, and Leon Bottou. 2024 · 2024
Closest in time.
Hopping too late: Exploring the limitations of large language models on multi-hop queries
Eden Biran, Daniela Gottesman, Sohee Yang, Mor Geva, and Amir Globerson. 2024 · 2024
Closest in time.
A mechanistic analysis of a transformer trained on a symbolic multi-step reasoning task
Jannik Brinkmann, Abhay Sheshadri, Victor Levoso, Paul Swoboda, and Christian Bartelt. 2024 · 2024
Closest in time.
What can transformer learn with varying depth? case studies on sequence learning tasks
Xingwu Chen and Difan Zou. 2024 · 2024
Closest in time.
How to think step-by-step: A mechanistic understanding of chain-of-thought reasoning
Subhabrata Dutta, Joykirat Singh, Soumen Chakrabarti, and Tanmoy Chakraborty. 2024 · 2024
Closest in time.
The evolution of statistical induction heads: In-context learning markov chains
Benjamin L Edelman, Ezra Edelman, Surbhi Goel, Eran Malach, and Nikolaos Tsilivis. 2024 · 2024
Closest in time.
Deepseek-coder: When the large language model meets programming–the rise of code intelligence
Daya Guo, Qihao Zhu, Dejian Yang, Zhenda Xie, Kai Dong, Wentao Zhang, Guanting Chen, Xiao Bi, Yu Wu, YK Li, and 1 others. 2024 · 2024
Closest in time.
A peek into token bias: Large language models are not yet genuine reasoners
Bowen Jiang, Yangxinyu Xie, Zhuoqun Hao, Xiaomeng Wang, Tanwi Mallick, Weijie J Su, Camillo J Taylor, and Dan Roth. 2024 · 2024
Closest in time.
Ii-mmr: Identifying and improving multi-modal multi-hop reasoning in visual question answering
Jihyung Kil, Farideh Tavazoee, Dongyeop Kang, and Joo-Kyung Kim. 2024 · 2024
Closest in time.
Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, and 1 others. 2024 · 2024
Closest in time.
How transformers learn causal structure with gradient descent
Eshaan Nichani, Alex Damian, and Jason D Lee. 2024 · 2024
Closest in time.
Mechanistic design and scaling of hybrid architectures
Michael Poli, Armin W Thomas, Eric Nguyen, Pragaash Ponnusamy, Björn Deiseroth, Kristian Kersting, Taiji Suzuki, Brian Hie, Stefano Ermon, Christopher Ré, and 1 others. 2024 · 2024
Closest in time.
Understanding the generalization benefits of late learning rate decay
Yinuo Ren, Chao Ma, and Lexing Ying. 2024 · 2024
Closest in time.
Distributional reasoning in llms: Parallel reasoning processes in multi-hop reasoning
Yuval Shalev, Amir Feder, and Ariel Goldstein. 2024 · 2024
Closest in time.
Solving olympiad geometry without human demonstrations
Trieu H Trinh, Yuhuai Wu, Quoc V Le, He He, and Thang Luong. 2024 · 2024
Closest in time.
Logicasker: Evaluating and improving the logical reasoning ability of large language models
Yuxuan Wan, Wenxuan Wang, Yiliu Yang, Youliang Yuan, Jen-tse Huang, Pinjia He, Wenxiang Jiao, and Michael R Lyu. 2024 · 2024
Closest in time.
Understanding the expressive power and mechanisms of transformer for sequence modeling
Mingze Wang and E Weinan. 2024 · 2024
Closest in time.
Do large language models latently perform multi-hop reasoning?
Sohee Yang, Elena Gribovskaya, Nora Kassner, Mor Geva, and Sebastian Riedel. 2024 · 2024
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan. 2024 · 2024
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, and 1 others. 2025 · 2025
Closest in time.