Fetching the paper…
Reading the bibliography…
Transformer-based language models have demonstrated impressive capabilities across a range of complex reasoning tasks.
Communication complexity
Christos H. Papadimitriou and Michael Sipser · 1984
Earlier work this paper cites.
Learning and development in neural networks: The importance of starting small
Jeffrey L Elman · 1993
Earlier work this paper cites.
Rounds in communication complexity revisited
Noam Nisan and Avi Wigderson · 1993
Earlier work this paper cites.
Efficient noise-tolerant learning from statistical queries
Michael Kearns · 1998
Earlier work this paper cites.
Curriculum learning
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston · 2009
Earlier work this paper cites.
Characterizing statistical query learning: simplified notions and proofs
Balázs Szörényi · 2009
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma · 2014
Earlier work this paper cites.
Jason Weston, Sumit Chopra, and Antoine Bordes · 2014
Earlier work this paper cites.
End-to-end memory networks
Sainbayar Sukhbaatar, Jason Weston, and Rob Fergus · 2015
Earlier work this paper cites.
Towards ai-complete question answering: A set of prerequisite toy tasks
Jason Weston, Antoine Bordes, Sumit Chopra, Alexander M Rush, Bart Van Merriënboer, Armand Joulin, and Tomas Mikolov · 2015
Earlier work this paper cites.
Interleaved group products
William Timothy Gowers and Emanuele Viola · 2019
Earlier work this paper cites.
Pointer chasing via triangular discrimination
Amir Yehudayoff · 2020
Earlier work this paper cites.
The staircase property: How hierarchical structure can guide deep learning
Emmanuel Abbe, Enric Boix-Adsera, Matthew S Brennan, Guy Bresler, and Dheeraj Nagaraj · 2021
Earlier work this paper cites.
Graph streaming lower bounds for parameter estimation and property testing via a streaming xor lemma
Sepehr Assadi and Vishvajeet N · 2021
Earlier work this paper cites.
Show your work: scratchpads for intermediate computation with language models
Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, Charles Sutton, and Augustus Odena · 2021
Earlier work this paper cites.
The merged-staircase property: a necessary and nearly sufficient condition for sgd learning of sparse functions on two-layer neural networks
Emmanuel Abbe, Enric Boix Adsera, and Theodor Misiakiewicz · 2022
Earlier work this paper cites.
Data distributional properties drive emergent in-context learning in transformers
Stephanie Chan, Adam Santoro, Andrew Lampinen, Jane Wang, Aaditya Singh, Pierre Richemond, James McClelland, and Felix Hill · 2022
Earlier work this paper cites.
Neural networks can learn representations with gradient descent
Alexandru Damian, Jason Lee, and Mahdi Soltanolkotabi · 2022
Earlier work this paper cites.
Vision transformers provably learn spatial structure
Samy Jelassi, Michael Sander, and Yuanzhi Li · 2022
Earlier work this paper cites.
Solving quantitative reasoning problems with language models
Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra · 2022
Cited alongside, same era.
In-context learning and induction heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2022
Cited alongside, same era.
Solving math word problems with process-and outcome-based feedback
Jonathan Uesato, Nate Kushman, Ramana Kumar, Francis Song, Noah Siegel, Lisa Wang, Antonia Creswell, Geoffrey Irving, and Irina Higgins · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou · 2022
Cited alongside, same era.
Understanding the interplay between parametric and contextual knowledge for large language models
Sitao Cheng, Liangming Pan, Xunjian Yin, Xinyi Wang, and William Yang Wang · 2024
Later among the works it cites.
From explicit cot to implicit cot: Learning to internalize cot step by step
Yuntian Deng, Yejin Choi, and Stuart Shieber · 2024
Later among the works it cites.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Later among the works it cites.
Towards revealing the mystery behind chain of thought: a theoretical perspective
Guhao Feng, Bohang Zhang, Yuntian Gu, Haotian Ye, Di He, and Liwei Wang · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Emmanuel Abbe, Elisabetta Cornacchia, and Aryo Lotfi · 2023
Cited alongside, same era.
Birth of a transformer: A memory viewpoint
Alberto Bietti, Vivien Cabannes, Diane Bouchacourt, Herve Jegou, and Leon Bottou · 2023
Cited alongside, same era.
Faith and fate: Limits of transformers on compositionality (2023)
Nouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li, Liwei Jiang, Bill Yuchen Lin, Peter West, Chandra Bhagavatula, Ronan Le Bras, Jena D. Hwang, Soumya Sanyal, Sean Welleck, Xiang Ren, Allyson Ettinger, Zaid Harchaoui, and Yejin Choi · 2023
Cited alongside, same era.
In-context convergence of transformers
Yu Huang, Yuan Cheng, and Yingbin Liang · 2023
Cited alongside, same era.
Massively parallel computation: Algorithms and applications
Sungjin Im, Ravi Kumar, Silvio Lattanzi, Benjamin Moseley, and Sergei Vassilvitskii · 2023
Cited alongside, same era.
How do transformers learn topic structure: Towards a mechanistic understanding
Yuchen Li, Yuanzhi Li, and Andrej Risteski · 2023
Cited alongside, same era.
Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe · 2023
Cited alongside, same era.
Transformers learn shortcuts to automata
Bingbin Liu, Jordan T. Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang · 2023
Cited alongside, same era.
Cheng Gao, Yuan Cao, Zihao Li, Yihan He, Mengdi Wang, Han Liu, Jason Matthew Klusowski, and Jianqing Fan · 2024
Later among the works it cites.
Better & faster large language models via multi-token prediction
Fabian Gloeckle, Badr Youbi Idrissi, Baptiste Rozière, David Lopez-Paz, and Gabriel Synnaeve · 2024
Later among the works it cites.
Repeat after me: Transformers are better than state space models at copying
Samy Jelassi, David Brandfonbrener, Sham M Kakade, and Eran Malach · 2024
Later among the works it cites.
Transformers provably solve parity efficiently with chain of thought
Juno Kim and Taiji Suzuki · 2024
Later among the works it cites.
Learning to reason and memorize with self-notes
Jack Lanchantin, Shubham Toshniwal, Jason Weston, and Sainbayar Sukhbaatar · 2024
Later among the works it cites.
Chain of thought empowers transformers to solve inherently serial problems
Zhiyuan Li, Hong Liu, Denny Zhou, and Tengyu Ma · 2024
Later among the works it cites.
Progressive distillation induces an implicit curriculum
Abhishek Panigrahi, Bingbin Liu, Sadhika Malladi, Andrej Risteski, and Surbhi Goel · 2024
Later among the works it cites.
On limitations of the transformer architecture
Binghui Peng, Srini Narayanan, and Christos Papadimitriou · 2024
Later among the works it cites.
Learning and transferring sparse contextual bigrams with linear transformers
Yunwei Ren, Zixuan Wang, and Jason D Lee · 2024
Later among the works it cites.
Transformers provably learn sparse token selection while fully-connected nets cannot
Zixuan Wang, Stanley Wei, Daniel Hsu, and Jason D. Lee · 2024
Later among the works it cites.
Kaiyue Wen, Huaqing Zhang, Hongzhou Lin, and Jingzhao Zhang · 2024
Later among the works it cites.
Doremi: Optimizing data mixtures speeds up language model pretraining
Sang Michael Xie, Hieu Pham, Xuanyi Dong, Nan Du, Hanxiao Liu, Yifeng Lu, Percy S Liang, Quoc V Le, Tengyu Ma, and Adams Wei Yu · 2024
Later among the works it cites.
Do large language models latently perform multi-hop reasoning?
Sohee Yang, Elena Gribovskaya, Nora Kassner, Mor Geva, and Sebastian Riedel · 2024
Later among the works it cites.
Tree of thoughts: deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan · 2024
Later among the works it cites.
Transformers learn to implement multi-step gradient descent with chain of thought
Jianhao Huang, Zixuan Wang, and Jason D Lee · 2025
Closest in time.