Fetching the paper…
Reading the bibliography…
We prove rich algebraic structures of the solution space for 2-layer neural networks with quadratic activation and $L_2$ loss, trained on reasoning tasks in Abelian group (e.g., modular addition).
Solutions to some functional equations and their applications to characterization of probability distributions
CG Khatri and C Radhakrishna Rao · 1968
Earlier work this paper cites.
Group representations in probability and statistics
Persi Diaconis · 1988
Earlier work this paper cites.
Representation theory of finite groups
Benjamin Steinberg · 2009
Earlier work this paper cites.
Characters of finite abelian groups
Keith Conrad · 2010
Earlier work this paper cites.
Representation theory: a first course , volume 129
William Fulton and Joe Harris · 2013
Earlier work this paper cites.
Lifelong learning with dynamically expandable networks
Jaehong Yoon, Eunho Yang, Jeongtae Lee, and Sung Ju Hwang · 2017
Earlier work this paper cites.
On the power of over-parametrization in neural networks with quadratic activation
Simon Du and Jason Lee · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton · 2018
Earlier work this paper cites.
Jiahui Yu, Linjie Yang, Ning Xu, Jianchao Yang, and Thomas Huang · 2018
Earlier work this paper cites.
Splitting steepest descent for growing neural architectures
Lemeng Wu, Dilin Wang, and Qiang Liu · 2019
Earlier work this paper cites.
Glu variants improve transformer
Noam Shazeer · 2020
Earlier work this paper cites.
Geometric deep learning: Grids, groups, graphs, geodesics, and gauges
Michael M Bronstein, Joan Bruna, Taco Cohen, and Petar Veličković · 2021
Earlier work this paper cites.
Primer: Searching for efficient transformers for language modeling
David R. So, Wojciech Manke, Hanxiao Liu, Zihang Dai, Noam Shazeer, and Quoc V. Le · 2021
Earlier work this paper cites.
Transformers learn shortcuts to automata
Bingbin Liu, Jordan T Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang · 2022
Earlier work this paper cites.
Grokking: Generalization beyond overfitting on small algorithmic datasets
Alethea Power, Yuri Burda, Harri Edwards, Igor Babuschkin, and Vedant Misra · 2022
Cited alongside, same era.
Emergent abilities of large language models
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al · 2022
Cited alongside, same era.
Backward feature correction: How deep learning performs deep (hierarchical) learning
Zeyuan Allen-Zhu and Yuanzhi Li · 2023
Cited alongside, same era.
The reversal curse: Llms trained on” a is b” fail to learn” b is a”
Lukas Berglund, Meg Tong, Max Kaufmann, Mikita Balesni, Asa Cooper Stickland, Tomasz Korbak, and Owain Evans · 2023
Cited alongside, same era.
Faith and fate: Limits of transformers on compositionality (2023)
Learning and leveraging world models in visual representation learning
Quentin Garrido, Mahmoud Assran, Nicolas Ballas, Adrien Bardes, Laurent Najman, and Yann LeCun · 2024
Closest in time.
Yufei Huang, Shengding Hu, Xu Han, Zhiyuan Liu, and Maosong Sun · 2024
Closest in time.
Emergent representations of program semantics in language models trained on programs, 2024
Charles Jin and Martin Rinard · 2024
Closest in time.
Llms can’t plan, but can help planning in llm-modulo frameworks, 2024
Subbarao Kambhampati, Karthik Valmeekam, Lin Guan, Mudit Verma, Kaya Stechly, Siddhant Bhambri, Lucas Saldyt, and Anil Murthy · 2024
Closest in time.
Chain of thought empowers transformers to solve inherently serial problems
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li, Liwei Jiang, Bill Yuchen Lin, Peter West, Chandra Bhagavatula, Ronan Le Bras, Jena D Hwang, et al · 2023
Cited alongside, same era.
Andrey Gromov · 2023
Cited alongside, same era.
Large language models cannot self-correct reasoning yet
Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, and Denny Zhou · 2023
Cited alongside, same era.
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed · 2023
Cited alongside, same era.
Feature emergence via margin maximization: case studies in algebraic tasks
Depen Morwani, Benjamin L Edelman, Costin-Andrei Oncescu, Rosie Zhao, and Sham Kakade · 2023
Cited alongside, same era.
Progress measures for grokking via mechanistic interpretability
Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt · 2023
Cited alongside, same era.
Counting and algorithmic generalization with transformers
Simon Ouellette, Rolf Pfister, and Hansueli Jud · 2023
Cited alongside, same era.
Explaining grokking through circuit efficiency
Vikrant Varma, Rohin Shah, Zachary Kenton, János Kramár, and Ramana Kumar · 2023
Cited alongside, same era.
Zhiyuan Li, Hong Liu, Denny Zhou, and Tengyu Ma · 2024
Closest in time.
Marianna Nezhurina, Lucia Cipolina-Kun, Mehdi Cherti, and Jenia Jitsev · 2024
Closest in time.
OpenAI · 2024
Closest in time.
Travelplanner: A benchmark for real-world planning with language agents, 2024
Jian Xie, Kai Zhang, Jiangjie Chen, Tinghui Zhu, Renze Lou, Yuandong Tian, Yanghua Xiao, and Yu Su · 2024
Closest in time.
Physics of language models: Part 2.1, grade-school math and the hidden reasoning process
Tian Ye, Zicheng Xu, Yuanzhi Li, and Zeyuan Allen-Zhu · 2024
Closest in time.
When can transformers count to n?
Gilad Yehudai, Haim Kaplan, Asma Ghandeharioun, Mor Geva, and Amir Globerson · 2024
Closest in time.
Relu 2 wins: Discovering efficient activation functions for sparse llms, 2024
Zhengyan Zhang, Yixin Song, Guanghui Yu, Xu Han, Yankai Lin, Chaojun Xiao, Chenyang Song, Zhiyuan Liu, Zeyu Mi, and Maosong Sun · 2024
Closest in time.
The clock and the pizza: Two stories in mechanistic explanation of neural networks
Ziqian Zhong, Ziming Liu, Max Tegmark, and Jacob Andreas · 2024
Closest in time.
Pre-trained large language models use fourier features to compute addition
Tianyi Zhou, Deqing Fu, Vatsal Sharan, and Robin Jia · 2024
Closest in time.