Fetching the paper…
Reading the bibliography…
Systematic, compositional generalization beyond the training distribution remains a core challenge in machine learning -- and a critical bottleneck for the emergent reasoning abilities of modern language models.
“Syntactic structures”
Noam Chomsky · 1957
Earlier work this paper cites.
“Computer science as empirical inquiry: Symbols and search”
Allen Newel and Herbert Simon · 1976
Earlier work this paper cites.
“Acquisition of cognitive skill.”
John Anderson · 1982
Earlier work this paper cites.
“Connectionism and cognitive architecture: A critical analysis”
Jerry Fodor and Zenon Pylyshyn · 1988
Earlier work this paper cites.
“The transfer of cognitive skill”
Mark Singley and John Anderson · 1989
Earlier work this paper cites.
“Finding structure in time”
Jeffrey Elman · 1990
Earlier work this paper cites.
“Recursive distributed representations”
Jordan Pollack · 1990
Earlier work this paper cites.
“Long short-term memory”
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
“Serial order: A parallel distributed processing approach”
Michael Jordan · 1997
Earlier work this paper cites.
“A model of inductive bias learning”
Jonathan Baxter · 2000
Earlier work this paper cites.
“DeBERTa: Decoding-enhanced BERT with Disentangled Attention”, 2021
Pengcheng He, Xiaodong Liu, Jianfeng Gao and Weizhu Chen · 2006
Earlier work this paper cites.
“Neural-symbolic cognitive reasoning”
Artur’Avila Garcez, Luis Lamb and Dov Gabbay · 2008
Earlier work this paper cites.
“Deep boltzmann machines”
Ruslan Salakhutdinov and Geoffrey Hinton · 2009
Earlier work this paper cites.
“A spike and slab restricted Boltzmann machine”
Aaron Courville, James Bergstra and Yoshua Bengio · 2011
Earlier work this paper cites.
“Semantic compositionality through recursive matrix-vector spaces”
Richard Socher, Brody Huval, Christopher Manning and Andrew Ng · 2012
Earlier work this paper cites.
“Learning Phrase Representations using RNN Encoder–Decoder for Statistical Machine Translation”
Kyunghyun Cho, Bart van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk and Yoshua Bengio · 2014
Earlier work this paper cites.
“Sequence to sequence learning with neural networks”
Ilya Sutskever, Oriol Vinyals and Quoc Le · 2014
Earlier work this paper cites.
“Soft-to-hard vector quantization for end-to-end learning compressible representations”
Eirikur Agustsson, Fabian Mentzer, Michael Tschannen, Lukas Cavigelli, Radu Timofte, Luca Benini and Luc Gool · 2017
Earlier work this paper cites.
“Geometric deep learning: going beyond euclidean data”
Michael Bronstein, Joan Bruna, Yann LeCun, Arthur Szlam and Pierre Vandergheynst · 2017
Earlier work this paper cites.
“Adaptive Computation Time for Recurrent Neural Networks”, 2017
Alex Graves · 2017
Earlier work this paper cites.
“Building machines that learn and think like people”
Brenden Lake, Tomer Ullman, Joshua Tenenbaum and Samuel Gershman · 2017
Earlier work this paper cites.
“Attention is all you need”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan Gomez, Łukasz Kaiser and Illia Polosukhin · 2017
Earlier work this paper cites.
“Measuring abstract reasoning in neural networks”
David Barrett, Felix Hill, Adam Santoro, Ari Morcos and Timothy Lillicrap · 2018
Earlier work this paper cites.
“Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks”
Brenden Lake and Marco Baroni · 2018
Cited alongside, same era.
“Neural Discrete Representation Learning”, 2018
Aaron van Oord, Oriol Vinyals and Koray Kavukcuoglu · 2018
Cited alongside, same era.
“Universal Transformers”, 2019
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit and Łukasz Kaiser · 2019
Cited alongside, same era.
“Compositionality decomposed: How do neural networks generalise?”
Dieuwke Hupkes, Verna Dankers, Mathijs Mul and Elia Bruni · 2020
Cited alongside, same era.
“PonderNet: Learning to Ponder”, 2021
Andrea Banino, Jan Balaguer and Charles Blundell · 2021
Cited alongside, same era.
“Towards Monosemanticity: Decomposing Language Models With Dictionary Learning”
Trenton Bricken et al · 2023
Later among the works it cites.
“Geometric deep learning and equivariant neural networks”
Jan Gerken, Jimmy Aronsson, Oscar Carlsson, Hampus Linander, Fredrik Ohlsson, Christoffer Petersson and Daniel Persson · 2023
Later among the works it cites.
“Length Generalization in Arithmetic Transformers”, 2023
Samy Jelassi, Stéphane d’Ascoli, Carles Domingo-Enrich, Yuhuai Wu, Yuanzhi Li and François Charton · 2023
Later among the works it cites.
“The Impact of Positional Encoding on Length Generalization in Transformers”
Amirhossein Kazemnejad, Inkit Padhi, Karthikeyan Natesan, Payel Das and Siva Reddy · 2023
Later among the works it cites.
“Let’s Verify Step by Step”, 2023
Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever and Karl Cobbe · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Michael. Bronstein, Joan Bruna, Taco Cohen and Petar Veličković · 2021
Cited alongside, same era.
“Training Verifiers to Solve Math Word Problems” arXiv:2110.14168 [cs]
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Łukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse and John Schulman · 2021
Cited alongside, same era.
“A Mathematical Framework for Transformer Circuits”
Nelson Elhage et al · 2021
Cited alongside, same era.
“Causal Abstractions of Neural Networks”
Atticus Geiger, Hanson Lu, Thomas Icard and Christopher Potts · 2021
Cited alongside, same era.
“Show Your Work: Scratchpads for Intermediate Computation with Language Models”, 2021
Maxwell Nye, Anders Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, Charles Sutton and Augustus Odena · 2021
Cited alongside, same era.
“Can You Learn an Algorithm? Generalizing from Easy to Hard Problems with Recurrent Networks”, 2021
Avi Schwarzschild, Eitan Borgnia, Arjun Gupta, Furong Huang, Uzi Vishkin, Micah Goldblum and Tom Goldstein · 2021
Cited alongside, same era.
“Neural algorithmic reasoning”
Petar Veličković and Charles Blundell · 2021
Cited alongside, same era.
Later among the works it cites.
“LogiCoT: Logical Chain-of-Thought Instruction Tuning”
Hanmeng Liu, Zhiyang Teng, Leyang Cui, Chaoli Zhang, Qiji Zhou and Yue Zhang · 2023
Later among the works it cites.
“Progress measures for grokking via mechanistic interpretability”
Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith and Jacob Steinhardt · 2023
Later among the works it cites.
“RoFormer: Enhanced Transformer with Rotary Position Embedding”, 2023
Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen and Yunfeng Liu · 2023
Later among the works it cites.
“Scaling instruction-finetuned language models”
Hyung Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani and Siddhartha Brahma · 2024
Later among the works it cites.
“To grok or not to grok: Disentangling generalization and memorization on corrupted algorithmic datasets”
Darshil Doshi, Aritra Das, Tianyu He and Andrey Gromov · 2024
Later among the works it cites.
“Looped Transformers for Length Generalization”, 2024
Ying Fan, Yilun Du, Kannan Ramchandran and Kangwook Lee · 2024
Later among the works it cites.
“Finding alignments between interpretable causal variables and distributed neural representations”
Atticus Geiger, Zhengxuan Wu, Christopher Potts, Thomas Icard and Noah Goodman · 2024
Later among the works it cites.
“DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models”, 2024
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y.. Li, Y. Wu and Daya Guo · 2024
Later among the works it cites.
“Mechanisms of Symbol Processing for In-Context Learning in Transformer Networks”, 2024
Paul Smolensky, Roland Fernandez, Zhenghao Zhou, Mattia Opper and Jianfeng Gao · 2024
Later among the works it cites.
“Chain of thoughtlessness? an analysis of cot in planning”
Kaya Stechly, Karthik Valmeekam and Subbarao Kambhampati · 2024
Later among the works it cites.
“Composing Global Optimizers to Reasoning Tasks via Algebraic Objects in Neural Nets”, 2024
Yuandong Tian · 2024
Later among the works it cites.
“Looped Transformers Are Better at Learning Learning Algorithms”, 2024
Liu Yang, Kangwook Lee, Robert Nowak and Dimitris Papailiopoulos · 2024
Later among the works it cites.
“Physics of Language Models: Part 2.1, Grade-School Math and the Hidden Reasoning Process”, 2024
Tian Ye, Zicheng Xu, Yuanzhi Li and Zeyuan Allen-Zhu · 2024
Later among the works it cites.
“What Algorithms can Transformers Learn? A Study in Length Generalization”
Hattie Zhou, Arwen Bradley, Etai Littwin, Noam Razin, Omid Saremi, Joshua. Susskind, Samy Bengio and Preetum Nakkiran · 2024
Later among the works it cites.
“Transformers Can Achieve Length Generalization But Not Robustly”, 2024
Yongchao Zhou, Uri Alon, Xinyun Chen, Xuezhi Wang, Rishabh Agarwal and Denny Zhou · 2024
Later among the works it cites.
“Circuit Tracing: Revealing Computational Graphs in Language Models”
Emmanuel Ameisen et al · 2025
Closest in time.
“Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach”, 2025
Jonas Geiping, Sean McLeish, Neel Jain, John Kirchenbauer, Siddharth Singh, Brian. Bartoldson, Bhavya Kailkhura, Abhinav Bhatele and Tom Goldstein · 2025
Closest in time.