Fetching the paper…
Reading the bibliography…
Transformers have demonstrated impressive capabilities across various tasks, yet their performance on compositional problems remains a subject of debate.
1906
Earlier work this paper cites.
J. A. Fodor and Z. W. Pylyshyn, “Connectionism and cognitive architecture: A critical analysis,” Cognition , vol. 28, no. 1-2, pp. 3–71, 1988
1988
Earlier work this paper cites.
G. F. Marcus, The algebraic mind: Integrating connectionism and cognitive science . MIT press, 2003
2003
Earlier work this paper cites.
M. Rudelson and R. Vershynin, “Sampling from large matrices: An approach through geometric functional analysis,” Journal of the ACM (JACM) , vol. 54, no. 4, pp. 21–es, 2007
2007
Earlier work this paper cites.
J. A. Tropp et al. , “An introduction to matrix concentration inequalities,” Foundations and Trends® in Machine Learning , vol. 8, no. 1-2, pp. 1–230, 2015
2015
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in neural information processing systems , 2017, pp. 5998–6008
2017
Earlier work this paper cites.
D. Arpit, S. Jastrzbski, N. Ballas, D. Krueger, E. Bengio, M. S. Kanwal, T. Maharaj, A. Fischer, A. Courville, Y. Bengio et al. , “A closer look at memorization in deep networks,” in Proceedings of the 34th International Conference on Machine Learning-Volume 70 , 2017, pp. 233–242
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
L. Chizat and F. Bach, “On the Global Convergence of Gradient Descent for Over-parameterized Models using Optimal Transport,” in Advances in Neural Information Processing Systems 31 , 2018, pp. 3036–3046
2018
Earlier work this paper cites.
A. Jacot, F. Gabriel, and C. Hongler, “Neural Tangent Kernel: Convergence and Generalization in Neural Networks,” in Advances in Neural Information Processing Systems 31 , 2018, pp. 8571–8580
2018
Earlier work this paper cites.
S. Mei, A. Montanari, and P.-M. Nguyen, “A mean field view of the landscape of two-layer neural networks,” Proceedings of the National Academy of Sciences , vol. 115, no. 33, pp. E7665–E7671, 2018
2018
Earlier work this paper cites.
G. Rotskoff and E. Vanden-Eijnden, “Parameters as interacting particles: long time convergence and asymptotic error scaling of neural networks,” in Advances in Neural Information Processing Systems 31 , 2018, pp. 7146–7155
2018
Earlier work this paper cites.
R. Vershynin, High-dimensional probability: An introduction with applications in data science . Cambridge university press, 2018, vol. 47
2018
Earlier work this paper cites.
B. Lake and M. Baroni, “Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks,” in International conference on machine learning . PMLR, 2018, pp. 2873–2882
2018
Earlier work this paper cites.
D. Keysers, N. Schärli, N. Scales, H. Buisman, D. Furrer, S. Kashubin, N. Momchev, D. Sinopalnikov, L. Stafiniak, T. Tihon et al. , “Measuring compositional generalization: A comprehensive method on realistic data,” in International Conference on Learning Representations , 2019
2019
Earlier work this paper cites.
S. Arora, S. S. Du, W. Hu, Z. Li, R. R. Salakhutdinov, and R. Wang, “On exact computation with an infinitely wide neural net,” in Advances in Neural Information Processing Systems , 2019, pp. 8141–8150
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Earlier work this paper cites.
N. Rahaman, D. Arpit, A. Baratin, F. Draxler, M. Lin, F. A. Hamprecht, Y. Bengio, and A. Courville, “On the spectral bias of deep neural networks,” International Conference on Machine Learning , 2019
2019
Earlier work this paper cites.
Z.-Q. J. Xu, Y. Zhang, and Y. Xiao, “Training Behavior of Deep Neural Network in Frequency Domain,” in Neural Information Processing , ser. Lecture Notes in Computer Science, 2019, pp. 264–274
2019
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Earlier work this paper cites.
L. Yu and A. Ettinger, “Assessing phrasal representation and composition in transformers,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , 2020, pp. 4896–4907
2020
Earlier work this paper cites.
D. Hupkes, V. Dankers, M. Mul, and E. Bruni, “Compositionality decomposed: How do neural networks generalise?” Journal of Artificial Intelligence Research , vol. 67, pp. 757–795, 2020
2020
Earlier work this paper cites.
N. Kim and T. Linzen, “Cogs: A compositional generalization challenge based on semantic interpretation,” in Proceedings of the 2020 conference on empirical methods in natural language processing (emnlp) , 2020, pp. 9087–9105
2020
Earlier work this paper cites.
W. E, C. Ma, and L. Wu, “A comparative analysis of optimization and generalization properties of two-layer neural network and random feature models under gradient descent dynamics.” Sci. China Math. , vol. 63, 2020
2020
Earlier work this paper cites.
J. Sirignano and K. Spiliopoulos, “Mean field analysis of neural networks: A central limit theorem,” Stochastic Processes and their Applications , vol. 130, no. 3, pp. 1820–1852, 2020
2020
Earlier work this paper cites.
X. S. Huang, F. Perez, J. Ba, and M. Volkovs, “Improving transformer optimization through better initialization,” in International Conference on Machine Learning . PMLR, 2020, pp. 4475–4483
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
Z.-Q. J. Xu, Y. Zhang, T. Luo, Y. Xiao, and Z. Ma, “Frequency principle: Fourier analysis sheds light on deep neural networks,” Communications in Computational Physics , vol. 28, no. 5, pp. 1746–1767, 2020
2020
Cited alongside, same era.
J. H. Choi, K. E. Hickman, A. B. Monahan, and D. Schwarcz, “Chatgpt goes to law school,” J. Legal Educ. , vol. 71, p. 387, 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
T. Luo, Z.-Q. J. Xu, Z. Ma, and Y. Zhang, “Phase diagram for two-layer relu neural networks at infinite-width limit,” Journal of Machine Learning Research , vol. 22, no. 71, pp. 1–47, 2021
2021
Cited alongside, same era.
2023
Later among the works it cites.
O. Press, M. Zhang, S. Min, L. Schmidt, N. A. Smith, and M. Lewis, “Measuring and narrowing the compositionality gap in language models,” in Findings of the Association for Computational Linguistics: EMNLP 2023 , 2023, pp. 5687–5711
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. Zhu, R. Ni, Z. Xu, K. Kong, W. R. Huang, and T. Goldstein, “Gradinit: Learning to initialize neural networks for stable and efficient training,” Advances in Neural Information Processing Systems , vol. 34, pp. 16 410–16 422, 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
Y. Zhang, Z. Zhang, T. Luo, and Z. J. Xu, “Embedding principle of loss landscape of deep neural networks,” Advances in Neural Information Processing Systems , vol. 34, pp. 14 848–14 859, 2021
2021
Cited alongside, same era.
P. Smolensky, R. McCoy, R. Fernandez, M. Goldrick, and J. Gao, “Neurocompositional computing: From the central paradox of cognition to a new generation of ai systems,” AI Magazine , vol. 43, no. 3, pp. 308–322, 2022
2022
Cited alongside, same era.
Y. Fu, H. Peng, and T. Khot, “How does gpt obtain its ability? tracing emergent abilities of language models to their sources,” Yao Fu’s Notion , 2022
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
E. Leivada, E. Murphy, and G. Marcus, “Dall· e 2 fails to reliably capture common syntactic processes,” Social Sciences & Humanities Open , vol. 8, no. 1, p. 100648, 2023
2023
Later among the works it cites.
Y. Du, C. Durkan, R. Strudel, J. B. Tenenbaum, S. Dieleman, R. Fergus, J. Sohl-Dickstein, A. Doucet, and W. S. Grathwohl, “Reduce, reuse, recycle: Compositional generation with energy-based diffusion models and mcmc,” in International conference on machine learning . PMLR, 2023, pp. 8489–8510
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
A. Trockman and J. Z. Kolter, “Mimetic initialization of self-attention layers,” in International Conference on Machine Learning . PMLR, 2023, pp. 34 456–34 468
2023
Later among the works it cites.
D. Soboleva, F. Al-Khateeb, R. Myers, J. R. Steeves, J. Hestness, and N. Dey, “SlimPajama: A 627B token cleaned and deduplicated version of RedPajama,” https://www.cerebras.net/blog/slimpajama-a-627b-token-cleaned-and-deduplicated-version-of-redpajama
2023
Later among the works it cites.
A. Saparov and H. He, “Language models are greedy reasoners: A systematic formal analysis of chain-of-thought,” in The Eleventh International Conference on Learning Representations , 2023
2023
Later among the works it cites.
2024
Later among the works it cites.
Z. Zhang, P. Lin, Z. Wang, Y. Zhang, and Z.-Q. J. Xu, “Initialization is critical to whether transformers fit composite functions by reasoning or memorizing,” in The Thirty-eighth Annual Conference on Neural Information Processing Systems , 2024
2024
Later among the works it cites.
N. Dziri, X. Lu, M. Sclar, X. L. Li, L. Jiang, B. Y. Lin, S. Welleck, P. West, C. Bhagavatula, R. Le Bras et al. , “Faith and fate: Limits of transformers on compositionality,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
M. Okawa, E. S. Lubana, R. Dick, and H. Tanaka, “Compositional abilities emerge multiplicatively: Exploring diffusion models on a synthetic task,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
Z. Zhang and Z.-Q. J. Xu, “Implicit regularization of dropout,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2024
2024
Later among the works it cites.
P. Gopalani, E. S. Lubana, and W. Hu, “How do transformers fill in the blanks? a case study on matrix completion,” in ICML 2024 Workshop on Mechanistic Interpretability , 2024
2024
Later among the works it cites.
H. Wang, S. Ma, L. Dong, S. Huang, D. Zhang, and F. Wei, “Deepnet: Scaling transformers to 1,000 layers,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2024
2024
Later among the works it cites.
Y. Hu, X. Tang, H. Yang, and M. Zhang, “Case-based or rule-based: How do transformers do the math?” in Forty-first International Conference on Machine Learning , 2024
2024
Later among the works it cites.