Fetching the paper…
Reading the bibliography…
Many recent studies have found evidence for emergent reasoning capabilities in large language models (LLMs), but debate persists concerning the robustness of these capabilities, and the extent to which they depend on structured reasoning mechanisms.
Reasoning about a rule
Wason, P. C · 1968
Earlier work this paper cites.
Connectionism and cognitive architecture: A critical analysis
Fodor, J. A. and Pylyshyn, Z. W · 1988
Earlier work this paper cites.
Tensor product variable binding and the representation of symbolic structures in connectionist systems
Smolensky, P · 1990
Earlier work this paper cites.
The copycat project: A model of mental fluidity and analogy-making
Hofstadter, D. R., Mitchell, M., et al · 1995
Earlier work this paper cites.
Rule learning by seven-month-old infants
Marcus, G. F., Vijayan, S., Bandi Rao, S., and Vishton, P. M · 1999
Earlier work this paper cites.
The algebraic mind: Integrating connectionism and cognitive science
Marcus, G. F · 2001
Earlier work this paper cites.
A symbolic-connectionist theory of relational inference and generalization
Hummel, J. E. and Holyoak, K. J · 2003
Earlier work this paper cites.
Representational similarity analysis-connecting the branches of systems neuroscience
Kriegeskorte, N., Mur, M., and Bandettini, P. A · 2008
Earlier work this paper cites.
Indirection and symbol-like processing in the prefrontal cortex and basal ganglia
Kriete, T., Noelle, D. C., Cohen, J. D., and O’Reilly, R. C · 2013
Earlier work this paper cites.
Hybrid computing using a neural network with dynamic external memory
Graves, A., Wayne, G., Reynolds, M., Harley, T., Danihelka, I., Grabska-Barwińska, A., Colmenarejo, S. G., Grefenstette, E., Ramalho, T., Agapiou, J., et al · 2016
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., et al · 2017
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
On the dangers of stochastic parrots: Can language models be too big?
Bender, E. M., Gebru, T., McMillan-Major, A., and Shmitchell, S · 2021
Earlier work this paper cites.
A mathematical framework for transformer circuits
Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., DasSarma, N., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2021
Earlier work this paper cites.
Emergent symbols through binding in external memory
Webb, T. W., Sinha, I., and Cohen, J · 2021
Earlier work this paper cites.
Symbols and mental programs: a hypothesis about human singularity
Dehaene, S., Al Roumi, F., Lakretz, Y., Planton, S., and Sablé-Meyer, M · 2022
Earlier work this paper cites.
On neural architecture inductive biases for relational tasks
Kerg, G., Mittal, S., Rolnick, D., Bengio, Y., Richards, B., and Lajoie, G · 2022
Earlier work this paper cites.
Locating and editing factual associations in gpt
Meng, K., Bau, D., Andonian, A., and Belinkov, Y · 2022
Cited alongside, same era.
In-context learning and induction heads
Olsson, C., Elhage, N., Nanda, N., Joseph, N., DasSarma, N., Henighan, T., Mann, B., Askell, A., Bai, Y., Chen, A., et al · 2022
Cited alongside, same era.
Direct and indirect effects
Pearl, J · 2022
Cited alongside, same era.
Emergent abilities of large language models
Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., et al · 2022
Cited alongside, same era.
McCoy, R. T., Yao, S., Friedman, D., Hardy, M., and Griffiths, T. L · 2023
Cited alongside, same era.
Gemma 2: Improving open language models at a practical size
Gemma Team · 2024
Later among the works it cites.
Language models, like humans, show content effects on reasoning tasks
Lampinen, A. K., Dasgupta, I., Chan, S. C., Sheahan, H. R., Creswell, A., Kumaran, D., McClelland, J. L., and Hill, F · 2024
Later among the works it cites.
Evaluating the robustness of analogical reasoning in large language models
Lewis, M. and Mitchell, M · 2024
Later among the works it cites.
Evaluating cognitive maps and planning in large language models with cogeval
Momennejad, I., Hasanbeig, H., Vieira Frujeri, F., Sharma, H., Jojic, N., Palangi, H., Ness, R., and Larson, J · 2024
Later among the works it cites.
Semantic structure-mapping in llm and human analogical reasoning
Musker, S., Duchnowski, A., Millière, R., and Pavlick, E · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A mechanism for solving relational tasks in transformer language models
Merullo, J., Eickhoff, C., and Pavlick, E · 2023
Cited alongside, same era.
Large language models as general pattern machines
Mirchandani, S., Xia, F., Florence, P., Ichter, B., Driess, D., Arenas, M. G., Rao, K., Sadigh, D., and Zeng, A · 2023
Cited alongside, same era.
Interpretability in the wild: a circuit for indirect object identification in gpt-2 small
Wang, K. R., Variengien, A., Conmy, A., Shlegeris, B., and Steinhardt, J · 2023
Cited alongside, same era.
Emergent analogical reasoning in large language models
Webb, T., Holyoak, K. J., and Lu, H · 2023
Cited alongside, same era.
Wong, L., Grand, G., Lew, A. K., Goodman, N. D., Mansinghka, V. K., Andreas, J., and Tenenbaum, J. B · 2023
Cited alongside, same era.
In-context language learning: architectures and algorithms
Akyürek, E., Wang, B., Kim, Y., and Andreas, J · 2024
Cited alongside, same era.
Approximation of relation functions and attention mechanisms
Altabaa, A. and Lafferty, J · 2024
Cited alongside, same era.
Later among the works it cites.
Function vectors in large language models
Todd, E., Li, M., Sharma, A. S., Mueller, A., Wallace, B. C., and Bau, D · 2024
Later among the works it cites.
The relational bottleneck as an inductive bias for efficient abstraction
Webb, T. W., Frankland, S. M., Altabaa, A., Segert, S., Krishnamurthy, K., Campbell, D., Russin, J., Giallanza, T., O’Reilly, R., Lafferty, J., et al · 2024
Later among the works it cites.
Reasoning or reciting? exploring the capabilities and limitations of language models through counterfactual tasks
Wu, Z., Qiu, L., Ross, A., Akyürek, E., Chen, B., Wang, B., Kim, N., Andreas, J., and Kim, Y · 2024
Later among the works it cites.
Emergence of symbolic abstraction heads for in-context learning in large language models
Al-Saeedi, A. and Härmä, A · 2025
Closest in time.
Emergent symbol-like number variables in artificial neural networks
Grant, S., Goodman, N. D., and McClelland, J. L · 2025
Closest in time.
Analogical reasoning inside large language models: Concept vectors and the limits of abstraction
Opiełka, G., Rosenbusch, H., and Stevenson, C. E · 2025
Closest in time.
Qwen Team · 2025
Closest in time.
An explainable transformer circuit for compositional generalization
Tang, C., Lake, B., and Jazayeri, M · 2025
Closest in time.
Evidence from counterfactual tasks supports emergent analogical reasoning in large language models
Webb, T. W., Holyoak, K. J., and Lu, H · 2025
Closest in time.
How do transformers learn variable binding in symbolic programs?
Wu, Y., Geiger, A., and Millière, R · 2025
Closest in time.
Which attention heads matter for in-context learning?
Yin, K. and Steinhardt, J · 2025
Closest in time.