Fetching the paper…
Reading the bibliography…
For text-based AI systems to interact in the real world, causal reasoning is an essential skill.
Axioms of causal relevance
Galles, D. and Pearl, J · 1997
Earlier work this paper cites.
Causality: Models, Reasoning and Inference
Pearl, J · 2009
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2018
Earlier work this paper cites.
On the ability and limitations of transformers to recognize formal languages, 2020
Bhattamishra, S., Ahuja, K., and Goyal, N · 2020
Earlier work this paper cites.
Compositionality decomposed: how do neural networks generalise?, 2020
Hupkes, D., Dankers, V., Mul, M., and Bruni, E · 2020
Earlier work this paper cites.
Transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Le Scao, T., Gugger, S., Drame, M., Lhoest, Q., and Rush, A · 2020
Earlier work this paper cites.
The devil is in the detail: Simple tricks improve systematic generalization of transformers
Csordás, R., Irie, K., and Schmidhuber, J · 2021
Earlier work this paper cites.
Compositional generalization in semantic parsing: Pre-training vs. specialized architectures, 2021
Furrer, D., van Zee, M., Scales, N., and Schärli, N · 2021
Earlier work this paper cites.
On pearl’s hierarchy and the foundations of causal inference (1st edition)
Bareinboim, E., Correa, J., Ibeling, D., and Icard, T · 2022
Earlier work this paper cites.
Transformer language models without positional encodings still learn positional information
Haviv, A., Ram, O., Press, O., Izsak, P., and Levy, O · 2022
Earlier work this paper cites.
Making transformers solve compositional tasks, 2022
Ontañón, S., Ainslie, J., Cvicek, V., and Fisher, Z · 2022
Earlier work this paper cites.
Probing for correlations of causal facts: Large language models and causality
Willig, M., Zečević, M., Dhami, D. S., and Kersting, K · 2022
Earlier work this paper cites.
Ban, T., Chen, L., Wang, X., and Chen, H · 2023
Cited alongside, same era.
Gemini: A family of highly capable multimodal models
Google DeepMind and Google Research · 2023
Cited alongside, same era.
Textbooks are all you need, 2023
Gunasekar, S., Zhang, Y., Aneja, J., Mendes, C. C. T., Giorno, A. D., Gopi, S., Javaheripi, M., Kauffmann, P., de Rosa, G., Saarikivi, O., Salim, A., Shah, S., Behl, H. S., Wang, X., Bubeck, S., Eldan, R., Kalai, A. T., Lee, Y. T., and Li, Y · 2023
Cited alongside, same era.
Cladder: Assessing causal reasoning in language models
Jin, Z., Chen, Y., Leeb, F., Gresele, L., Kamal, O., Lyu, Z., Blin, K., Adauto, F. G., Kleiman-Weiner, M., Sachan, M., and Schölkopf, B · 2023
Cited alongside, same era.
The impact of positional encoding on length generalization in transformers
Kazemnejad, A., Padhi, I., Natesan Ramamurthy, K., Das, P., and Reddy, S · 2023
Cited alongside, same era.
Roformer: Enhanced transformer with rotary position embedding
Su, J., Ahmed, M., Lu, Y., Pan, S., Bo, W., and Liu, Y · 2023
Later among the works it cites.
Causal parrots: Large language models may talk causality but are not causal
Zečević, M., Willig, M., Dhami, D. S., and Kersting, K · 2023
Later among the works it cites.
Unveiling transformers with lego: a synthetic reasoning task, 2023
Zhang, Y., Backurs, A., Bubeck, S., Eldan, R., Gunasekar, S., and Wagner, T · 2023
Later among the works it cites.
The Llama 3 herd of models, 2024
2024
Closest in time.
Phi-3 technical report: A highly capable language model locally on your phone
Abdin, M. et al · 2024
Closest in time.
The reversal curse: LLMs trained on "a is b" fail to learn "b is a", 2024
Berglund, L., Tong, M., Kaufmann, M., Balesni, M., Stickland, A. C., Korbak, T., and Evans, O · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Causal reasoning and large language models: Opening a new frontier for causality
Kıcıman, E., Ness, R., Sharma, A., and Tan, C · 2023
Cited alongside, same era.
Passive learning of active causal strategies in agents and language models
Lampinen, A., Chan, S., Dasgupta, I., Nam, A., and Wang, J · 2023
Cited alongside, same era.
Textbooks are all you need ii: phi-1.5 technical report, 2023
Li, Y., Bubeck, S., Eldan, R., Giorno, A. D., Gunasekar, S., and Lee, Y. T · 2023
Cited alongside, same era.
Causal discovery with language models as imperfect experts
Long, S., Piché, A., Zantedeschi, V., Schuster, T., and Drouin, A · 2023
Cited alongside, same era.
GPT-4 technical report
OpenAI · 2023
Cited alongside, same era.
Injecting structural hints: Using language models to study inductive biases in language learning
Papadimitriou, I. and Jurafsky, D · 2023
Cited alongside, same era.
Testing the general deductive reasoning capacity of large language models using ood examples, 2023
Saparov, A., Pang, R. Y., Padmakumar, V., Joshi, N., Kazemi, S. M., Kim, N., and He, H · 2023
Cited alongside, same era.
Closest in time.
CLEAR: Can language models really understand causal graphs?
Chen, S., Xu, M., Wang, K., Zeng, X., Zhao, R., Zhao, S., and Lu, C · 2024
Closest in time.
Can large language models infer causation from correlation?
Jin, Z., Liu, J., Lyu, Z., Poff, S., Sachan, M., Mihalcea, R., Diab, M., and Schölkopf, B · 2024
Closest in time.
Enhancing reasoning capabilities of llms via principled synthetic logic corpus
Morishita, T., Morio, G., Yamaguchi, A., and Sogawa, Y · 2024
Closest in time.
Axiomatization of interventional probability distributions
Sadeghi, K. and Soo, T · 2024
Closest in time.
Solving olympiad geometry without human demonstrations
Trinh, T. H., Wu, Y., Le, Q. V., He, H., and Luong, T · 2024
Closest in time.
Cognitive behaviors that enable self-improving reasoners, or, four habits of highly effective stars, 2025
Gandhi, K., Chakravarthy, A., Singh, A., Lile, N., and Goodman, N. D · 2025
Closest in time.
Causal order: The key to leveraging imperfect experts in causal inference
Vashishtha, A., Reddy, A. G., Kumar, A., Bachu, S., Balasubramanian, V. N., and Sharma, A · 2025
Closest in time.