Fetching the paper…
Reading the bibliography…
Recent breakthroughs in artificial intelligence have driven a paradigm shift, where large language models (LLMs) with billions or trillions of parameters are trained on vast datasets, achieving unprecedented success across a series of language tasks.
M. Smithson, “Conflict aversion: preference for ambiguity vs conflict in sources and evidence,” Organizational behavior and human decision processes , vol. 79, no. 3, pp. 179–198, 1999
1999
Earlier work this paper cites.
K. B. Korb, L. R. Hope, A. E. Nicholson, and K. Axnick, “Varieties of causal intervention,” in PRICAI 2004: Trends in Artificial Intelligence: 8th Pacific Rim International Conference on Artificial Intelligence, Auckland, New Zealand, August 9-13, 2004. Proceedings 8 . Springer, 2004, pp. 322–331
2004
Earlier work this paper cites.
H. Bang and J. M. Robins, “Doubly robust estimation in missing data and causal inference models,” Biometrics , vol. 61, no. 4, pp. 962–973, 2005
2005
Earlier work this paper cites.
J. Pearl, Causality . Cambridge university press, 2009
2009
Earlier work this paper cites.
F. Johansson, U. Shalit, and D. Sontag, “Learning representations for counterfactual inference,” in International conference on machine learning . PMLR, 2016, pp. 3020–3029
2016
Earlier work this paper cites.
A. Vaswani, “Attention is all you need,” Advances in Neural Information Processing Systems , 2017
2017
Earlier work this paper cites.
U. Shalit, F. D. Johansson, and D. Sontag, “Estimating individual treatment effect: generalization bounds and algorithms,” in International Conference on Machine Learning . PMLR, 2017, pp. 3076–3085
2017
Earlier work this paper cites.
S. Zhao, Q. Wang, S. Massung, B. Qin, T. Liu, B. Wang, and C. Zhai, “Constructing and embedding abstract event causality networks from text snippets,” in Proceedings of the Tenth ACM International Conference on Web Search and Data Mining , 2017, pp. 335–344
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei, “Deep reinforcement learning from human preferences,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
K. Kuang, P. Cui, S. Athey, R. Xiong, and B. Li, “Stable prediction across unknown environments,” in proceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data mining , 2018, pp. 1617–1626
2018
Earlier work this paper cites.
M. Rojas-Carulla, B. Schölkopf, R. Turner, and J. Peters, “Invariant models for causal transfer learning,” Journal of Machine Learning Research , vol. 19, no. 36, pp. 1–34, 2018
2018
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Earlier work this paper cites.
R. Zmigrod, S. J. Mielke, H. Wallach, and R. Cotterell, “Counterfactual data augmentation for mitigating gender stereotypes in languages with rich morphology,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , 2019, pp. 1651–1661
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” Journal of machine learning research , vol. 21, no. 140, pp. 1–67, 2020
2020
Earlier work this paper cites.
T. B. Brown, “Language models are few-shot learners,” arXiv preprint arXiv:2005.14165 , 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
D. Kaushik, E. Hovy, and Z. Lipton, “Learning the difference that makes a difference with counterfactually-augmented data,” in International Conference on Learning Representations , 2020. [Online]. Available: https://openreview.net/forum?id=Sklgs0NFvr
2020
Earlier work this paper cites.
D. Zhang, H. Zhang, J. Tang, X.-S. Hua, and Q. Sun, “Causal intervention for weakly-supervised semantic segmentation,” Advances in Neural Information Processing Systems , vol. 33, pp. 655–666, 2020
2020
Earlier work this paper cites.
M. Kaneko and D. Bollegala, “Debiasing pre-trained contextualised embeddings,” in Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume , 2021, pp. 1256–1266
2021
Earlier work this paper cites.
T. Wang, C. Zhou, Q. Sun, and H. Zhang, “Causal attention for unbiased visual recognition,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 3091–3100
2021
Earlier work this paper cites.
X. Yang, H. Zhang, G. Qi, and J. Cai, “Causal attention for vision-language tasks,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2021, pp. 9847–9857
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
L. Yao, Z. Chu, S. Li, Y. Li, J. Gao, and A. Zhang, “A survey on causal inference,” ACM Transactions on Knowledge Discovery from Data (TKDD) , vol. 15, no. 5, pp. 1–46, 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
L. Liao, Z. Fu, Z. Yang, Y. Wang, M. Kolar, and Z. Wang, “Instrumental variable value iteration for causal offline reinforcement learning,” stat , vol. 1050, p. 13, 2021
2021
Earlier work this paper cites.
Y. Guo, Y. Yang, and A. Abbasi, “Auto-debias: Debiasing masked language models with automated biased prompts,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2022, pp. 1012–1023
2022
Earlier work this paper cites.
J. He, M. Xia, C. Fellbaum, and D. Chen, “Mabel: Attenuating gender bias using textual entailment data,” in Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , 2022, pp. 9681–9702
2022
Earlier work this paper cites.
J. Zhang, H. Zhang, W. Su, and D. Roth, “Rock: Causal inference principles for reasoning about commonsense causality,” in International Conference on Machine Learning . PMLR, 2022, pp. 26 750–26 771
2022
Earlier work this paper cites.
J. Frohberg and F. Binder, “Crass: A novel data set and benchmark to test counterfactual reasoning of large language models,” in Proceedings of the Thirteenth Language Resources and Evaluation Conference , 2022, pp. 2126–2140
2022
Earlier work this paper cites.
T. Lin, Y. Wang, X. Liu, and X. Qiu, “A survey of transformers,” AI open , vol. 3, pp. 111–132, 2022
2022
Earlier work this paper cites.
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “Lora: Low-rank adaptation of large language models,” in International Conference on Learning Representations , 2022. [Online]. Available: https://openreview.net/forum?id=nZeVKeeFYf9
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
A. Wu, J. Yuan, K. Kuang, B. Li, R. Wu, Q. Zhu, Y. Zhuang, and F. Wu, “Learning decomposed representations for treatment effect estimation,” IEEE Transactions on Knowledge and Data Engineering , vol. 35, no. 5, pp. 4989–5001, 2022
2022
Earlier work this paper cites.
N. Meade, E. Poole-Dayan, and S. Reddy, “An empirical survey of the effectiveness of debiasing techniques for pre-trained language models,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2022, pp. 1878–1898
2022
Earlier work this paper cites.
T. Le Scao, A. Fan, C. Akiki, E. Pavlick, S. Ilić, D. Hesslow, R. Castagné, A. S. Luccioni, F. Yvon, M. Gallé et al. , “Bloom: A 176b-parameter open-access multilingual language model,” 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn, “Direct preference optimization: Your language model is secretly a reward model,” in Thirty-seventh Conference on Neural Information Processing Systems , 2023. [Online]. Available: https://openreview.net/forum?id=HPuSIXJaa9
2023
Later among the works it cites.
2023
Later among the works it cites.
Y. Chang, X. Wang, J. Wang, Y. Wu, L. Yang, K. Zhu, H. Chen, X. Yi, C. Wang, Y. Wang et al. , “A survey on evaluation of large language models,” ACM Transactions on Intelligent Systems and Technology , vol. 15, no. 3, pp. 1–45, 2024
2024
Closest in time.
2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
2023
Cited alongside, same era.
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann et al. , “Palm: Scaling language modeling with pathways,” Journal of Machine Learning Research , vol. 24, no. 240, pp. 1–113, 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
H. Li, Q. Ai, J. Chen, Q. Dong, Y. Wu, Y. Liu, C. Chen, and Q. Tian, “Sailer: structure-aware pre-trained language model for legal case retrieval,” in Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval , 2023, pp. 1035–1044
2023
Cited alongside, same era.
J. Mökander, J. Schuett, H. R. Kirk, and L. Floridi, “Auditing large language models: a three-layered approach,” AI and Ethics , pp. 1–31, 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Closest in time.
M. A. K. Raiaan, M. S. H. Mukta, K. Fatema, N. M. Fahad, S. Sakib, M. M. J. Mim, J. Ahmad, M. E. Ali, and S. Azam, “A review on large language models: Architectures, applications, taxonomies, open issues and challenges,” IEEE Access , 2024
2024
Closest in time.
I. O. Gallegos, R. A. Rossi, J. Barrow, M. M. Tanjim, S. Kim, F. Dernoncourt, T. Yu, R. Zhang, and N. K. Ahmed, “Bias and fairness in large language models: A survey,” Computational Linguistics , pp. 1–79, 2024
2024
Closest in time.
2024
Closest in time.
N. Guha, J. Nyarko, D. Ho, C. Ré, A. Chilton, A. Chohlas-Wood, A. Peters, B. Waldon, D. Rockmore, D. Zambrano et al. , “Legalbench: A collaboratively built benchmark for measuring legal reasoning in large language models,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
Z. Jin, Y. Chen, F. Leeb, L. Gresele, O. Kamal, Z. Lyu, K. Blin, F. Gonzalez Adauto, M. Kleiman-Weiner, M. Sachan et al. , “Cladder: A benchmark to assess causal reasoning capabilities of language models,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
A. Bhattacharjee, R. Moraffah, J. Garland, and H. Liu, “Towards llm-guided causal explainability for black-box text classifiers,” in AAAI 2024 Workshop on Responsible Language Models , Vancouver, BC, Canada, 2024
2024
Closest in time.
R. Y. Rohekar, Y. Gurwicz, and S. Nisimov, “Causal interpretation of self-attention in pre-trained transformers,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
J. Zhang, J. Jennings, A. Hilmkil, N. Pawlowski, C. Zhang, and C. Ma, “Towards causal foundation model: on duality between optimal balancing and attention,” in Forty-first International Conference on Machine Learning , 2024. [Online]. Available: https://openreview.net/forum?id=cFDaYtZR4u
2024
Closest in time.
K. Zhang, D. Zhang, L. Wu, R. Hong, Y. Zhao, and M. Wang, “Label-aware debiased causal reasoning for natural language inference,” AI Open , vol. 5, pp. 70–78, 2024
2024
Closest in time.
A. Feder, Y. Wald, C. Shi, S. Saria, and D. Blei, “Causal-structure driven augmentations for text ood generalization,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
A. Nie, Y. Zhang, A. S. Amdekar, C. Piech, T. B. Hashimoto, and T. Gerstenberg, “Moca: Measuring human-language model alignment on causal and moral judgment tasks,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
R. Ohtani, Y. Sakurai, and S. Oyama, “Does metacognitive prompting improve causal inference in large language models?” in 2024 IEEE Conference on Artificial Intelligence (CAI) . IEEE, 2024, pp. 458–459
2024
Closest in time.
Z. Jin, J. Liu, L. Zhiheng, S. Poff, M. Sachan, R. Mihalcea, M. T. Diab, and B. Schölkopf, “Can large language models infer causation from correlation?” in The Twelfth International Conference on Learning Representations , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
H. Duong Le, X. Xia, and Z. Chen, “Multi-agent causal discovery using large language models,” arXiv e-prints , pp. arXiv–2407, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
H. Bao, L. Dong, W. Wang, N. Yang, S. Piao, and F. Wei, “Fine-tuning pretrained transformer encoders for sequence-to-sequence learning,” International Journal of Machine Learning and Cybernetics , vol. 15, no. 5, pp. 1711–1728, 2024
2024
Closest in time.
T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer, “Qlora: Efficient finetuning of quantized llms,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn, “Direct preference optimization: Your language model is secretly a reward model,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
S. Pan, L. Luo, Y. Wang, C. Chen, J. Wang, and X. Wu, “Unifying large language models and knowledge graphs: A roadmap,” IEEE Transactions on Knowledge and Data Engineering , 2024
2024
Closest in time.
E. Fedorenko, S. T. Piantadosi, and E. A. Gibson, “Language is primarily a tool for communication rather than thought,” Nature , vol. 630, no. 8017, pp. 575–586, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.