Fetching the paper…
Reading the bibliography…
Causal reasoning is viewed as crucial for achieving human-level machine intelligence.
Multiple causation and damage
Peaslee, R. J · 1934
Earlier work this paper cites.
Two sides of the same coin: Exploiting the impact of identifiers in neural code comprehension
Gao, S., Gao, C., Wang, C., Sun, J., Lo, D., and Yu, Y · 1945
Earlier work this paper cites.
Limitations of the application of fourfold table analysis to hospital data
Berkson, J · 1946
Earlier work this paper cites.
A mathematical theory of communication
Shannon, C. E · 1948
Earlier work this paper cites.
Speech production and the predictability of words in context
Goldman-Eisler, F · 1958
Earlier work this paper cites.
The perceptron: a probabilistic model for information storage and organization in the brain
Rosenblatt, F · 1958
Earlier work this paper cites.
Continuous speech recognition by statistical methods
Jelinek, F · 1976
Earlier work this paper cites.
A computational model for causal and diagnostic reasoning in inference systems
Kim, J. and Pearl, J · 1983
Earlier work this paper cites.
Norm theory: Comparing reality to its alternatives
Kahneman, D. and Miller, D. T · 1986
Earlier work this paper cites.
Evaluating the econometric evaluations of training programs with experimental data
LaLonde, R. J · 1986
Earlier work this paper cites.
Probabilistic reasoning in intelligent systems: networks of plausible inference
Pearl, J · 1988
Earlier work this paper cites.
Causal diagrams for empirical research
Pearl, J · 1995
Earlier work this paper cites.
Identification of causal effects using instrumental variables
Angrist, J. D., Imbens, G. W., and Rubin, D. B · 1996
Earlier work this paper cites.
Statistical methods for speech recognition
Jelinek, F · 1998
Earlier work this paper cites.
The analects of Confucius: A philosophical translation
Ames, R. T. and Rosemont Jr, H · 1999
Earlier work this paper cites.
A neural probabilistic language model
Bengio, Y., Ducharme, R., and Vincent, P · 2000
Earlier work this paper cites.
Two decades of statistical language modeling: Where do we go from here?
Rosenfeld, R · 2000
Earlier work this paper cites.
Causation, prediction, and search
Spirtes, P., Glymour, C. N., and Scheines, R · 2000
Earlier work this paper cites.
Direct and indirect effects
Pearl, J · 2001
Earlier work this paper cites.
A general identification condition for causal effects
Tian, J. and Pearl, J · 2002
Earlier work this paper cites.
Efficient estimation of average treatment effects using the estimated propensity score
Hirano, K., Imbens, G. W., and Ridder, G · 2003
Earlier work this paper cites.
The appraisal basis of anger: specificity, necessity and sufficiency of components
Kuppens, P., Van Mechelen, I., Smits, D. J., and De Boeck, P · 2003
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Lin, C.-Y · 2004
Earlier work this paper cites.
Improved estimation of controlled direct effects in the presence of unmeasured confounding of intermediate variables
Kaufman, S., Kaufman, J. S., MacLehose, R. F., Greenland, S., and Poole, C · 2005
Earlier work this paper cites.
Timebank 1.2 documentation
Pustejovsky, J., Littman, J., Saurí, R., and Verhagen, M · 2006
Earlier work this paper cites.
The rational imagination: How people create alternatives to reality
Byrne, R. M · 2007
Earlier work this paper cites.
Neural-symbolic cognitive reasoning
Garcez, A. S., Lamb, L. C., and Gabbay, D. M · 2008
Earlier work this paper cites.
Probabilistic logic networks: A comprehensive framework for uncertain inference
Goertzel, B., Iklé, M., Goertzel, I. F., and Heljakka, A · 2008
Earlier work this paper cites.
Complete identification methods for the causal hierarchy
Shpitser, I. and Pearl, J · 2008
Earlier work this paper cites.
Explaining away: A model of affective adaptation
Wilson, T. D. and Gilbert, D. T · 2008
Earlier work this paper cites.
Causality
Pearl, J · 2009
Earlier work this paper cites.
Illustrating bias due to conditioning on a collider
Cole, S. R., Platt, R. W., Schisterman, E. F., Chu, H., Westreich, D., Richardson, D., and Poole, C · 2010
Earlier work this paper cites.
Recurrent neural network based language model
Mikolov, T., Karafiát, M., Burget, L., Cernocký, J., and Khudanpur, S · 2010
Earlier work this paper cites.
Infobert: Improving robustness of language models from an information theoretic perspective
Wang, B., Wang, S., Cheng, Y., Gan, Z., Jia, R., Li, B., and Liu, J · 2010
Earlier work this paper cites.
Bias-corrected matching estimators for average treatment effects
Abadie, A. and Imbens, G. W · 2011
Earlier work this paper cites.
Comparison of values of pearson’s and spearman’s correlation coefficients on the same sets of data
Hauke, J. and Kossowski, T · 2011
Earlier work this paper cites.
Bayesian nonparametric modeling for causal inference
Hill, J. L · 2011
Earlier work this paper cites.
Glottolog/langdoc: Defining dialects, languages, and language families as collections of resources
Nordhoff, S. and Hammarström, H · 2011
Earlier work this paper cites.
Choice of plausible alternatives: An evaluation of commonsense causal reasoning
Roemmele, M., Bejan, C. A., and Gordon, A. S · 2011
Earlier work this paper cites.
Instrumental variables in sociology and the social sciences
Bollen, K. A · 2012
Earlier work this paper cites.
Socioeconomic status and cognitive functioning: moving from correlation to causation
Duncan, G. J. and Magnuson, K · 2012
Earlier work this paper cites.
Generalized propensity score for estimating the average treatment effect of multiple treatments
Feng, P., Zhou, X.-H., Zou, Q.-M., Fan, M.-Y., and Li, X.-S · 2012
Earlier work this paper cites.
Pearson’s correlation coefficient
Sedgwick, P · 2012
Earlier work this paper cites.
Graphical causal models
Elwert, F · 2013
Earlier work this paper cites.
The cancer genome atlas pan-cancer analysis project
Weinstein, J. N., Collisson, E. A., Mills, G. B., Shaw, K. R., Ozenberger, B. A., Ellrott, K., Shmulevich, I., Sander, C., and Stuart, J. M · 2013
Earlier work this paper cites.
Endogenous selection bias: The problem of conditioning on a collider variable
Elwert, F. and Winship, C · 2014
Earlier work this paper cites.
Attribution theory in the organizational sciences: The road traveled and the path ahead
Harvey, P., Madison, K., Martinko, M., Crook, T. R., and Crook, T. A · 2014
Earlier work this paper cites.
Annotating causality in the tempeval-3 corpus
Mirza, P., Sprugnoli, R., Tonelli, S., and Speranza, M · 2014
Earlier work this paper cites.
Doubly robust estimation of causal effects with multivalued treatments: an application to the returns to schooling
Uysal, S. D · 2015
Earlier work this paper cites.
Actual Causality
Halpern, J. Y · 2016
Earlier work this paper cites.
Robust text classification in the presence of confounding bias
Landeiro, V. and Culotta, A · 2016
Earlier work this paper cites.
Causal inference in statistics: A primer
Pearl, J., Glymour, M., and Jewell, N. P · 2016
Earlier work this paper cites.
Causal discovery and inference: concepts and recent methodological advances
Spirtes, P. and Zhang, K · 2016
Earlier work this paper cites.
Video summarization with long short-term memory
Zhang, K., Chao, W.-L., Sha, F., and Grauman, K · 2016
Earlier work this paper cites.
The event storyline corpus: A new benchmark for causal and temporal relation extraction
Caselli, T. and Vossen, P · 2017
Earlier work this paper cites.
Counterfactual fairness
Kusner, M. J., Loftus, J., Russell, C., and Silva, R · 2017
Earlier work this paper cites.
Causal effect inference with deep latent-variable models
Louizos, C., Shalit, U., Mooij, J. M., Sontag, D., Zemel, R., and Welling, M · 2017
Earlier work this paper cites.
Elements of causal inference: foundations and learning algorithms
Peters, J., Janzing, D., and Schölkopf, B · 2017
Earlier work this paper cites.
The Oxford handbook of causal reasoning
Waldmann, M · 2017
Earlier work this paper cites.
Generalized adjustment under confounding and selection biases
Correa, J., Tian, J., and Bareinboim, E · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
Randomised controlled trials—the gold standard for effectiveness research
Hariton, E. and Locascio, J. J · 2018
Earlier work this paper cites.
Robust text classification under confounding shift
Landeiro, V. and Culotta, A · 2018
Earlier work this paper cites.
Collider scope: when selection bias can substantially influence observed associations
Munafò, M. R., Tilling, K., Taylor, A. E., Evans, D. M., and Davey Smith, G · 2018
Earlier work this paper cites.
The book of why: the new science of cause and effect
Pearl, J. and Mackenzie, D · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et al · 2018
Earlier work this paper cites.
From correlation to causation: analysis of metabolomics data using systems biology approaches
Rosato, A., Tenori, L., Cascante, M., De Atauri Carulla, P. R., Martins dos Santos, V. A., and Saccenti, E · 2018
Earlier work this paper cites.
Benchmarking framework for performance-evaluation of causal inference analysis
Shimoni, Y., Yanover, C., Karavani, E., and Goldschmnidt, Y · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2018
Earlier work this paper cites.
What caused what? a quantitative account of actual causation using dynamical causal networks
Albantakis, L., Marshall, W., Hoel, E., and Tononi, G · 2019
Earlier work this paper cites.
Modeling document-level causal structures for event causal relation identification
Gao, L., Choubey, P. K., and Huang, R · 2019
Earlier work this paper cites.
Review of causal discovery methods based on graphical models
Glymour, C., Zhang, K., and Spirtes, P · 2019
Earlier work this paper cites.
The seven tools of causal inference, with reflections on machine learning
Pearl, J · 2019
Earlier work this paper cites.
Detecting and quantifying causal associations in large nonlinear time series datasets
Runge, J., Nowack, P., Kretschmer, M., Flaxman, S., and Sejdinovic, D · 2019
Earlier work this paper cites.
Universal adversarial triggers for attacking and analyzing nlp
Wallace, E., Feng, S., Kandpal, N., Gardner, M., and Singh, S · 2019
Earlier work this paper cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
Wang, A., Pruksachatkun, Y., Nangia, N., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S · 2019
Earlier work this paper cites.
A causally formulated hazard ratio estimation through backdoor adjustment on structural causal model
Adib, R., Griffin, P., Ahamed, S. I., and Adibuzzaman, M · 2020
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Relation-aware collaborative learning for unified aspect-based sentiment analysis
Chen, Z. and Qian, T · 2020
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Gehman, S., Gururangan, S., Sap, M., Choi, Y., and Smith, N. A · 2020
Earlier work this paper cites.
An attributional theory of motivation
Graham, S · 2020
Earlier work this paper cites.
Learning individual causal effects from networked observational data
Guo, R., Li, J., and Liu, H · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Earlier work this paper cites.
Potential outcome and directed acyclic graph approaches to causality: Relevance for empirical practice in economics
Imbens, G. W · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., et al · 2020
Cited alongside, same era.
Conducting sensitivity analysis for unmeasured confounding in observational studies using e-values: the evalue package
Linden, A., Mathur, M. B., and VanderWeele, T. J · 2020
Cited alongside, same era.
Explainable reinforcement learning through a causal lens
Madumal, P., Miller, T., Sonenberg, L., and Vetere, F · 2020
Cited alongside, same era.
Profiling compliers and noncompliers for instrumental-variable analysis
Marbach, M. and Hangartner, D · 2020
Cited alongside, same era.
Estimating causal effects in linear regression models with observational data: The instrumental variables regression model
Maydeu-Olivares, A., Shi, D., and Fairchild, A. J · 2020
Cited alongside, same era.
Least-to-most prompting enables complex reasoning in large language models
Zhou, D., Schärli, N., Hou, L., Wei, J., Scales, N., Wang, X., Schuurmans, D., Cui, C., Bousquet, O., Le, Q., et al · 2022
Later among the works it cites.
Model card and evaluations for claude models
Anthropic · 2023
Later among the works it cites.
Applying the structural causal model framework for observational causal inference in ecology
Arif, S. and MacNeil, M. A · 2023
Later among the works it cites.
Llemma: An open language model for mathematics
Azerbayev, Z., Schoelkopf, H., Paster, K., Santos, M. D., McAleer, S., Jiang, A. Q., Deng, J., Biderman, S., and Welleck, S · 2023
Later among the works it cites.
Bai, J., Bai, S., Chu, Y., Cui, Z., Dang, K., Deng, X., Fan, Y., Ge, W., Han, Y., Huang, F., Hui, B., Ji, L., Li, M., Lin, J., Lin, R., Liu, D., Liu, G., Lu, C., Lu, K., Ma, J., Men, R., Ren, X., Ren, X., Tan, C., Tan, S., Tu, J., Wang, P., Wang, S., Wang, W., Wu, S., Xu, B., Xu, J., Yang, A., Yang, H., Yang, J., Yang, S., Yao, Y., Yu, B., Yuan, H., Yuan, Z., Zhang, J., Zhang, X., Zhang, Y., Zhang, Z., Zhou, C., Zhou, J., Zhou, X., and Zhu, T · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Effectual and causal reasoning in the adoption of marketing automation
Mero, J., Tarkiainen, A., and Tobon, J · 2020
Cited alongside, same era.
Causal interpretability for machine learning-problems, methods and evaluation
Moraffah, R., Karami, M., Guo, R., Raglin, A., and Liu, H · 2020
Cited alongside, same era.
Generative causal explanations of black-box classifiers
O’Shaughnessy, M., Canal, G., Connor, M., Rozell, C., and Davenport, M · 2020
Cited alongside, same era.
Robust predictors for seasonal atlantic hurricane activity identified with causal effect networks
Pfleiderer, P., Schleussner, C.-F., Geiger, T., and Kretschmer, M · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Cited alongside, same era.
Improving the accuracy of medical diagnosis with causal machine learning
Richens, J. G., Lee, C. M., and Johri, S · 2020
Cited alongside, same era.
Housing as a social determinant of health and wellbeing: Developing an empirically-informed realist theoretical framework
Rolfe, S., Garnham, L., Godwin, J., Anderson, I., Seaman, P., and Donaldson, C · 2020
Cited alongside, same era.
Later among the works it cites.
Baichuan 2: Open large-scale language models
Baichuan · 2023
Later among the works it cites.
Ban, T., Chen, L., Wang, X., and Chen, H · 2023
Later among the works it cites.
Sparks of artificial general intelligence: Early experiments with gpt-4
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., et al · 2023
Later among the works it cites.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., and Xing, E. P · 2023
Later among the works it cites.
Robust sentiment classification based on the backdoor adjustment
Dai, L. and Han, M · 2023
Later among the works it cites.
Dao, X.-Q. and Le, N.-B · 2023
Later among the works it cites.
Two-way fixed effects and differences-in-differences with heterogeneous treatment effects: A survey
De Chaisemartin, C. and d’Haultfoeuille, X · 2023
Later among the works it cites.
Multilingual jailbreak challenges in large language models
Deng, Y., Zhang, W., Pan, S. J., and Bing, L · 2023
Later among the works it cites.
Chatgpt outperforms humans in emotional awareness evaluations
Elyoseph, Z., Hadar-Shoval, D., Asraf, K., and Lvovsky, M · 2023
Later among the works it cites.
Bias and fairness in large language models: A survey
Gallegos, I. O., Rossi, R. A., Barrow, J., Tanjim, M. M., Kim, S., Dernoncourt, F., Yu, T., Zhang, R., and Ahmed, N. K · 2023
Later among the works it cites.
Gan, C., Zhang, Q., and Mori, T · 2023
Later among the works it cites.
Is chatgpt a good causal reasoner? a comprehensive evaluation
Gao, J., Ding, X., Qin, B., and Liu, T · 2023
Later among the works it cites.
Koala: A dialogue model for academic research
Geng, X., Gudibande, A., Liu, H., Wallace, E., Abbeel, P., Levine, S., and Song, D · 2023
Later among the works it cites.
More than you’ve asked for: A comprehensive analysis of novel prompt injection threats to application-integrated large language models
Greshake, K., Abdelnabi, S., Mishra, S., Endres, C., Holz, T., and Fritz, M · 2023
Later among the works it cites.
Solving math word problems by combining language models with symbolic solvers
He-Yueya, J., Poesia, G., Wang, R., and Goodman, N · 2023
Later among the works it cites.
Causalainer: Causal explainer for automatic video summarization
Huang, J.-H., Yang, C.-H. H., Chen, P.-Y., Chen, M.-H., and Worring, M · 2023
Later among the works it cites.
Mathprompter: Mathematical reasoning using large language models
Imani, S., Du, L., and Shrivastava, H · 2023
Later among the works it cites.
Benchmarking and explaining large language model-based code generation: A causality-centric approach
Ji, Z., Ma, P., Li, Z., and Wang, S · 2023
Later among the works it cites.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al · 2023
Later among the works it cites.
How novices use llm-based code generators to solve cs1 coding tasks in a self-paced learning environment
Kazemitabaar, M., Hou, X., Henley, A., Ericson, B. J., Weintrop, D., and Grossman, T · 2023
Later among the works it cites.
Causal reasoning and large language models: Opening a new frontier for causality
Kıcıman, E., Ness, R., Sharma, A., and Tan, C · 2023
Later among the works it cites.
Multi-step jailbreaking privacy attacks on chatgpt
Li, H., Guo, D., Fan, W., Xu, M., Huang, J., Meng, F., and Song, Y · 2023
Later among the works it cites.
Halueval: A large-scale hallucination evaluation benchmark for large language models
Li, J., Cheng, X., Zhao, X., Nie, J.-Y., and Wen, J.-R · 2023
Later among the works it cites.
Let’s verify step by step
Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., and Cobbe, K · 2023
Later among the works it cites.
Wizardcoder: Empowering code large language models with evol-instruct, 2023
Luo, Z., Xu, C., Zhao, P., Sun, Q., Geng, X., Hu, W., Tao, C., Ma, J., Lin, Q., and Jiang, D · 2023
Later among the works it cites.
Augmented language models: a survey
Mialon, G., Dessi, R., Lomeli, M., Nalmpantis, C., Pasunuru, R., Raileanu, R., Roziere, B., Schick, T., Dwivedi-Yu, J., Celikyilmaz, A., et al · 2023
Later among the works it cites.
Beyond accuracy: Evaluating self-consistency of code llms
Min, M. J., Ding, Y., Buratti, L., Pujar, S., Kaiser, G., Jana, S., and Ray, B · 2023
Later among the works it cites.
Gpt-4 technical report, 2023
OpenAI · 2023
Later among the works it cites.
Proving test set contamination for black-box language models
Oren, Y., Meister, N., Chatterji, N. S., Ladhak, F., and Hashimoto, T · 2023
Later among the works it cites.
Art: Automatic multi-step reasoning and tool-use for large language models
Paranjape, B., Lundberg, S., Singh, S., Hajishirzi, H., Zettlemoyer, L., and Ribeiro, M. T · 2023
Later among the works it cites.
Nested markov properties for acyclic directed mixed graphs
Richardson, T. S., Evans, R. J., Robins, J. M., and Shpitser, I · 2023
Later among the works it cites.
Benchmarking causal study to interpret large language models for source code
Rodriguez-Cardenas, D., Palacio, D. N., Khati, D., Burke, H., and Poshyvanyk, D · 2023
Later among the works it cites.
What’s trending in difference-in-differences? a synthesis of the recent econometrics literature
Roth, J., Sant’Anna, P. H., Bilinski, A., and Poe, J · 2023
Later among the works it cites.
Code llama: Open foundation models for code
Roziere, B., Gehring, J., Gloeckle, F., Sootla, S., Gat, I., Tan, X. E., Adi, Y., Liu, J., Remez, T., Rapin, J., et al · 2023
Later among the works it cites.
Causal inference for time series
Runge, J., Gerhardus, A., Varando, G., Eyring, V., and Camps-Valls, G · 2023
Later among the works it cites.
Arb: Advanced reasoning benchmark for large language models
Sawada, T., Paleka, D., Havrilla, A., Tadepalli, P., Vidas, P., Kranias, A., Nay, J., Gupta, K., and Komatsuzaki, A · 2023
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., et al · 2023
Later among the works it cites.
Challenging BIG-bench tasks and whether chain-of-thought can solve them
Suzgun, M., Scales, N., Schärli, N., Gehrmann, S., Tay, Y., Chung, H. W., Chowdhery, A., Le, Q., Chi, E., Zhou, D., and Wei, J · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Team, G., Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., et al · 2023
Later among the works it cites.
Internlm: A multilingual language model with progressively enhanced capabilities, 2023
Team, I · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Causal-discovery performance of chatgpt in the context of neuropathic pain diagnosis
Tu, R., Ma, C., and Zhang, C · 2023
Later among the works it cites.
Causal inference using llm-guided discovery
Vashishtha, A., Reddy, A. G., Kumar, A., Bachu, S., Balasubramanian, V. N., and Sharma, A · 2023
Later among the works it cites.
Prompt engineering
Weng, L · 2023
Later among the works it cites.
Glue-x: Evaluating natural language understanding models from an out-of-distribution generalization perspective
Yang, L., Zhang, S., Qin, L., Li, Y., Wang, Y., Liu, H., Wang, J., Xie, X., and Zhang, Y · 2023
Later among the works it cites.
Metamath: Bootstrap your own mathematical questions for large language models
Yu, L., Jiang, W., Shi, H., Jincheng, Y., Liu, Z., Zhang, Y., Kwok, J., Li, Z., Weller, A., and Liu, W · 2023
Later among the works it cites.
Causal parrots: Large language models may talk causality but are not causal
Zečević, M., Willig, M., Dhami, D. S., and Kersting, K · 2023
Later among the works it cites.
Towards generic image manipulation detection with weakly-supervised self-consistency learning
Zhai, Y., Luan, T., Doermann, D., and Yuan, J · 2023
Later among the works it cites.
Causality in the time of llms: Round table discussion results of clear 2023
Zhang, C., Janzing, D., van der Schaar, M., Locatello, F., and Spirtes, P · 2023
Later among the works it cites.
A survey of large language models
Zhao, W. X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., Min, Y., Zhang, B., Zhang, J., Dong, Z., et al · 2023
Later among the works it cites.
Safety and ethical concerns of large language models
Zhiheng, X., Rui, Z., and Tao, G · 2023
Later among the works it cites.
A study on robustness and reliability of large language model code generation
Zhong, L. and Wang, Z · 2023
Later among the works it cites.
Securing large language models: Threats, vulnerabilities and responsible practices
Abdali, S., Anarfi, R., Barberan, C., and He, J · 2024
Closest in time.
Large language models for mathematical reasoning: Progresses and challenges
Ahn, J., Verma, R., Lou, R., Liu, D., Zhang, R., and Yin, W · 2024
Closest in time.
Introducing the next generation of claude
Anthropic · 2024
Closest in time.
Graph of thoughts: Solving elaborate problems with large language models
Besta, M., Blach, N., Kubicek, A., Gerstenberger, R., Podstawski, M., Gianinazzi, L., Gajda, J., Lehmann, T., Niewiadomski, H., Nyczyk, P., et al · 2024
Closest in time.
Quantifying and mitigating unimodal biases in multimodal large language models: A causal perspective
Chen, M., Cao, Y., Zhang, Y., and Lu, C · 2024
Closest in time.
Security and privacy challenges of large language models: A survey
Das, B. C., Amini, M. H., and Wu, Y · 2024
Closest in time.
Qlora: Efficient finetuning of quantized llms
Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L · 2024
Closest in time.
Xiezhi: An ever-updating benchmark for holistic domain knowledge evaluation
Gu, Z., Zhu, X., Ye, H., Zhang, L., Wang, J., Zhu, Y., Jiang, S., Xiong, Z., Li, Z., Wu, W., et al · 2024
Closest in time.
Can large language models infer causation from correlation?
Jin, Z., Liu, J., LYU, Z., Poff, S., Sachan, M., Mihalcea, R., Diab, M. T., and Schölkopf, B · 2024
Closest in time.
Task contamination: Language models may not be few-shot anymore
Li, C. and Flanigan, J · 2024
Closest in time.
Image content generation with causal reasoning
Li, X., Fan, B., Zhang, R., Jin, L., Wang, D., Guo, Z., Zhao, Y., and Li, R · 2024
Closest in time.
Lu, C., Qian, C., Zheng, G., Fan, H., Gao, H., Zhang, J., Shao, J., Deng, J., Fu, J., Huang, K., et al · 2024
Closest in time.
Self-refine: Iterative refinement with self-feedback
Madaan, A., Tandon, N., Gupta, P., Hallinan, S., Gao, L., Wiegreffe, S., Alon, U., Dziri, N., Prabhumoye, S., Yang, Y., et al · 2024
Closest in time.
Meta llama 3, 2024
Meta · 2024
Closest in time.
Gpt-4v(ision) system card
OpenAI · 2024
Closest in time.
Mathematical discoveries from program search with large language models
Romera-Paredes, B., Barekatain, M., Novikov, A., Balog, M., Kumar, M. P., Dupont, E., Ruiz, F. J., Ellenberg, J. S., Wang, P., Fawzi, O., et al · 2024
Closest in time.
Toolformer: Language models can teach themselves to use tools
Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Hambro, E., Zettlemoyer, L., Cancedda, N., and Scialom, T · 2024
Closest in time.
Prompt injection vs jailbreaking: What is the difference?
Schulhoff, S. V · 2024
Closest in time.
Towards understanding sycophancy in language models
Sharma, M., Tong, M., Korbak, T., Duvenaud, D., Askell, A., Bowman, S. R., DURMUS, E., Hatfield-Dodds, Z., Johnston, S. R., Kravec, S. M., Maxwell, T., McCandlish, S., Ndousse, K., Rausch, O., Schiefer, N., Yan, D., Zhang, M., and Perez, E · 2024
Closest in time.
Solving olympiad geometry without human demonstrations
Trinh, T. H., Wu, Y., Le, Q. V., He, H., and Luong, T · 2024
Closest in time.
Autodev: Automated ai-driven development
Tufano, M., Agarwal, A., Jang, J., Moghaddam, R. Z., and Sundaresan, N · 2024
Closest in time.
Jailbroken: How does llm safety training fail?
Wei, A., Haghtalab, N., and Steinhardt, J · 2024
Closest in time.
Deciphering spatio-temporal graph forecasting: A causal lens and treatment
Xia, Y., Liang, Y., Wen, H., Liu, X., Wang, K., Zhou, Z., and Zimmermann, R · 2024
Closest in time.
Invariant learning via probability of sufficient and necessary causes
Yang, M., Zhang, Y., Fang, Z., Du, Y., Liu, F., Ton, J.-F., Wang, J., and Wang, J · 2024
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T., Cao, Y., and Narasimhan, K · 2024
Closest in time.
KoLA: Carefully benchmarking world knowledge of large language models
Yu, J., Wang, X., Tu, S., Cao, S., Zhang-Li, D., Lv, X., Peng, H., Yao, Z., Zhang, X., Li, H., Li, C., Zhang, Z., Bai, Y., Liu, Y., Xin, A., Yun, K., GONG, L., Lin, N., Chen, J., Wu, Z., Qi, Y., Li, W., Guan, Y., Zeng, K., Qi, J., Jin, H., Liu, J., Gu, Y., Yao, Y., Ding, N., Hou, L., Liu, Z., Bin, X., Tang, J., and Li, J · 2024
Closest in time.
Evaluating and improving tool-augmented computation-intensive math reasoning
Zhang, B., Zhou, K., Wei, X., Zhao, X., Sha, J., Wang, S., and Wen, J.-R · 2024
Closest in time.
Mathattack: Attacking large language models towards math solving ability
Zhou, Z., Wang, Q., Jin, M., Yao, J., Ye, J., Liu, W., Wang, W., Huang, X., and Huang, K · 2024
Closest in time.