Fetching the paper…
Reading the bibliography…
Foundation models have brought changes to the landscape of machine learning, demonstrating sparks of human-level intelligence across a diverse array of tasks.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020) · 1901
Earlier work this paper cites.
Generating long sequences with sparse transformers
Child, R., Gray, S., Radford, A., and Sutskever, I. (2019) · 1904
Earlier work this paper cites.
Gradient-based neural dag learning
Lachapelle, S., Brouillard, P., Deleu, T., and Lacoste-Julien, S. (2019) · 1906
Earlier work this paper cites.
Comment: Neyman (1923) and causal inference in experiments and observational studies
Rubin, D. B. (1990) · 1923
Earlier work this paper cites.
On the evolution of random graphs
Erdős, P. and Rényi, A. (1960) · 1960
Earlier work this paper cites.
The central role of the propensity score in observational studies for causal effects
Rosenbaum, P. R. and Rubin, D. B. (1983) · 1983
Earlier work this paper cites.
Decision making and postdecision surprises
Harrison, J. R. and March, J. G. (1984) · 1984
Earlier work this paper cites.
Statistics and causal inference
Holland, P. W. (1986) · 1986
Earlier work this paper cites.
Evaluating the econometric evaluations of training programs with experimental data
LaLonde, R. J. (1986) · 1986
Earlier work this paper cites.
Model-based direct adjustment
Rosenbaum, P. R. (1987) · 1987
Earlier work this paper cites.
On the strong universal consistency of nearest neighbor regression function estimates
Devroye, L., Gyorfi, L., Krzyzak, A., and Lugosi, G. (1994) · 1994
Earlier work this paper cites.
Infant mortality statistics from the 1996 period linked birth/infant death data set
MacDorman, M. F. and Atkinson, J. O. (1998) · 1996
Earlier work this paper cites.
Causal effects in nonexperimental studies: Reevaluating the evaluation of training programs
Dehejia, R. H. and Wahba, S. (1999) · 1999
Earlier work this paper cites.
Reformer: The efficient transformer
Kitaev, N., Kaiser, Ł., and Levskaya, A. (2020) · 2001
Earlier work this paper cites.
Convex optimization
Boyd, S. P. and Vandenberghe, L. (2004) · 2004
Earlier work this paper cites.
Nonparametric estimation of average treatment effects under exogeneity: A review
Imbens, G. W. (2004) · 2004
Earlier work this paper cites.
The costs of low birth weight
Almond, D., Chay, K. Y., and Lee, D. S. (2005) · 2005
Earlier work this paper cites.
Comment: Performance of double-robust estimators when” inverse probability” weights are highly variable
Robins, J., Sued, M., Lei-Gomez, Q., and Rotnitzky, A. (2007) · 2007
Earlier work this paper cites.
The prognostic analogue of the propensity score
Hansen, B. B. (2008) · 2008
Earlier work this paper cites.
Evaluating uses of data mining techniques in propensity score estimation: a simulation study
Setoguchi, S., Schneeweiss, S., Brookhart, M. A., Glynn, R. J., and Cook, E. F. (2008) · 2008
Earlier work this paper cites.
The elements of statistical learning: data mining, inference, and prediction
Hastie, T., Tibshirani, R., Friedman, J. H., and Friedman, J. H. (2009) · 2009
Earlier work this paper cites.
Nonparametric estimation of conditional expectation
Li, J. and Tran, L. T. (2009) · 2009
Earlier work this paper cites.
Causal inference in statistics: An overview
Pearl, J. (2009) · 2009
Earlier work this paper cites.
Improving propensity score weighting using machine learning
Lee, B. K., Lessler, J., and Stuart, E. A. (2010) · 2010
Earlier work this paper cites.
Doubly robust policy evaluation and learning
Dudík, M., Langford, J., and Li, L. (2011) · 2011
Earlier work this paper cites.
Bayesian nonparametric modeling for causal inference
Hill, J. L. (2011) · 2011
Earlier work this paper cites.
Realcause: Realistic causal inference benchmarking
Neal, B., Huang, C.-W., and Raghupathi, S. (2020) · 2011
Earlier work this paper cites.
Scikit-learn: Machine learning in python
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., et al. (2011) · 2011
Earlier work this paper cites.
Entropy balancing for causal effects: A multivariate reweighting method to produce balanced samples in observational studies
Hainmueller, J. (2012) · 2012
Earlier work this paper cites.
New evidence on the finite sample properties of propensity score reweighting and matching estimators
Busso, M., DiNardo, J., and McCrary, J. (2014) · 2014
Earlier work this paper cites.
Covariate balancing propensity score
Imai, K. and Ratkovic, M. (2014) · 2014
Cited alongside, same era.
Causal inference in statistics, social, and biomedical sciences
Imbens, G. W. and Rubin, D. B. (2015) · 2015
Cited alongside, same era.
Npci: Non-parametrics for causal inference
Dorie, V. (2016) · 2016
Cited alongside, same era.
Learning representations for counterfactual inference
Johansson, F., Shalit, U., and Sontag, D. (2016) · 2016
Cited alongside, same era.
Bayesian inference of individualized treatment effects using multi-task gaussian processes
Alaa, A. M. and Van Der Schaar, M. (2017) · 2017
Cited alongside, same era.
Deep counterfactual networks with propensity-dropout
Alaa, A. M., Weisz, M., and Van Der Schaar, M. (2017) · 2017
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al. (2021) · 2021
Later among the works it cites.
Estimating average treatment effects with support vector machines
Tarr, A. and Imai, K. (2021) · 2021
Later among the works it cites.
The causal-neural connection: Expressiveness, learnability, and inference
Xia, K., Lee, K.-Z., Bengio, Y., and Bareinboim, E. (2021) · 2021
Later among the works it cites.
27 on pearl’s hierarchy and the foundations of causal inference
Bareinboim, E., Correa, J. D., Ibeling, D., and Icard, T. (2022) · 2022
Later among the works it cites.
Riesznet and forestriesz: Automatic debiased machine learning with neural nets and random forests
Chernozhukov, V., Newey, W., Quintas-Martınez, V. M., and Syrgkanis, V. (2022) · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Causalgan: Learning causal implicit generative models with adversarial training
Kocaoglu, M., Snyder, C., Dimakis, A. G., and Vishwanath, S. (2017) · 2017
Cited alongside, same era.
Causal effect inference with deep latent-variable models
Louizos, C., Shalit, U., Mooij, J. M., Sontag, D., Zemel, R., and Welling, M. (2017) · 2017
Cited alongside, same era.
Estimating individual treatment effect: generalization bounds and algorithms
Shalit, U., Johansson, F. D., and Sontag, D. (2017) · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017) · 2017
Cited alongside, same era.
Double/debiased machine learning for treatment and structural parameters
Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C., Newey, W., and Robins, J. (2018) · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2018) · 2018
Cited alongside, same era.
Causal inference in natural language processing: Estimation, prediction, interpretation and beyond
Feder, A., Keith, K. A., Manzoor, E., Pryzant, R., Sridhar, D., Wood-Doughty, Z., Eisenstein, J., Grimmer, J., Reichart, R., Roberts, M. E., et al. (2022) · 2022
Later among the works it cites.
Deep end-to-end causal inference
Geffner, T., Antoran, J., Foster, A., Gong, W., Ma, C., Kiciman, E., Sharma, A., Lamb, A., Kukla, M., Pawlowski, N., et al. (2022) · 2022
Later among the works it cites.
Emergent world representations: Exploring a sequence model trained on a synthetic task
Li, K., Hopkins, A. K., Bau, D., Viégas, F., Pfister, H., and Wattenberg, M. (2022) · 2022
Later among the works it cites.
Empirical analysis of model selection for heterogenous causal effect estimation
Mahajan, D., Mitliagkas, I., Neal, B., and Syrgkanis, V. (2022) · 2022
Later among the works it cites.
Causal transformer for estimating counterfactual outcomes
Melnychuk, V., Frauen, D., and Feuerriegel, S. (2022) · 2022
Later among the works it cites.
A primal-dual framework for transformers and neural networks
Nguyen, T. M., Nguyen, T. M., Ho, N., Bertozzi, A. L., Baraniuk, R., and Osher, S. (2022) · 2022
Later among the works it cites.
Neural causal models for counterfactual identification and estimation
Xia, K., Pan, Y., and Bareinboim, E. (2022) · 2022
Later among the works it cites.
Ban, T., Chen, L., Wang, X., and Chen, H. (2023) · 2023
Closest in time.
Sparks of artificial general intelligence: Early experiments with gpt-4
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., et al. (2023) · 2023
Closest in time.
A decoder-only foundation model for time-series forecasting
Das, A., Kong, W., Sen, R., and Zhou, Y. (2023) · 2023
Closest in time.
Towards foundation models for knowledge graph reasoning
Galkin, M., Yuan, X., Mostafa, H., Tang, J., and Zhu, Z. (2023) · 2023
Closest in time.
Can large language models infer causation from correlation?
Jin, Z., Liu, J., Lyu, Z., Poff, S., Sachan, M., Mihalcea, R., Diab, M., and Schölkopf, B. (2023) · 2023
Closest in time.
Causal reasoning and large language models: Opening a new frontier for causality
Kıcıman, E., Ness, R., Sharma, A., and Tan, C. (2023) · 2023
Closest in time.
Dissociating language and thought in large language models: a cognitive perspective
Mahowald, K., Ivanova, A. A., Blank, I. A., Kanwisher, N., Tenenbaum, J. B., and Fedorenko, E. (2023) · 2023
Closest in time.
Nilforoshan, H., Moor, M., Roohani, Y., Chen, Y., Šurina, A., Yasunaga, M., Oblak, S., and Leskovec, J. (2023) · 2023
Closest in time.
Gpt-4 technical report
OpenAI (2023) · 2023
Closest in time.
The linear representation hypothesis and the geometry of large language models
Park, K., Choe, Y. J., and Veitch, V. (2023) · 2023
Closest in time.
Retentive network: A successor to transformer for large language models
Sun, Y., Dong, L., Huang, S., Ma, S., Xia, Y., Xue, J., Wang, J., and Wei, F. (2023) · 2023
Closest in time.
Transformers as support vector machines
Tarzanagh, D. A., Li, Y., Thrampoulidis, C., and Oymak, S. (2023) · 2023
Closest in time.
Transfer learning enables predictions in network biology
Theodoris, C. V., Xiao, L., Chopra, A., Chaffin, M. D., Al Sayed, Z. R., Hill, M. C., Mantineo, H., Brydon, E. M., Zeng, Z., Liu, X. S., et al. (2023) · 2023
Closest in time.
Causal-discovery performance of chatgpt in the context of neuropathic pain diagnosis
Tu, R., Ma, C., and Zhang, C. (2023) · 2023
Closest in time.
Wolfram— alpha as the way to bring computational knowledge superpowers to chatgpt
Wolfram, S. (2023) · 2023
Closest in time.
Causal parrots: Large language models may talk causality but are not causal
Zečević, M., Willig, M., Dhami, D. S., and Kersting, K. (2023) · 2023
Closest in time.