Fetching the paper…
Reading the bibliography…
We introduce LArge Model Based Data Agent (LAMBDA), a novel open-source, code-free multi-agent data analysis system that leverages the power of large language models.
Framingham heart study dataset
FHS (1948) · 1948
Earlier work this paper cites.
Heart Disease
Janosi, A., Steinbrunn, W., Pfisterer, M., and Detrano, R. (1988) · 1988
Earlier work this paper cites.
The Wine dataset
Aeberhard, S. and Forina, M. (1991) · 1991
Earlier work this paper cites.
Breast Cancer Wisconsin (Diagnostic)
Wolberg, W., Mangasarian, O., Street, N., and Street, W. (1995) · 1995
Earlier work this paper cites.
A trial comparing nucleoside monotherapy with combination therapy in hiv-infected adults with cd4 cell counts from 200 to 500 per cubic millimeter. aids clinical trials group study 175 study team
Hammer, S. M., Katzenstein, D. A., Hughes, M. D., Gundacker, H., Schooley, R. T., Haubrich, R. H., et al. (1996) · 1996
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P. (1998) · 1998
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., et al. (2020) · 2005
Earlier work this paper cites.
A quadratically convergent newton method for computing the nearest correlation matrix
Qi, H. and Sun, D. (2006) · 2006
Earlier work this paper cites.
Concrete Compressive Strength
Yeh, I.-C. (2007) · 2007
Earlier work this paper cites.
Contributions to the study of sms spam filtering: New collection and results
Almeida, T. A., Hidalgo, J. M. G., and Yamakami, A. (2011) · 2011
Earlier work this paper cites.
Angiogenic mrna and microrna gene expression signature predicts a novel subtype of serous ovarian cancer
Bentink, S., Haibe-Kains, B., Risch, T., Fan, J. B., Hirsch, M. S., Holton, K., Rubio, R., April, C., Chen, J., Wang, J., Lu, Y., Wickham-Garcia, E., Liu, J., Culhane, A. C., Drapkin, R., Quackenbush, J., and Birrer, M. J. (2012) · 2012
Earlier work this paper cites.
Validating the impact of a molecular subtype in ovarian cancer on outcomes: a study of the ovcad consortium
Pils, D., Hager, G., Tong, D., Aust, S., et al. (2012) · 2012
Earlier work this paper cites.
Airfoil Self-Noise
Brooks, T., Pope, D., and Marcolini, M. (2014) · 2014
Earlier work this paper cites.
National Health and Nutrition Health Survey 2013-2014 (NHANES) Age Prediction Subset
Dinh, A., Miertschin, S., Young, A., and Mohanty, S. D. (2023) · 2014
Earlier work this paper cites.
Combined Cycle Power Plant
Tfekci, P. and Kaya, H. (2014) · 2014
Earlier work this paper cites.
Tcgabiolinks: An r/bioconductor package for integrative analysis of tcga data
Colaprico, A., Silva, T. C., Olsen, C., Garofano, L., Cava, C., Garolini, D., Sabedot, T. S., and et al. (2015) · 2015
Earlier work this paper cites.
Deep learning
LeCun, Y., Bengio, Y., and Hinton, G. (2015) · 2015
Earlier work this paper cites.
Sas/stat® 14.1 user’s guide
SAS Institute Inc. (2015) · 2015
Earlier work this paper cites.
Student admission dataset
Kaggle SAD (2016) · 2016
Cited alongside, same era.
Reinventing biostatistics education for basic scientists
Weissgerber, T. L., Garovic, V. D., Milin-Lazovic, J. S., Winham, S. J., Obradovic, Z., Trzeciakowski, J. P., and Milic, N. M. (2016) · 2016
Cited alongside, same era.
Data science: the impact of statistics
Weihs, C. and Ickstadt, K. (2018) · 2018
Cited alongside, same era.
Sentence-bert: Sentence embeddings using siamese bert-networks
Reimers, N. and Gurevych, I. (2019) · 2019
Cited alongside, same era.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., et al. (2020) · 2020
Cited alongside, same era.
Facilitating knowledge sharing from domain experts to data scientists for building nlp models
Park, S., Wang, A. Y., Kawas, B., Liao, Q. V., Piorkowski, D., and Danilevsky, M. (2021) · 2021
Augmented language models: a survey
Mialon, G., Dessì, R., Lomeli, M., Nalmpantis, C., Pasunuru, R., Raileanu, R., Rozière, B., Schick, T., Dwivedi-Yu, J., Celikyilmaz, A., et al. (2023) · 2023
Later among the works it cites.
OpenAI (2023) · 2023
Later among the works it cites.
Python: A programming language
Python Software Foundation (2023) · 2023
Later among the works it cites.
Taskweaver: A code-first agent framework
Qiao, B., Li, L., Zhang, X., He, S., Kang, Y., Zhang, C., et al. (2023) · 2023
Later among the works it cites.
R: A language and environment for statistical computing
R Core Team (2023) · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Fixed points of nonnegative neural networks
Piotrowski, T. J., Cavalcante, R. L., and Gabor, M. (2021) · 2021
Cited alongside, same era.
Improving language models by retrieving from trillions of tokens
Borgeaud, S., Mensch, A., Hoffmann, J., Cai, T., Rutherford, E., Millican, K., et al. (2022) · 2022
Cited alongside, same era.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., et al. (2022) · 2022
Cited alongside, same era.
A review of some techniques for inclusion of domain-knowledge into deep neural networks
Dash, T., Chitlangia, S., Ahuja, A., and Srinivasan, A. (2022) · 2022
Cited alongside, same era.
Glm: General language model pretraining with autoregressive blank infilling
Du, Z., Qian, Y., Liu, X., Ding, M., Qiu, J., Yang, Z., and Tang, J. (2022) · 2022
Cited alongside, same era.
Three high-dimensional genomic datasets
Anh, P. (2023) · 2023
Cited alongside, same era.
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., et al. (2023) · 2023
Later among the works it cites.
Autogen: Enabling next-gen llm applications via multi-agent conversation
Wu, Q., Bansal, G., Zhang, J., Wu, Y., Li, B., Zhu, E., et al. (2023) · 2023
Later among the works it cites.
Mlcopilot: Unleashing the power of large language models in solving machine learning tasks
Zhang, L., Zhang, Y., Ren, K., Li, D., and Yang, Y. (2023) · 2023
Later among the works it cites.
Agents: An open-source framework for autonomous language agents
Zhou, W., Jiang, Y. E., Li, L., Wu, J., Wang, T., Qiu, S., et al. (2023) · 2023
Later among the works it cites.
Ethical concerns around privacy and data security in ai health monitoring for parkinson’s disease: Insights from patients, family members, and healthcare professionals
Bavli, I., Ho, A., Mahal, R., and McKeown, M. J. (2024) · 2024
Closest in time.
Spider2-v: How far are multimodal agents from automating data science and engineering workflows?
Cao, R., Lei, F., Wu, H., Chen, J., Fu, Y., Gao, H., Xiong, X., Zhang, H., Hu, W., Mao, Y., Xie, T., Xu, H., Zhang, D., Wang, S., Sun, R., Yin, P., Xiong, C., Ni, A., Liu, Q., Zhong, V., Chen, L., Yu, K., and Yu, T. (2024) · 2024
Closest in time.
Large language model based multi-agents: A survey of progress and challenges
Guo, T., Chen, X., Wang, Y., Chang, R., Pei, S., Chawla, N. V., Wiest, O., and Zhang, X. (2024) · 2024
Closest in time.
Data interpreter: An llm agent for data science
Hong, S., Lin, Y., Liu, B., Wu, B., Li, D., Chen, J., et al. (2024) · 2024
Closest in time.
Building domain-specific machine learning workflows: A conceptual framework for the state of the practice
Oakes, B. J., Famelis, M., and Sahraoui, H. (2024) · 2024
Closest in time.
Pami: An open-source python library for pattern mining
Rage, U. K., Pamalla, V., Toyoda, M., and Kitsuregawa, M. (2024) · 2024
Closest in time.
A survey on large language model-based agents for statistics and data science
Sun, M., Han, R., Jiang, B., Qi, H., Sun, D., Yuan, Y., and Huang, J. (2024) · 2024
Closest in time.
What should data science education do with large language models?
Tu, X., Zou, J., Su, W., and Zhang, L. (2024) · 2024
Closest in time.
Raft: Adapting language model to domain specific rag
Zhang, T., Patil, S. G., Jain, N., Shen, S., Zaharia, M., Stoica, I., and Gonzalez, J. E. (2024) · 2024
Closest in time.