Fetching the paper…
Reading the bibliography…
Large AI models, or foundation models, are models recently emerging with massive scales both parameter-wise and data-wise, the magnitudes of which can reach beyond billions.
K. Wüthrich, “The way to nmr structures of proteins,” Nature structural biology , vol. 8, no. 11, pp. 923–925, 2001
2001
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in CVPR . IEEE, 2009, pp. 248–255
2009
Earlier work this paper cites.
J. Yosinski et al. , “How transferable are features in deep neural networks?” NeurIPS , vol. 27, 2014
2014
Earlier work this paper cites.
J. Sohl-Dickstein et al. , “Deep unsupervised learning using nonequilibrium thermodynamics,” in ICML . PMLR, 2015, pp. 2256–2265
2015
Earlier work this paper cites.
X.-C. Bai et al. , “How cryo-em is revolutionizing structural biology,” Trends in biochemical sciences , vol. 40, no. 1, pp. 49–57, 2015
2015
Earlier work this paper cites.
B. E. Suzek et al. , “Uniref clusters: a comprehensive and scalable alternative for improving sequence similarity searches,” Bioinformatics , vol. 31, no. 6, pp. 926–932, 2015
2015
Earlier work this paper cites.
W. J. Hall, M. V. Chapman, K. M. Lee et al. , “Implicit racial/ethnic bias among health care professionals and its influence on health care outcomes: a systematic review,” American journal of public health , vol. 105, no. 12, pp. e60–e76, 2015
2015
Earlier work this paper cites.
V. Gulshan, L. Peng, M. Coram et al. , “Development and validation of a deep learning algorithm for detection of diabetic retinopathy in retinal fundus photographs,” jama , vol. 316, no. 22, pp. 2402–2410, 2016
2016
Earlier work this paper cites.
A. Vaswani et al. , “Attention is all you need,” NeurIPS , vol. 30, 2017
2017
Earlier work this paper cites.
P. F. Christiano, J. Leike, T. Brown et al. , “Deep reinforcement learning from human preferences,” NeurIPS , vol. 30, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
C. Sun et al. , “Revisiting unreasonable effectiveness of data in deep learning era,” in ICCV , 2017, pp. 843–852
2017
Earlier work this paper cites.
R. Shokri, M. Stronati, C. Song, and V. Shmatikov, “Membership inference attacks against machine learning models,” in 2017 IEEE symposium on security and privacy (SP) . IEEE, 2017, pp. 3–18
2017
Earlier work this paper cites.
A. Radford et al. , “Improving language understanding by generative pre-training,” 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
D. Mahajan et al. , “Exploring the limits of weakly supervised pretraining,” in ECCV , 2018, pp. 181–196
2018
Earlier work this paper cites.
J. M. Grimes et al. , “Where is crystallography going?” Acta Crystallographica Section D: Structural Biology , vol. 74, no. 2, pp. 152–166, 2018
2018
Earlier work this paper cites.
M. Steinegger and J. Söding, “Clustering huge protein sequence sets in linear time,” Nature communications , vol. 9, no. 1, p. 2542, 2018
2018
Earlier work this paper cites.
D. S. Char, N. H. Shah, and D. Magnus, “Implementing machine learning in health care—addressing ethical challenges,” The New England journal of medicine , vol. 378, no. 11, p. 981, 2018
2018
Earlier work this paper cites.
E. F. Villaronga, P. Kieseberg, and T. Li, “Humans forget, machines remember: Artificial intelligence and the right to be forgotten,” Computer Law & Security Review , vol. 34, no. 2, pp. 304–313, 2018
2018
Earlier work this paper cites.
M. Tan et al. , “Efficientnet: Rethinking model scaling for convolutional neural networks,” in ICML , 2019, pp. 6105–6114
2019
Earlier work this paper cites.
Y. Huang et al. , “Gpipe: Efficient training of giant neural networks using pipeline parallelism,” NeurIPS , vol. 32, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
E. Alsentzer et al. , “Publicly available clinical bert embeddings,” arXiv , 2019
2019
Earlier work this paper cites.
B. M. Popkin, C. Corvalan, and L. M. Grummer-Strawn, “Dynamics of the double burden of malnutrition and the changing nutrition reality,” The Lancet , 2019
2019
Earlier work this paper cites.
Z. Obermeyer, B. Powers, C. Vogeli, and S. Mullainathan, “Dissecting racial bias in an algorithm used to manage the health of populations,” Science , vol. 366, no. 6464, pp. 447–453, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
M. Chen, A. Radford, R. Child et al. , “Generative Pretraining From Pixels,” in ICML . PMLR, Nov. 2020
2020
Earlier work this paper cites.
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A simple framework for contrastive learning of visual representations,” in ICML , 2020, pp. 1597–1607
2020
Earlier work this paper cites.
A. Kolesnikov et al. , “Big transfer (bit): General visual representation learning,” in ECCV , 2020, pp. 491–507
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
Y. Li, S. Rao, J. R. A. Solares et al. , “Behrt: transformer for electronic health records,” Scientific reports , vol. 10, no. 1, pp. 1–12, 2020
2020
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder et al. , “Language models are few-shot learners,” NeurIPS , vol. 33, pp. 1877–1901, 2020
2020
Earlier work this paper cites.
J. Lee et al. , “Biobert: a pre-trained biomedical language representation model for biomedical text mining,” Bioinformatics , vol. 36, no. 4, pp. 1234–1240, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
E. Chen, K. Lerman, E. Ferrara et al. , “Tracking social media discourse about the covid-19 pandemic: Development of a public coronavirus twitter data set,” JMIR public health and surveillance , vol. 6, no. 2, p. e19273, 2020
2020
Earlier work this paper cites.
D. Cirillo, S. Catuara-Solarz, C. Morey et al. , “Sex and gender differences and biases in artificial intelligence for biomedicine and healthcare,” NPJ digital medicine , vol. 3, no. 1, p. 81, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
R. Bommasani et al. , “On the opportunities and risks of foundation models,” arXiv:2108.07258 , 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
H. Chen, Y. Wang, T. Guo et al. , “Pre-trained image processing transformer,” in CVPR , 2021, pp. 12 299–12 310
2021
Earlier work this paper cites.
A. Dosovitskiy, L. Beyer, A. Kolesnikov et al. , “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,” in ICLR , 2021
2021
Earlier work this paper cites.
K. Han, A. Xiao, E. Wu et al. , “Transformer in Transformer,” in NeurIPS , vol. 34, 2021, pp. 15 908–15 919
2021
Earlier work this paper cites.
Z. Liu et al. , “Swin transformer: Hierarchical vision transformer using shifted windows,” in ICCV , 2021, pp. 10 012–10 022
2021
Earlier work this paper cites.
Z. Dai, H. Liu, Q. V. Le, and M. Tan, “Coatnet: Marrying convolution and attention for all data sizes,” NeurIPS , vol. 34, pp. 3965–3977, 2021
2021
Earlier work this paper cites.
S. d’Ascoli et al. , “Convit: Improving vision transformers with soft convolutional inductive biases,” in ICML . PMLR, 2021, pp. 2286–2296
2021
Earlier work this paper cites.
——, “Efficientnetv2: Smaller models and faster training,” in ICML , 2021, pp. 10 096–10 106
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
A. Radford et al. , “Learning transferable visual models from natural language supervision,” in ICML . PMLR, 2021, pp. 8748–8763
2021
Earlier work this paper cites.
C. Jia, Y. Yang, Y. Xia et al. , “Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision,” in ICML , Jul. 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
H. Pham, Z. Dai, G. Ghiasi et al. , “Combined scaling for open-vocabulary image classification,” arXiv e-prints , pp. arXiv–2111, 2021
2021
Earlier work this paper cites.
A. Ramesh, M. Pavlov, G. Goh et al. , “Zero-Shot Text-to-Image Generation,” in ICML , Jul. 2021
2021
Earlier work this paper cites.
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” 2021
2021
Earlier work this paper cites.
J. Jumper, R. Evans, A. Pritzel et al. , “Highly accurate protein structure prediction with alphafold,” Nature , vol. 596, no. 7873, pp. 583–589, 2021
2021
Earlier work this paper cites.
R. Evans et al. , “Protein complex prediction with alphafold-multimer,” BioRxiv , pp. 2021–10, 2021
2021
Earlier work this paper cites.
A. Elnaggar et al. , “Prottrans: Towards cracking the language of lifes code through self-supervised deep learning and high performance computing,” IEEE TPAMI , vol. PP, pp. 1–1, 07 2021
2021
Earlier work this paper cites.
M. Baek et al. , “Accurate prediction of protein structures and interactions using a three-track neural network,” Science , vol. 373, no. 6557, pp. 871–876, 2021
2021
Earlier work this paper cites.
“Rnacentral 2021: secondary structure integration, improved sequence search and new member databases,” Nucleic acids research , vol. 49, no. D1, pp. D212–D220, 2021
2021
Earlier work this paper cites.
L. Rasmy, Y. Xiang, Z. Xie et al. , “Med-bert: pretrained contextualized embeddings on large-scale structured electronic health records for disease prediction,” NPJ digital medicine , vol. 4, no. 1, p. 86, 2021
2021
Earlier work this paper cites.
L. Rasmy, Y. Xiang, Z. Xie, C. Tao, and D. Zhi, “Med-bert: pretrained contextualized embeddings on large-scale structured electronic health records for disease prediction,” NPJ digital medicine , vol. 4, no. 1, p. 86, 2021
2021
Earlier work this paper cites.
K. raj Kanakarajan, B. Kundumani, and M. Sankarasubbu, “Bioelectra: pretrained biomedical text encoder using discriminators,” in Proceedings of the 20th Workshop on Biomedical Language Processing , 2021, pp. 143–154
2021
Cited alongside, same era.
Y. Gu, R. Tinn, H. Cheng et al. , “Domain-specific language model pretraining for biomedical natural language processing,” HEALTH , vol. 3, no. 1, pp. 1–23, 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
D. M. Korngiebel and S. D. Mooney, “Considering the possibilities and pitfalls of generative pre-trained transformer 3 (gpt-3) in healthcare delivery,” NPJ Digital Medicine , vol. 4, no. 1, p. 93, 2021
2021
Cited alongside, same era.
Y. LeCun, “A path towards autonomous machine intelligence version 0.9. 2, 2022-06-27,” 2022
2022
Later among the works it cites.
S. A. Siddiqui, N. Rajkumar, T. Maharaj et al. , “Metadata archaeology: Unearthing data subsets by leveraging training dynamics,” in ICLR , 2022
2022
Later among the works it cites.
A. Kirillov et al. , “Segment anything,” arXiv preprint arXiv:2304.02643 , 2023
2023
Closest in time.
O. (2023), “Gpt-4 technical report,” arXiv:2303.08774 , 2023
2023
Closest in time.
P. Lee, C. Goldberg, and I. Kohane, The AI Revolution in Medicine: GPT-4 and Beyond . Pearson Education, Limited, 2023
2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
N. Carlini, F. Tramer, E. Wallace, M. Jagielski, A. Herbert-Voss, K. Lee, A. Roberts, T. B. Brown, D. Song, U. Erlingsson et al. , “Extracting training data from large language models.” in USENIX Security Symposium , vol. 6, 2021
2021
Cited alongside, same era.
N. Elhage, N. Nanda, C. Olsson et al. , “A mathematical framework for transformer circuits,” Transformer Circuits Thread , vol. 1, 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
OpenAI, “Chatgpt: Optimizing language models for dialogue,” 2022. [Online]. Available: https://openai.com/blog/chatgpt/
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2023
Closest in time.
2023
Closest in time.
A. Susano Pinto, A. Kolesnikov, Y. Shi et al. , “Tuning computer vision models with task rewards,” arXiv e-prints , pp. arXiv–2302, 2023
2023
Closest in time.
M. Assran, Q. Duval, I. Misra et al. , “Self-supervised learning from images with a joint-embedding predictive architecture,” in CVPR , 2023, pp. 15 619–15 629
2023
Closest in time.
2023
Closest in time.
W. Wang, J. Dai, Z. Chen et al. , “Internimage: Exploring large-scale vision foundation models with deformable convolutions,” in CVPR , 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
R. Girdhar, A. El-Nouby, Z. Liu et al. , “Imagebind: One embedding space to bind them all,” in CVPR , 2023, pp. 15 180–15 190
2023
Closest in time.
2023
Closest in time.
Z. Lin et al. , “Evolutionary-scale prediction of atomic-level protein structure with a language model,” Science , vol. 379, no. 6637, pp. 1123–1130, 2023
2023
Closest in time.
B. Chen, X. Cheng, L. ao Gengyang et al. , “xtrimopglm: Unified 100b-scale pre-trained transformer for deciphering the language of protein,” bioRxiv , 2023
2023
Closest in time.
A. Elnaggar et al. , “Ankh: Optimized protein language model unlocks general-purpose modelling,” bioRxiv , pp. 2023–01, 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Y. Li, Z. Li, K. Zhang et al. , “Chatdoctor: A medical chat model fine-tuned on a large language model meta-ai (llama) using medical domain knowledge,” Cureus , vol. 15, no. 6, 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
A. Vaid, J. Jiang, A. Sawant et al. , “A foundational vision transformer improves diagnostic performance for electrocardiograms,” NPJ Digital Medicine , vol. 6, no. 1, p. 108, 2023
2023
Closest in time.
P. Shi, J. Qiu, S. M. D. Abaxi et al. , “Generalist vision foundation models for medical imaging: A case study of segment anything model on zero-shot medical segmentation,” Diagnostics , vol. 13, no. 11, p. 1947, 2023
2023
Closest in time.
2023
Closest in time.
Z. Huang et al. , “A visual–language foundation model for pathology image analysis using medical twitter,” Nature Medicine , pp. 1–10, 2023
2023
Closest in time.
2023
Closest in time.
“Pubmed abstract,” 2023. [Online]. Available: https://pubmed.ncbi.nlm.nih.gov/download/
2023
Closest in time.
“Pubmed central,” 2023. [Online]. Available: https://www.ncbi.nlm.nih.gov/pmc/
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
S. B. Patel and K. Lam, “Chatgpt: the future of discharge summaries?” The Lancet Digital Health , 2023
2023
Closest in time.
“Fairway health - process prior authorization faster,” 2023. [Online]. Available: https://www.ycombinator.com/launches/IIu-fairway-health-process-prior-authorization-faster
2023
Closest in time.
T. H. Kung, M. Cheatham, A. Medenilla et al. , “Performance of chatgpt on usmle: Potential for ai-assisted medical education using large language models,” PLOS Digital Health , vol. 2, no. 2, p. e0000198, 2023
2023
Closest in time.
E. Shue, L. Liu, B. Li, Z. Feng, X. Li, and G. Hu, “Empowering beginners in bioinformatics with chatgpt,” bioRxiv , pp. 2023–03, 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
J. Qiu et al. , “Egocentric image captioning for privacy-preserved passive dietary intake monitoring,” IEEE Transactions on Cybernetics , 2023
2023
Closest in time.
2023
Closest in time.
K. Bi, L. Xie, H. Zhang et al. , “Accurate medium-range global weather forecasting with 3d neural networks,” Nature , pp. 1–6, 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
S. Vemprala, R. Bonatti, A. Bucker, and A. Kapoor, “Chatgpt for robotics: Design principles and model abilities,” Microsoft, Tech. Rep. MSR-TR-2023-8, February 2023. [Online]. Available: https://www.microsoft.com/en-us/research/publication/chatgpt-for-robotics-design-principles-and-model-abilities/
2023
Closest in time.
2023
Closest in time.
N. Ding, Y. Qin, G. Yang et al. , “Parameter-efficient fine-tuning of large-scale pre-trained language models,” Nature Machine Intelligence , vol. 5, no. 3, pp. 220–235, 2023
2023
Closest in time.
S. Gilbert, H. Harvey, T. Melvin et al. , “Large language model ai chatbots require approval as medical devices,” Nature Medicine , pp. 1–3, 2023
2023
Closest in time.
2023
Closest in time.
P. Lee, S. Bubeck, and J. Petro, “Benefits, limits, and risks of gpt-4 as an ai chatbot for medicine,” New England Journal of Medicine , vol. 388, no. 13, pp. 1233–1239, 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
L. Li and M. W. Spratling, “Data augmentation alone can improve adversarial training,” in ICLR , 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
“How your data is used to improve model performance,” 2023. [Online]. Available: https://help.openai.com/en/articles/5722486-how-your-data-is-used-to-improve-model-performance
2023
Closest in time.
“March 20 chatgpt outage: Here’s what happened,” 2023. [Online]. Available: https://openai.com/blog/march-20-chatgpt-outage
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
P. Liu, W. Yuan, J. Fu et al. , “Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing,” ACM Computing Surveys , vol. 55, no. 9, pp. 1–35, 2023
2023
Closest in time.