Fetching the paper…
Reading the bibliography…
Foundation models trained on patient electronic health records (EHRs) require tokenizing medical data into sequences of discrete vocabulary items.
International classification of diseases—ninth revision (icd-9)
Organization, W. H. et al · 1988
Earlier work this paper cites.
Comorbidities, complications, and coding bias: does the number of diagnosis codes matter in predicting in-hospital mortality?
Foley, S. M., Daley, J., Hughes, J., Fisher, E. S., Heeren, T., et al · 1992
Earlier work this paper cites.
A new drug classification for computer systems: the atc extension code
Miller, G. and Britt, H · 1995
Earlier work this paper cites.
Development of the icd-10 procedure coding system (icd-10-pcs)
Averill, R. F., Mullin, R. L., Steinbeck, B. A., Goldfield, N. I., and Grant, T. M · 2001
Earlier work this paper cites.
The unified medical language system (umls): integrating biomedical terminology
Bodenreider, O · 2004
Earlier work this paper cites.
International Statistical Classification of Diseases and related health problems: Alphabetical index , volume 3
Organization, W. H · 2004
Earlier work this paper cites.
Snomed-ct: The advanced terminology and coding system for ehealth
Donnelly, K. et al · 2006
Earlier work this paper cites.
What is ndc?
Palmer, E · 2006
Earlier work this paper cites.
Normalized names for clinical drugs: Rxnorm at 6 years
Nelson, S. J., Zeng, K., Kilbourne, J., Powell, T., and Moore, R · 2011
Earlier work this paper cites.
Song, X., Salcianu, A., Song, Y., Dopson, D., and Zhou, D · 2012
Earlier work this paper cites.
Cpt® codes: what are they, why are they necessary, and how are they developed?, 2013
Dotson, P · 2013
Earlier work this paper cites.
Risk prediction with electronic health records: The importance of model validation and clinical context
Goldstein, B. A., Navar, A. M., and Pencina, M. J · 2016
Earlier work this paper cites.
MIMIC-III Clinical Database (version 1.4), 2016
Johnson, A., Pollard, T., and Mark, R · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Sennrich, R., Haddow, B., and Birch, A · 2016
Earlier work this paper cites.
Neural discrete representation learning
Van Den Oord, A., Vinyals, O., et al · 2017
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
Kudo, T. and Richardson, J · 2018
Earlier work this paper cites.
Drugbank 5.0: a major update to the drugbank database for 2018
Wishart, D. S., Feunang, Y. D., Guo, A. C., Lo, E. J., Marcu, A., Grant, J. R., Sajed, T., Johnson, D., Li, C., Sayeeda, Z., et al · 2018
Earlier work this paper cites.
vq-wav2vec: Self-supervised learning of discrete speech representations
Baevski, A., Schneider, S., and Auli, M · 2019
Earlier work this paper cites.
Multitask learning and benchmarking with clinical time series data
Harutyunyan, H., Khachatrian, H., Kale, D. C., Ver Steeg, G., and Galstyan, A · 2019
Earlier work this paper cites.
The new international classification of diseases 11th edition: a comparative analysis with icd-10 and icd-10-cm
Fung, K. W., Xu, J., and Bodenreider, O · 2020
Earlier work this paper cites.
Behrt: transformer for electronic health records
Li, Y., Rao, S., Solares, J. R. A., Hassaine, A., Ramakrishnan, R., Canoy, D., Zhu, Y., Rahimi, K., and Salimi-Khorshidi, G · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2021
Cited alongside, same era.
Vector-quantized image modeling with improved vqgan
Yu, J., Li, X., Koh, J. Y., Zhang, H., Pang, R., Qin, J., Ku, A., Xu, Y., Baldridge, J., and Wu, Y · 2021
Cited alongside, same era.
Soundstream: An end-to-end neural audio codec
Zeghidour, N., Luebs, A., Omran, A., Skoglund, J., and Tagliasacchi, M · 2021
Cited alongside, same era.
Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering
Pal, A., Umapathi, L. K., and Sankarasubbu, M · 2022
Cited alongside, same era.
Vector-quantized image modeling with improved vqgan
Yu, J., Li, X., Koh, J. Y., Zhang, H., Pang, R., Qin, J., Ku, A., Xu, Y., Baldridge, J., and Wu, Y · 2022
Cited alongside, same era.
Digital twins for health: a scoping review
Katsoulakis, E., Wang, Q., Wu, H., Shahriyari, L., Fletcher, R., Liu, J., Achenie, L., Liu, H., Jackson, P., Xiao, Y., Syeda-Mahmood, T., Tuli, R., and Deng, J · 2024
Later among the works it cites.
Foresight—a generative pretrained transformer for modelling of patient timelines using electronic health records: a retrospective modelling study
Kraljevic, Z., Bean, D., Shek, A., Bendayan, R., Hemingway, H., Yeung, J. A., Deng, A., Balston, A., Ross, J., Idowu, E., Teo, J. T., and Dobson, R. J. B · 2024
Later among the works it cites.
Minixhofer, B., Ponti, E. M., and Vulić, I · 2024
Later among the works it cites.
Afrimed-qa: A pan-african, multi-specialty, medical question-answering benchmark dataset
Olatunji, T., Nimo, C., Owodunni, A., Abdullahi, T., Ayodele, E., Sanni, M., Aka, C., Omofoye, F., Yuehgoh, F., Faniran, T., et al · 2024
Later among the works it cites.
Let your graph do the talking: Encoding structured data for llms
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
ibot: Image bert pre-training with online tokenizer
Zhou, J., Wei, C., Wang, H., Shen, W., Xie, C., Yuille, A., and Kong, T · 2022
Cited alongside, same era.
Mondo: Unifying diseases for the world, by the world
Balsa-Canto, E., Brush, M. H., Carbon, S., et al · 2023
Cited alongside, same era.
Building a knowledge graph to enable precision medicine
Chandak, P., Huang, K., and Zitnik, M · 2023
Cited alongside, same era.
Recommender systems with generative retrieval
Rajput, S., Mehta, N., Singh, A., Hulikal Keshavan, R., Vu, T., Heldt, L., Hong, L., Tay, Y., Tran, V., Samost, J., et al · 2023
Cited alongside, same era.
Large language models encode clinical knowledge
Singhal, K., Azizi, S., Tu, T., Mahdavi, S. S., Wei, J., Chung, H. W., Scales, N., Tanwani, A., Cole-Lewis, H., Pfohl, S., et al · 2023
Cited alongside, same era.
EHRSHOT: An EHR benchmark for few-shot evaluation of foundation models
Wornow, M., Thapa, R., Steinberg, E., Fries, J. A., and Shah, N · 2023
Cited alongside, same era.
Language model beats diffusion–tokenizer is key to visual generation
Yu, L., Lezama, J., Gundavarapu, N. B., Versari, L., Sohn, K., Minnen, D., Cheng, Y., Birodkar, V., Gupta, A., Gu, X., et al · 2023
Cited alongside, same era.
Perozzi, B., Fatemi, B., Zelle, D., Tsitsulin, A., Kazemi, M., Al-Rfou, R., and Halcrow, J · 2024
Later among the works it cites.
Graph transformers on ehrs: Better representation improves downstream performance
Poulain, R. and Beheshti, R · 2024
Later among the works it cites.
Model decides how to tokenize: Adaptive dna sequence tokenization with mxdna
Qiao, L., Ye, P., Ren, Y., Bai, W., Liang, C., Ma, X., Dong, N., and Ouyang, W · 2024
Later among the works it cites.
Towards building multilingual language model for medicine
Qiu, P., Wu, C., Zhang, X., Lin, W., Wang, H., Zhang, Y., Wang, Y., and Xie, W · 2024
Later among the works it cites.
Tokenflow: Unified image tokenizer for multimodal understanding and generation
Qu, L., Zhang, H., Liu, Y., Wang, X., Jiang, Y., Gao, Y., Ye, H., Du, D. K., Yuan, Z., and Wu, X · 2024
Later among the works it cites.
Zero shot health trajectory prediction using transformer
Renc, P., Jia, Y., Samir, A. E., Was, J., Li, Q., Bates, D. W., and Sitek, A · 2024
Later among the works it cites.
Knowledge graph based agent for complex, knowledge-intensive qa in medicine
Su, X., Wang, Y., Gao, S., Liu, X., Giunchiglia, V., Clevert, D.-A., and Zitnik, M · 2024
Later among the works it cites.
Learning to tokenize for generative retrieval
Sun, W., Yan, L., Chen, Z., Wang, S., Zhu, H., Ren, P., Chen, Z., Yin, D., Rijke, M., and Ren, Z · 2024
Later among the works it cites.
Birna-bert allows efficient rna language modeling with adaptive tokenization
Tahmid, M. T., Shahgir, H. S., Mahbub, S., Dong, Y., and Bayzid, M. S · 2024
Later among the works it cites.
Towards generalist biomedical ai
Tu, T., Azizi, S., Driess, D., Schaekermann, M., Amin, M., Chang, P.-C., Carroll, A., Lau, C., Tanno, R., Ktena, I., et al · 2024
Later among the works it cites.
RAM-EHR: Retrieval augmentation meets clinical predictions on electronic health records
Xu, R., Shi, W., Yu, Y., Zhuang, Y., Jin, B., Wang, M. D., Ho, J., and Yang, C · 2024
Later among the works it cites.
Predict and interpret health risk using ehr through typical patients
Yu, Z., Zhang, C., Wang, Y., Tang, W., Wang, J., and Ma, L · 2024
Later among the works it cites.
Language-guided image tokenization for generation
Zha, K., Yu, L., Fathi, A., Ross, D. A., Schmid, C., Katabi, D., and Gu, X · 2024
Later among the works it cites.
Emerge: Enhancing multimodal electronic health records predictive modeling with retrieval-augmented generation
Zhu, Y., Ren, C., Wang, Z., Zheng, X., Xie, S., Feng, J., Zhu, X., Li, Z., Ma, L., and Pan, C · 2024
Later among the works it cites.
Cosmos world foundation model platform for physical ai
Agarwal, N., Ali, A., Bala, M., Balaji, Y., Barker, E., Cai, T., Chattopadhyay, P., Chen, Y., Cui, Y., Ding, Y., et al · 2025
Closest in time.
Toward expert-level medical question answering with large language models
Singhal, K., Tu, T., Gottweis, J., Sayres, R., Wulczyn, E., Amin, M., Hou, L., Clark, K., Pfohl, S. R., Cole-Lewis, H., et al · 2025
Closest in time.
Analysis of free text in electronic health records for identification of cancer patient trajectories
Jensen, K., Soguero-Ruiz, C., Oyvind Mikalsen, K., Lindsetmo, R.-O., Kouskoumvekaki, I., Girolami, M., Olav Skrovseth, S., and Augestad, K. M · 2045
Closest in time.