Fetching the paper…
Reading the bibliography…
Medical large language models (LLMs) research often makes bold claims, from encoding clinical knowledge to reasoning like a physician.
Technical recommendations for psychological tests and diagnostic techniques
APA · 1954
Earlier work this paper cites.
Construct validity in psychological tests
Cronbach, L. J. and Meehl, P. E · 1955
Earlier work this paper cites.
Taxonomy of educational objectives: The classification of educational goals. Handbook 1: Cognitive domain
Bloom, B. S., Engelhart, M. D., Furst, E. J., Hill, W. H., Krathwohl, D. R., et al · 1956
Earlier work this paper cites.
Convergent and discriminant validation by the multitrait-multimethod matrix
Campbell, D. T. and Fiske, D. W · 1959
Earlier work this paper cites.
An inventory for measuring depression
Beck, A. T., Ward, C. H., Mendelson, M., Mock, J., and Erbaugh, J · 1961
Earlier work this paper cites.
Concurrent and predictive validity designs: A critical reanalysis
Barrett, G. V., Phillips, J. S., and Alexander, R. A · 1981
Earlier work this paper cites.
A note on concurrent and predictive validity designs: A critical reanalysis
Guion, R. M. and Cranny, C · 1982
Earlier work this paper cites.
A review of current assessment tools for monitoring changes in depression
Moran, P. W · 1982
Earlier work this paper cites.
Kiddie schedule for affective disorders and schizophrenia
Puig-Antich, J. and Ryan, N · 1986
Earlier work this paper cites.
Construct validation after thirty years
Cronbach, L. J · 1989
Earlier work this paper cites.
Meaning and values in test validation: The science and ethics of assessment
Messick, S · 1989
Earlier work this paper cites.
The assessment of clinical skills/competence/performance
Miller, G. E · 1990
Earlier work this paper cites.
Concurrent validity and psychometric properties of the beck depression inventory in outpatient adolescents
Ambrosini, P. J., Metz, C., Bianchi, M. D., Rabinovich, H., and Undie, A · 1991
Earlier work this paper cites.
The big five personality dimensions and job performance: a meta-analysis
Barrick, M. R. and Mount, M. K · 1991
Earlier work this paper cites.
Review of validity research on the stanford-binet intelligence scale
Laurent, J., Swerdlik, M., and Ryburn, M · 1992
Earlier work this paper cites.
The mismeasure of man, 1996
Gould, S. J · 1996
Earlier work this paper cites.
Retention of basic science information by fourth-year medical students
Swanson, D. B., Case, S. M., Luecht, R. M., and Dillon, G. F · 1996
Earlier work this paper cites.
Construct validity of the beck depression inventory in a depressive population
Schotte, C., Maes, M., Cluydts, R., De Doncker, D., and Cosyns, P · 1997
Earlier work this paper cites.
Test validity: A matter of consequence
Messick, S · 1998
Earlier work this paper cites.
On the validity of the beck depression inventory: A review
Richter, P., Werner, J., Heerlein, A., Kraus, A., and Sauer, H · 1998
Earlier work this paper cites.
Current concerns in validity theory
Kane, M. T · 2001
Cited alongside, same era.
The unified medical language system (umls): integrating biomedical terminology
Bodenreider, O · 2004
Cited alongside, same era.
Construct validity: Advances in theory and methodology
Strauss, M. E. and Smith, G. T · 2009
Cited alongside, same era.
The “meaningful use” regulation for electronic health records
Blumenthal, D. and Tavenner, M · 2010
Cited alongside, same era.
Mayo clinical text analysis and knowledge extraction system (ctakes): architecture, component evaluation and applications
Savova, G. K., Masanz, J. J., Ogren, P. V., Zheng, J., Sohn, S., Kipper-Schuler, K. C., and Chute, C. G · 2010
Cited alongside, same era.
Beyond diagnostic accuracy: the clinical utility of diagnostic tests
Bossuyt, P. M., Reitsma, J. B., Linnet, K., and Moons, K. G · 2012
Emergent abilities of large language models
Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., et al · 2022
Later among the works it cites.
Systematic review and meta-analysis of ai-based conversational agents for promoting mental health and well-being
Li, H., Zhang, R., Lee, Y.-C., Kraut, R. E., and Mohr, D. C · 2023
Later among the works it cites.
Towards accurate differential diagnosis with large language models
McDuff, D., Schaekermann, M., Tu, T., Palepu, A., Wang, A., Garrison, J., Singhal, K., Sharma, Y., Azizi, S., Kulkarni, K., et al · 2023
Later among the works it cites.
Large language models can be easily distracted by irrelevant context
Shi, F., Chen, X., Misra, K., Scales, N., Dohan, D., Chi, E. H., Schärli, N., and Zhou, D · 2023
Later among the works it cites.
Large language models encode clinical knowledge
Singhal, K., Azizi, S., Tu, T., Mahdavi, S. S., Wei, J., Chung, H. W., Scales, N., Tanwani, A., Cole-Lewis, H., Pfohl, S., et al · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
All validity is construct validity. or is it?
Kane, M · 2012
Cited alongside, same era.
Ets contributions to the quantitative assessment of item, test, and score fairness
Dorans, N. J · 2013
Cited alongside, same era.
The evolution of the united states medical licensing examination (usmle): enhancing assessment of practice-related competencies
Haist, S. A., Katsufrakis, P. J., and Dillon, G. F · 2013
Cited alongside, same era.
A guide on the use of factor analysis in the assessment of construct validity
Kang, H · 2013
Cited alongside, same era.
The relationship between licensing examination performance and the outcomes of care by international medical school graduates
Norcini, J. J., Boulet, J. R., Opalek, A., and Dauphinee, W. D · 2014
Cited alongside, same era.
Electronic health record adoption in us hospitals: progress continues, but challenges persist
Adler-Milstein, J., DesRoches, C. M., Kralovec, P., Foster, G., Worzala, C., Charles, D., Searcy, T., and Jha, A. K · 2015
Cited alongside, same era.
Later among the works it cites.
Superhuman performance of a large language model on the reasoning tasks of a physician
Brodeur, P. G., Buckley, T. A., Kanjee, Z., Goh, E., Ling, E. B., Jain, P., Cabral, S., Abdulnour, R.-E., Haimovich, A., Freed, J. A., et al · 2024
Later among the works it cites.
Data science at the singularity
Donoho, D · 2024
Later among the works it cites.
Does progress on imagenet transfer to real-world datasets?
Fang, A., Kornblith, S., and Schmidt, L · 2024
Later among the works it cites.
Ai-generated clinical summaries require more than accuracy
Goodman, K. E., Paul, H. Y., and Morgan, D. J · 2024
Later among the works it cites.
Evaluation and mitigation of the limitations of large language models in clinical decision-making
Hager, P., Jungmann, F., Holland, R., Bhagat, K., Hubrecht, I., Knauer, M., Vielhauer, J., Makowski, M., Braren, R., Kaissis, G., et al · 2024
Later among the works it cites.
Chatgpt vs. neurologists: a cross-sectional study investigating preference, satisfaction ratings and perceived empathy in responses among people living with multiple sclerosis
Maida, E., Moccia, M., Palladino, R., Borriello, G., Affinito, G., Clerico, M., Repice, A. M., Di Sapio, A., Iodice, R., Spiezia, A. L., et al · 2024
Later among the works it cites.
Evaluating large language models as agents in the clinic
Mehandru, N., Miao, B. Y., Almaraz, E. R., Sushil, M., Butte, A. J., and Alaa, A · 2024
Later among the works it cites.
Ai as a sport: On the competitive epistemologies of benchmarking
Orr, W. and Kang, E. B · 2024
Later among the works it cites.
Open medical llm leaderboard
Pal, A., Minervini, P., Motzfeldt, A. G., and Alex, B · 2024
Later among the works it cites.
The mechanics of frictionless reproducibility
Recht, B · 2024
Later among the works it cites.
Diagnostic reasoning prompts reveal the potential for large language model interpretability in medicine
Savage, T., Nayak, A., Gallo, R., Rangan, E., and Chen, J. H · 2024
Later among the works it cites.
Adapted large language models can outperform medical experts in clinical text summarization
Van Veen, D., Van Uden, C., Blankemeier, L., Delbrouck, J.-B., Aali, A., Bluethgen, C., Pareek, A., Polacin, M., Reis, E. P., Seehofnerová, A., et al · 2024
Later among the works it cites.
An evaluation framework for clinical use of large language models in patient interaction tasks
Johri, S., Jeong, J., Tran, B. A., Schlessinger, D. I., Wongvibulsin, S., Barnes, L. A., Zhou, H.-Y., Cai, Z. R., Van Allen, E. M., Kim, D., et al · 2025
Closest in time.
It’s time to bench the medical exam benchmark, 2025
Raji, I. D., Daneshjou, R., and Alsentzer, E · 2025
Closest in time.
Toward expert-level medical question answering with large language models
Singhal, K., Tu, T., Gottweis, J., Sayres, R., Wulczyn, E., Amin, M., Hou, L., Clark, K., Pfohl, S. R., Cole-Lewis, H., et al · 2025
Closest in time.