2022

TabText: Language-Based Representations of Tabular Health Data for Predictive Modelling

Carballo, Kimberly Villalobos, Na, Liangyuan, Ma, Yu et al.

Understand

Tabular medical records remain the most readily available data format for applying machine learning in healthcare.

  • However, traditional data preprocessing ignores valuable contextual information in tables and requires substantial manual cleaning and harmonisation, creating a bottleneck for model development.
  • We introduce TabText, a preprocessing and feature extraction method that leverages contextual information and streamlines the curation of tabular medical data.
  • This method converts tables into contextual language and applies pretrained large language models (LLMs) to generate task-independent numerical representations.

Reading the bibliography…