Skip to main navigation Skip to search Skip to main content

A natural language processing pipeline for identifying pediatric long COVID symptoms and functional impacts in freeform clinical notes: a RECOVER study

  • RECOVER Consortium
  • Applied Clinical Research Center
  • Ann & Robert H. Lurie Children's Hospital
  • School of Medicine, Stanford University
  • Seattle Children's
  • Stanford University
  • Abigail Wexner Research Institute
  • O'Donnell School of Public Health
  • Seattle Children’s Research Institute
  • University of Colorado School of Medicine and Children's Hospital Colorado
  • Department of Research
  • The Children's Hospital of Philadelphia
  • Cincinnati Children's Hospital Medical Center
  • Department of Biostatistics at the Medical College of Wisconsin
  • RECOVER Patient

Research output: Contribution to journalArticlepeer-review

1 Scopus citations

Abstract

OBJECTIVE: To develop a natural language processing (NLP) pipeline for unstructured electronic health record (EHR) data to identify symptoms and functional impacts associated with Long COVID in children.

MATERIALS AND METHODS: We analyzed 48 287 outpatient progress notes from 10 618 pediatric patients from 12 institutions. We evaluated notes obtained 28 to 179 days after a COVID-19 diagnosis or positive test. Two samples were examined: patients with evidence of Long COVID and patients with acute COVID but no evidence of Long COVID based on diagnostic codes. The pipeline identified clinical concepts associated with 21 symptoms and 4 functional impact categories. Subject matter experts (SMEs) screened a sample of 4586 terms from the NLP output to assess pipeline accuracy. Prevalence and concordance of each of the 25 concepts was compared between the 2 patient samples.

RESULTS: A binary assertion measure comparing SME and NLP assertions showed moderate accuracy (N = 4133; F1 = .80) and improved substantially when only high-confidence SME assertions were considered (N = 2043; F1 = .90). Overall, the 25 Long COVID concept categories were markedly more prevalent in the presumptive Long COVID cohort, and differences were noted between concepts identified in notes versus structured data.

DISCUSSION: This preliminary analysis illustrates the additional insight into a syndrome such as Long COVID gained from incorporating notes data, characterizing symptoms and functional impacts.

CONCLUSION: These data support the importance of incorporating NLP methodology when possible into designing computable phenotypes and to accurately characterize patients with Long COVID.

Original languageEnglish
Article numberooaf089
Pages (from-to)ooaf089
JournalJAMIA Open
Volume8
Issue number5
StatePublished - 1 Oct 2025

Keywords

  • NLP
  • PEDSnet
  • RECOVER
  • pediatrics

Fingerprint

Dive into the research topics of 'A natural language processing pipeline for identifying pediatric long COVID symptoms and functional impacts in freeform clinical notes: a RECOVER study'. Together they form a unique fingerprint.

Cite this