Background: Much of the symptom burden in colorectal cancer (CRC) patients is documented in unstructured discharge-note narrative, and manual extraction is not scalable.
Whether large language models (LLMs) outperform rule-based and named entity recognition (NER) methods has not been rigorously benchmarked. Objective: To benchmark rule-based, NER, and zero-shot LLM methods for extracting 46 cancer-related symptoms from CRC discharge notes against an adjudicated ground truth. Methods: We analyzed 2,704 discharge notes from CRC patients in MIMIC-IV. A 46-symptom target list was built from the Memorial Symptom Assessment Scale and the EORTC QLQ-CR29.
Four approaches -- dictionary-based rule matching, pretrained clinical NER, and zero-shot Claude Haiku and Gemini 3.5 Flash -- plus two hybrid variants (LLM output with post-hoc rule-based negation filtering) were evaluated against a 200-note gold standard adjudicated by two raters (pooled kappa=0.71, macro kappa=0.49), using Macro/Micro F1, precision, and recall. Results: Gemini 3.5 Flash performed best (Macro F1=0.70, Micro F1=0.86, Macro Precision=0.74), followed by Claude Haiku (Macro F1=0.63, Macro Recall=0.71); both substantially outperformed rule-based (Macro F1=0.44) and NER (Macro F1=0.38) methods.
Post-hoc negation filtering paradoxically degraded LLM performance (Gemini+Hybrid Macro F1=0.58; Claude+Hybrid Macro F1=0.54) by overriding correct predictions through rigid, fixed-window matching. Conclusions: Zero-shot LLMs substantially outperform rule-based and NER approaches for CRC symptom extraction; post-hoc negation correction should not be applied to LLM outputs without syntactic scope validation. Implications for Practice: Zero-shot LLM extraction offers a scalable, accurate alternative to manual chart review and traditional NLP pipelines for oncology symptom surveillance, without institution-specific rule development or model training.
Inicia sesión o regístrate para acceder al texto completo
¡Aún no hay comentarios. Sé el primero en comentar!