Multimodal Large Language Models vs. Medical Doctors in Degenerative Lumbar Spine Surgery: A Retrospective Decision Concordance Study of 147 Patients
OBJECTIVE: To evaluate decision concordance between commercially available multimodal large language models (LLMs), resident doctors, and senior-surgeon ground truth for surgical indication and spinal level in degenerative lumbar spine disease.
METHODS: Retrospective analysis of 147 consecutive patients (76 operative, 71 conservative). Each case included clinical documentation and MRI presented as two composite PNG images (axial and sagittal overviews). Two resident doctors and three frontier multimodal LLMs (GPT 5.5, Claude Sonnet 4.6, Gemini 3.1 Pro) independently assessed operative versus conservative management and, if operative, the exact surgical level. Analyses utilized Cochran’s Q, McNemar tests with Holm-Bonferroni correction, and analytical Bayesian Beta-Binomial models (95% Credible Intervals).
RESULTS & PRINCIPAL FINDING: LLMs achieved higher overall therapy-decision accuracy (66.0%–68.0%; 97–100/147) than residents (54.4%; 80/147) but over-recommended surgery. Conditional level accuracy when surgery was correctly indicated was 71.4% (20/28) for resident doctors versus only 33.3%–41.1% for LLMs. Bayesian analysis confirmed with >99% posterior probability that medical doctors maintain clear spatial and anatomical superiority.
CONCLUSION: Off-the-shelf multimodal LLMs approximate human performance for binary surgical indication but remain inferior for precise level localization. These results establish a practice-relevant baseline of spatial reasoning limitations for AI tools in clinical neurosurgery.