Measuring the practice of shared-decision making (OPTION12): An Investigation into Open-sourced Smaller LLMs (OS-sLLMs) for Better Privacy and Sustainability
For researchers and practitioners in clinical communication assessment, this work provides an initial benchmark for using privacy-preserving smaller LLMs in shared decision making evaluation, though the results are preliminary and incremental.
This paper presents the first study of open-source smaller LLMs for automated assessment of shared decision making using the OPTION12 framework, finding that general-domain models outperform medical-domain models, with Gemma3:12b achieving the strongest agreement (Pearson r=0.51, Spearman ρ=0.59). The results suggest OS-sLLMs offer a promising foundation for privacy-preserving human-in-the-loop SDM assessment but cannot yet replace human annotators.
We present LLM4SDM, the first study of open-source smaller language models (OS-sLLMs) for automated assessment of shared decision making (SDM) using the Observer OPTION12 framework. Unlike previous work that relies on large commercial models and the shorter OPTION5 instrument, our study focuses on privacy-preserving locally deployable models and Dutch melanoma consultation transcripts. Using expert-annotated clinical consultations, we evaluate three general-domain and two medical-domain OS-sLLMs during a development-phase pilot study. Results show that general-domain models outperform medical-domain models, which exhibit substantial hallucination and instruction-following failures. Gemma3:12b achieves the strongest agreement with human annotations (Pearson r=0.51, Spearman \r{ho}=0.59). Item-level and qualitative analyses reveal systematic challenges related to temporal discourse reasoning, conversational role attribution, and evidence grounding. We further introduce a Judge-LLM consensus framework designed to support disagreement resolution among multiple models. Our findings suggest that while current OS-sLLMs cannot replace human annotators, they offer a promising foundation for privacy-preserving human-in-the-loop SDM assessment.