Shanghai AI Laboratory with Tongji and Sun Yat-sen (the MedXpertQA pattern: a tracked lab's medical eval via academic collaboration) shifts medical LLM evaluation from final-answer accuracy to guideline adherence: can a model produce the clinical decision pathway a guideline prescribes? MEGA-CDP is built by a guideline-to-case pipeline over 2,274 English and Chinese clinical practice guidelines, yielding 42,353 cases with explicit reference pathways, evaluated in both single-turn vignette and multi-turn interactive settings with a pathway-consistency scoring framework.

Paper

medicalbenchmarkevaluationresearch

Related