HealMed: Multilingual Evaluation of Large Language Models in Medicine
Ranking
Overall
79
Content
95
Popularity
42
Observed public metrics from 1 member.
Merged summary
TL;DR - HealMed is an expert-reviewed benchmark evaluating medical LLMs across nine languages and three task formats. It reveals substantial multilingual performance gaps, especially in low-resource languages, and shows that medical specialization does not guarantee cross-language robustness.
- Includes 1,000 examples per language from nine datasets, spanning multiple-choice QA, natural language inference, and open-ended QA.
- Developed over two years by 23 physicians and medical experts across nine countries and regions, with each translation reviewed by two bilingual experts.
- Strong proprietary models were generally more stable across languages than many open-source and medically specialized models.
- Expert translation revisions sometimes raised and sometimes lowered scores, demonstrating that translation quality materially affects multilingual evaluation.
Sources (1)
HealMed: Multilingual Evaluation of Large Language Models in Medicine
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - HealMed is an expert-reviewed benchmark evaluating medical LLMs across nine languages and three task formats. It reveals substantial multilingual performance gaps, especially in low-resource languages, and shows that medical specialization does not guarantee cross-language robustness.
- Includes 1,000 examples per language from nine datasets, spanning multiple-choice QA, natural language inference, and open-ended QA.
- Developed over two years by 23 physicians and medical experts across nine countries and regions, with each translation reviewed by two bilingual experts.
- Strong proprietary models were generally more stable across languages than many open-source and medically specialized models.
- Expert translation revisions sometimes raised and sometimes lowered scores, demonstrating that translation quality materially affects multilingual evaluation.