🛰️ Daily AI Frontier
‹ back to 2026-08-22

HealMed: Multilingual Evaluation of Large Language Models in Medicine

arXiv cs.CL Medical/Healthcare AI Yingjian Chen, Fan Gao, Sherry T. Tong, Haoyu Zhang, Aosong Feng, Kevin W. Jin, Xing Wu, Jinghui Lu, Abdul Samad, Akbar Faruqi, Cesar Caraballo, Cibele Brandão, Dhruva, Gupta, Eunji Jeon, Gabriel Madera-Santiago, Geon Lee, Hugo Toshio Itikawa, Insook Cho, Isabelli Martins, Isarar Siddique, Israr Ahmed, Jihyo Kwak, Kanyakorn Veerakanjana, Luis Guilherme Cardoso, Minjin Kim, Piyalitt Ittichaiwong, Renee Dua, Santiago Gudiño-Rosales, Xiujie Chen, Zeo Lapalus, Zixin Xu, Michihiro Yasunaga, Rex Ying, Heuiseok Lim, Jaewoo Kang, Chanjun Park, Hang Jiang, Ethan Goh, Hyunjae Kim, Edison Marrese-Taylor, Yusuke Iwasawa, Yutaka Matsuo, Qingyu Chen, Irene Li 2026-08-20

TL;DR - HealMed is an expert-reviewed benchmark evaluating medical LLMs across nine languages and three task formats. It reveals substantial multilingual performance gaps, especially in low-resource languages, and shows that medical specialization does not guarantee cross-language robustness.

  • Includes 1,000 examples per language from nine datasets, spanning multiple-choice QA, natural language inference, and open-ended QA.
  • Developed over two years by 23 physicians and medical experts across nine countries and regions, with each translation reviewed by two bilingual experts.
  • Strong proprietary models were generally more stable across languages than many open-source and medically specialized models.
  • Expert translation revisions sometimes raised and sometimes lowered scores, demonstrating that translation quality materially affects multilingual evaluation.

view merged work →