Cross-Lingual Clinical Annotation Projection as Constrained Text Generation: A Six-Language Study
Ranking
Overall
82
Content
95
Popularity
N/A
No observed public metrics; popularity remains neutral/archived.
Merged summary
TL;DR - This study frames cross-lingual clinical annotation projection as constrained text generation, using LLMs to insert entity tags without altering source documents. The approach outperformed prior projection methods across six languages while enabling deterministic validation and character-offset reconstruction.
- GLM 5.2 achieved a mean strict-match F1 of 0.9201 across 18 language–entity combinations; locally deployable Gemma4:31B reached 0.9133.
- The best LLM configuration surpassed the previous state of the art in all 18 settings, improving strict F1 by 0.0564–0.1512.
- The workflow projected Spanish Disease, Symptom, and Procedure annotations and produced 55,416 grounded mentions with reconstructed offsets.
- Local inference and deterministic validation could reduce expert effort and costs when building clinical NLP datasets for lower-resource languages.
Sources (1)
Cross-Lingual Clinical Annotation Projection as Constrained Text Generation: A Six-Language Study
Public signals
N/A
TL;DR - This study frames cross-lingual clinical annotation projection as constrained text generation, using LLMs to insert entity tags without altering source documents. The approach outperformed prior projection methods across six languages while enabling deterministic validation and character-offset reconstruction.
- GLM 5.2 achieved a mean strict-match F1 of 0.9201 across 18 language–entity combinations; locally deployable Gemma4:31B reached 0.9133.
- The best LLM configuration surpassed the previous state of the art in all 18 settings, improving strict F1 by 0.0564–0.1512.
- The workflow projected Spanish Disease, Symptom, and Procedure annotations and produced 55,416 grounded mentions with reconstructed offsets.
- Local inference and deterministic validation could reduce expert effort and costs when building clinical NLP datasets for lower-resource languages.