Toward Generalizable Cognitive Impairment Detection with Speech-Based Multimodal Large Language Models
Merged summary
TL;DR - A privacy-preserving multimodal framework combines speech acoustics and transcripts using open-source LLMs to detect cognitive impairment. It reports 92.4% accuracy and improved cross-dataset generalization, supporting scalable, non-invasive screening.
- Extracts acoustic embeddings from speech and textual embeddings from automatic transcripts.
- Concatenates modality-specific embeddings for classification without downstream access to raw patient data.
- Evaluated on the ADReSS20 and ADReSSo21 benchmark datasets.
- Consistently outperforms single-modality baselines and reports state-of-the-art cognitive-impairment identification.
Sources (1)
Toward Generalizable Cognitive Impairment Detection with Speech-Based Multimodal Large Language Models
TL;DR - A privacy-preserving multimodal framework combines speech acoustics and transcripts using open-source LLMs to detect cognitive impairment. It reports 92.4% accuracy and improved cross-dataset generalization, supporting scalable, non-invasive screening.
- Extracts acoustic embeddings from speech and textual embeddings from automatic transcripts.
- Concatenates modality-specific embeddings for classification without downstream access to raw patient data.
- Evaluated on the ADReSS20 and ADReSSo21 benchmark datasets.
- Consistently outperforms single-modality baselines and reports state-of-the-art cognitive-impairment identification.