🛰️ Daily AI Frontier
‹ back to 2026-08-30

DocTalkBN: A Novel Dataset of Expert Telemedicine Conversations in Bengali

Research Medical/Healthcare AI

Ranking

Overall 78
Content 90
Popularity N/A

No observed public metrics; popularity remains neutral/archived.

Merged summary

TL;DR - DocTalkBN is a large-scale multimodal dataset of authentic Bengali telemedicine conversations designed to support reliable medical conversational AI in a low-resource language. It provides clinically grounded data and benchmarks for triage, advice safety, and medical entity recognition.

  • Contains 557.63 hours of paired audio and text from 1,515 multi-turn patient calls and 10,274 host–doctor exchanges.
  • Covers 26 medical specialties and totals 1.7 million tokens from nationally broadcast consultations with board-certified physicians.
  • Preserves spontaneous spoken interactions rather than relying on medical forums, written content, or synthetic conversations.
  • Includes benchmark tasks for medical triage classification, advice safety evaluation, and medical named entity recognition, evaluated with multiple LLM and encoder-based baselines.

Sources (1)

DocTalkBN: A Novel Dataset of Expert Telemedicine Conversations in Bengali

arXiv cs.CL Anik Saha, Fahmida Sultana Naznin, Sadatul Islam Sadi, Ananya Shahrin Promi, Wahid Al Azad Navid, Rifat Shahriyar 2026-08-27 arXiv:2608.27110
Public signals N/A
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-26 14:17:41.482576 UTC

TL;DR - DocTalkBN is a large-scale multimodal dataset of authentic Bengali telemedicine conversations designed to support reliable medical conversational AI in a low-resource language. It provides clinically grounded data and benchmarks for triage, advice safety, and medical entity recognition.

  • Contains 557.63 hours of paired audio and text from 1,515 multi-turn patient calls and 10,274 host–doctor exchanges.
  • Covers 26 medical specialties and totals 1.7 million tokens from nationally broadcast consultations with board-certified physicians.
  • Preserves spontaneous spoken interactions rather than relying on medical forums, written content, or synthetic conversations.
  • Includes benchmark tasks for medical triage classification, advice safety evaluation, and medical named entity recognition, evaluated with multiple LLM and encoder-based baselines.
item →