🛰️ Daily AI Frontier
‹ back to 2026-09-08

WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data

Research Medical/Healthcare AI

Ranking

Overall 79
Content 85
Popularity 67

Observed public metrics from 1 member.

Representative image for WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data

Merged summary

TL;DR - WearableQA is a benchmark testing whether LLMs can reason over noisy, longitudinal health data from real wearable users. Results across 14 models show substantial capability differences, but most remain below 60% accuracy.

  • Includes 4,084 ten-option questions derived from wearable time series, blood biomarkers, and demographics for 200 users with up to 500 days of measurements.
  • Covers 16 question types spanning data computation versus health interpretation and single-signal versus cross-signal reasoning.
  • Uses literature-grounded findings and statistically validated population patterns to generate reliable questions at scale.
  • Model accuracy ranges from 19.6% to 72.9%, compared with a 10% chance baseline.

Sources (1)

WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data

arXiv cs.CL Ji Soo Lee, Xilun Chen, Pierce Chuang, Ashish Shenoy, Jason Wei, Dohwan Ko, Hyunwoo J. Kim, Benoit Corda 2026-09-04 arXiv:2609.05405
Public signals Hugging Face upvotes 43
Providers: Hugging Face · Upvotes 43 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-25 14:22:52.705979 UTC

TL;DR - WearableQA is a benchmark testing whether LLMs can reason over noisy, longitudinal health data from real wearable users. Results across 14 models show substantial capability differences, but most remain below 60% accuracy.

  • Includes 4,084 ten-option questions derived from wearable time series, blood biomarkers, and demographics for 200 users with up to 500 days of measurements.
  • Covers 16 question types spanning data computation versus health interpretation and single-signal versus cross-signal reasoning.
  • Uses literature-grounded findings and statistically validated population patterns to generate reliable questions at scale.
  • Model accuracy ranges from 19.6% to 72.9%, compared with a 10% chance baseline.
item →