WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data
Ranking
Overall
79
Content
85
Popularity
67
Observed public metrics from 1 member.
Merged summary
TL;DR - WearableQA is a benchmark testing whether LLMs can reason over noisy, longitudinal health data from real wearable users. Results across 14 models show substantial capability differences, but most remain below 60% accuracy.
- Includes 4,084 ten-option questions derived from wearable time series, blood biomarkers, and demographics for 200 users with up to 500 days of measurements.
- Covers 16 question types spanning data computation versus health interpretation and single-signal versus cross-signal reasoning.
- Uses literature-grounded findings and statistically validated population patterns to generate reliable questions at scale.
- Model accuracy ranges from 19.6% to 72.9%, compared with a 10% chance baseline.
Sources (1)
WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data
Public signals
Hugging Face upvotes 43
TL;DR - WearableQA is a benchmark testing whether LLMs can reason over noisy, longitudinal health data from real wearable users. Results across 14 models show substantial capability differences, but most remain below 60% accuracy.
- Includes 4,084 ten-option questions derived from wearable time series, blood biomarkers, and demographics for 200 users with up to 500 days of measurements.
- Covers 16 question types spanning data computation versus health interpretation and single-signal versus cross-signal reasoning.
- Uses literature-grounded findings and statistically validated population patterns to generate reliable questions at scale.
- Model accuracy ranges from 19.6% to 72.9%, compared with a 10% chance baseline.