WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data
TL;DR - WearableQA is a benchmark testing whether LLMs can reason over noisy, longitudinal health data from real wearable users. Results across 14 models show substantial capability differences, but most remain below 60% accuracy.
- Includes 4,084 ten-option questions derived from wearable time series, blood biomarkers, and demographics for 200 users with up to 500 days of measurements.
- Covers 16 question types spanning data computation versus health interpretation and single-signal versus cross-signal reasoning.
- Uses literature-grounded findings and statistically validated population patterns to generate reliable questions at scale.
- Model accuracy ranges from 19.6% to 72.9%, compared with a 10% chance baseline.