🛰️ Daily AI Frontier
‹ back to 2026-09-08

WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data

arXiv cs.CL Medical/Healthcare AI Ji Soo Lee, Xilun Chen, Pierce Chuang, Ashish Shenoy, Jason Wei, Dohwan Ko, Hyunwoo J. Kim, Benoit Corda 2026-09-04
Representative image for WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data

TL;DR - WearableQA is a benchmark testing whether LLMs can reason over noisy, longitudinal health data from real wearable users. Results across 14 models show substantial capability differences, but most remain below 60% accuracy.

  • Includes 4,084 ten-option questions derived from wearable time series, blood biomarkers, and demographics for 200 users with up to 500 days of measurements.
  • Covers 16 question types spanning data computation versus health interpretation and single-signal versus cross-signal reasoning.
  • Uses literature-grounded findings and statistically validated population patterns to generate reliable questions at scale.
  • Model accuracy ranges from 19.6% to 72.9%, compared with a 10% chance baseline.

view merged work →