🛰️ Daily AI Frontier
‹ back to 2026-08-29

BALMS: Benchmarking Agentic LLMs for Longitudinal Mental Health Sensing

arXiv cs.CL Medical/Healthcare AI Yu Yvonne Wu, Arvind Pillai, Yuliang Chen, Yuwei Zhang, Sudarshan Regmi, Tess Z. Griffin, Michael V. Heinz, Lisa A. Marsch, Nicholas C. Jacobson, Andrew Campbell 2026-08-27

TL;DR - BALMS is a benchmark for evaluating LLM agents that use longitudinal wearable data to predict mental-wellbeing scores and generate evidence-grounded rationales. Results show that current agents often fail to beat a simple mean baseline, highlighting weaknesses in temporal grounding and numerical reasoning.

  • Covers three real-world longitudinal datasets, two task families, three agentic paradigms, and five open- and closed-source LLM backbones.
  • Zero-shot agents generally underperform the mean baseline unless paired with stronger models or compact, semantically meaningful features.
  • Chain-of-thought prompting helps reasoning-oriented models but does not reliably ensure temporal grounding or numerical accuracy.
  • The findings motivate agents that selectively retrieve history and reason over interpretable behavioral features.

view merged work →