🛰️ Daily AI Frontier
‹ back to 2026-08-29

BALMS: Benchmarking Agentic LLMs for Longitudinal Mental Health Sensing

Research Medical/Healthcare AI

Ranking

Overall 78
Content 95
Popularity 37

Observed public metrics from 1 member.

Merged summary

TL;DR - BALMS is a benchmark for evaluating LLM agents that use longitudinal wearable data to predict mental-wellbeing scores and generate evidence-grounded rationales. Results show that current agents often fail to beat a simple mean baseline, highlighting weaknesses in temporal grounding and numerical reasoning.

  • Covers three real-world longitudinal datasets, two task families, three agentic paradigms, and five open- and closed-source LLM backbones.
  • Zero-shot agents generally underperform the mean baseline unless paired with stronger models or compact, semantically meaningful features.
  • Chain-of-thought prompting helps reasoning-oriented models but does not reliably ensure temporal grounding or numerical accuracy.
  • The findings motivate agents that selectively retrieve history and reason over interpretable behavioral features.

Sources (1)

BALMS: Benchmarking Agentic LLMs for Longitudinal Mental Health Sensing

arXiv cs.CL Yu Yvonne Wu, Arvind Pillai, Yuliang Chen, Yuwei Zhang, Sudarshan Regmi, Tess Z. Griffin, Michael V. Heinz, Lisa A. Marsch, Nicholas C. Jacobson, Andrew Campbell 2026-08-27 arXiv:2608.27219
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-15 14:24:45.972861 UTC

TL;DR - BALMS is a benchmark for evaluating LLM agents that use longitudinal wearable data to predict mental-wellbeing scores and generate evidence-grounded rationales. Results show that current agents often fail to beat a simple mean baseline, highlighting weaknesses in temporal grounding and numerical reasoning.

  • Covers three real-world longitudinal datasets, two task families, three agentic paradigms, and five open- and closed-source LLM backbones.
  • Zero-shot agents generally underperform the mean baseline unless paired with stronger models or compact, semantically meaningful features.
  • Chain-of-thought prompting helps reasoning-oriented models but does not reliably ensure temporal grounding or numerical accuracy.
  • The findings motivate agents that selectively retrieve history and reason over interpretable behavioral features.
item →