BALMS: Benchmarking Agentic LLMs for Longitudinal Mental Health Sensing
Ranking
Overall
78
Content
95
Popularity
37
Observed public metrics from 1 member.
Merged summary
TL;DR - BALMS is a benchmark for evaluating LLM agents that use longitudinal wearable data to predict mental-wellbeing scores and generate evidence-grounded rationales. Results show that current agents often fail to beat a simple mean baseline, highlighting weaknesses in temporal grounding and numerical reasoning.
- Covers three real-world longitudinal datasets, two task families, three agentic paradigms, and five open- and closed-source LLM backbones.
- Zero-shot agents generally underperform the mean baseline unless paired with stronger models or compact, semantically meaningful features.
- Chain-of-thought prompting helps reasoning-oriented models but does not reliably ensure temporal grounding or numerical accuracy.
- The findings motivate agents that selectively retrieve history and reason over interpretable behavioral features.
Sources (1)
BALMS: Benchmarking Agentic LLMs for Longitudinal Mental Health Sensing
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - BALMS is a benchmark for evaluating LLM agents that use longitudinal wearable data to predict mental-wellbeing scores and generate evidence-grounded rationales. Results show that current agents often fail to beat a simple mean baseline, highlighting weaknesses in temporal grounding and numerical reasoning.
- Covers three real-world longitudinal datasets, two task families, three agentic paradigms, and five open- and closed-source LLM backbones.
- Zero-shot agents generally underperform the mean baseline unless paired with stronger models or compact, semantically meaningful features.
- Chain-of-thought prompting helps reasoning-oriented models but does not reliably ensure temporal grounding or numerical accuracy.
- The findings motivate agents that selectively retrieve history and reason over interpretable behavioral features.