🛰️ Daily AI Frontier
‹ back to 2026-07-23

JANUS: Foreseeing Latent Risk for Long-Horizon Agent Safety

arXiv cs.AI LLM Agents Yuan Xiong, Linji Hao, Shizhu He, Yequan Wang, Lijun Li 2026-07-22

TL;DR - JANUS trains an agent guard to anticipate delayed risks from partial trajectories and block unsafe actions before execution. Its Vanguard model improves both protection and benign task completion across four safety benchmarks.

  • Uses multi-agent simulation to synthesize diverse, long-horizon trajectories.
  • Jointly trains future-risk anticipation and safety adjudication with CoAA-RL.
  • Improves average protection by 15.9 percentage points over baseline guards.
  • Increases benign task completion by 5.1 percentage points.

view merged work →