🛰️ Daily AI Frontier
‹ back to 2026-07-23

JANUS: Foreseeing Latent Risk for Long-Horizon Agent Safety

Research LLM Agents

Merged summary

TL;DR - JANUS trains an agent guard to anticipate delayed risks from partial trajectories and block unsafe actions before execution. Its Vanguard model improves both protection and benign task completion across four safety benchmarks.

  • Uses multi-agent simulation to synthesize diverse, long-horizon trajectories.
  • Jointly trains future-risk anticipation and safety adjudication with CoAA-RL.
  • Improves average protection by 15.9 percentage points over baseline guards.
  • Increases benign task completion by 5.1 percentage points.

Sources (1)

JANUS: Foreseeing Latent Risk for Long-Horizon Agent Safety

arXiv cs.AI Yuan Xiong, Linji Hao, Shizhu He, Yequan Wang, Lijun Li 2026-07-22 arXiv:2607.19913

TL;DR - JANUS trains an agent guard to anticipate delayed risks from partial trajectories and block unsafe actions before execution. Its Vanguard model improves both protection and benign task completion across four safety benchmarks.

  • Uses multi-agent simulation to synthesize diverse, long-horizon trajectories.
  • Jointly trains future-risk anticipation and safety adjudication with CoAA-RL.
  • Improves average protection by 15.9 percentage points over baseline guards.
  • Increases benign task completion by 5.1 percentage points.
item →