🛰️ Daily AI Frontier
‹ back to 2026-08-13

LoongReflect: Boosting Long-Horizon Reflection in Search Agents via Global Perspective Distillation

arXiv cs.LG LLM Agents Zhixin Zhang, Xinke Jiang, Zhibang Yang, Weixuan Xu, Guohong Qiu, Xu Chu, Junfeng Zhao, Yasha Wang 2026-08-12
Representative image for LoongReflect: Boosting Long-Horizon Reflection in Search Agents via Global Perspective Distillation

TL;DR - LoongReflect trains search agents to make better long-horizon reflection and backtracking decisions by combining global teacher supervision with outcome-based reinforcement learning. It improves multi-hop retrieval and mathematical reasoning over outcome-only RL and self-distillation baselines.

  • Models reflection as memory control over a reversible trajectory tree.
  • Reflection updates working memory with verified facts, missing evidence, and branch risks.
  • Backtracking removes unreliable branches while retaining concise corrective lessons.
  • Coordinates targeted teacher distillation with trajectory-level GRPO using a look-ahead, extragradient-style mechanism.

view merged work →