LoongReflect: Boosting Long-Horizon Reflection in Search Agents via Global Perspective Distillation
Ranking
Overall
77
Content
95
Popularity
34
Observed public metrics from 1 member.
Merged summary
TL;DR - LoongReflect trains search agents to make better long-horizon reflection and backtracking decisions by combining global teacher supervision with outcome-based reinforcement learning. It improves multi-hop retrieval and mathematical reasoning over outcome-only RL and self-distillation baselines.
- Models reflection as memory control over a reversible trajectory tree.
- Reflection updates working memory with verified facts, missing evidence, and branch risks.
- Backtracking removes unreliable branches while retaining concise corrective lessons.
- Coordinates targeted teacher distillation with trajectory-level GRPO using a look-ahead, extragradient-style mechanism.
Sources (1)
LoongReflect: Boosting Long-Horizon Reflection in Search Agents via Global Perspective Distillation
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - LoongReflect trains search agents to make better long-horizon reflection and backtracking decisions by combining global teacher supervision with outcome-based reinforcement learning. It improves multi-hop retrieval and mathematical reasoning over outcome-only RL and self-distillation baselines.
- Models reflection as memory control over a reversible trajectory tree.
- Reflection updates working memory with verified facts, missing evidence, and branch risks.
- Backtracking removes unreliable branches while retaining concise corrective lessons.
- Coordinates targeted teacher distillation with trajectory-level GRPO using a look-ahead, extragradient-style mechanism.