🛰️ Daily AI Frontier
‹ back to 2026-09-09

Environments as Scaffold: Enriching Feedback to Bootstrap Self-Evolving Agents in Long-Horizon Tasks

arXiv cs.LG LLM Agents Hongbang Yuan, Zhuoran Jin, Yixin Cao 2026-09-08

TL;DR - This paper introduces Feedback-Enriched Environments (FEEs), which add richer observations during reinforcement learning to help autonomous LLM agents overcome sparse rewards in long-horizon tasks. Across SciWorld and BFCL, FEEs consistently improve performance while stabilizing training and encouraging exploration.

  • Shifts agent bootstrapping from supervised fine-tuning toward environment-side feedback adaptation.
  • Transitions from action guidance to observation enrichment during later intra-episode exploration and inter-episode evolution.
  • Demonstrates gains across multiple Qwen3 scales and RL algorithms, including GRPO, GSPO, and DAPO.
  • Finds that feedback is internalized into policy weights and that intra-group feedback consistency is important for stable optimization.

view merged work →