🛰️ Daily AI Frontier
‹ back to 2026-07-22

Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning

Research LLMs & Foundation Models

Ranking

Overall 82
Content 100
Popularity 40

Observed public metrics from 1 member.

Merged summary

TL;DR - The paper identifies repetitive prompt copying as a widespread long-context reasoning failure caused by poor evidence grounding. Its evidence-aware RL method, GEAR, improves benchmark scores by up to 4.6 points while producing shorter, less repetitive reasoning traces.

  • Repetitive copying worsens as context length increases and correlates with incorrect answers.
  • GEAR rewards overlap with relevant evidence and penalizes copying from irrelevant distractors.
  • An automated pipeline creates evidence-annotated training data from arbitrary documents.
  • Improvements are consistent across model scales and benchmarks, with larger gains on longer contexts.

Sources (1)

Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning

arXiv cs.CL Lizhe Fang, Weizhou Shen, Tianyi Tang, Yisen Wang 2026-07-21 arXiv:2607.19345
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-08-10 02:56:42.391184 UTC

TL;DR - The paper identifies repetitive prompt copying as a widespread long-context reasoning failure caused by poor evidence grounding. Its evidence-aware RL method, GEAR, improves benchmark scores by up to 4.6 points while producing shorter, less repetitive reasoning traces.

  • Repetitive copying worsens as context length increases and correlates with incorrect answers.
  • GEAR rewards overlap with relevant evidence and penalizes copying from irrelevant distractors.
  • An automated pipeline creates evidence-annotated training data from arbitrary documents.
  • Improvements are consistent across model scales and benchmarks, with larger gains on longer contexts.
item →