AREX: Towards a Recursively Self-Improving Agent for Deep Research
TL;DR - AREX is a family of deep-research agents that recursively audits provisional answers and targets unresolved constraints. This approach improves long-horizon research performance while remaining competitive with models using more activated parameters.
- Alternates evidence-gathering with constraint-wise verification and targeted follow-up research.
- Learns a context-update tool that compresses history while preserving verified evidence and open constraints.
- Uses agentic mid-training and long-horizon reinforcement learning with emphasis on decisive or corrective steps.
- Dense 4B and 122B-A10B MoE variants outperform comparable-scale baselines across multiple research, reasoning, and tool-use benchmarks.