Cross-Domain Hybrid OPD for Generalizable Search Agents
Ranking
Overall
67
Content
80
Popularity
36
Observed public metrics from 1 member.
Merged summary
TL;DR - A technical report on the training framework behind the Yuanbao search agent, which combines agentic RL with cross-domain expert on-policy distillation to build a specialized search agent without paying the usual "alignment tax" on general capabilities.
- Built on the Hunyuan3 architecture; agentic RL trains autonomous multi-step planning and iterative retrieval over dynamic information sources.
- The core contribution is a cross-domain expert On-Policy Distillation (OPD) pipeline: experts covering complementary general-purpose domains are distilled into the search-specialized student to restore and further improve broad capability.
- Framing: specialization and generality are treated as jointly optimizable rather than competing, directly targeting the alignment tax seen when tuning LLMs for narrow search behaviors.
- Reported experiments claim competitive search performance alongside consistent general-capability gains; the report does not provide specific benchmark numbers in the abstract, so magnitude of improvement is unverified here.
Sources (1)
Cross-Domain Hybrid OPD for Generalizable Search Agents
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - A technical report on the training framework behind the Yuanbao search agent, which combines agentic RL with cross-domain expert on-policy distillation to build a specialized search agent without paying the usual "alignment tax" on general capabilities.
- Built on the Hunyuan3 architecture; agentic RL trains autonomous multi-step planning and iterative retrieval over dynamic information sources.
- The core contribution is a cross-domain expert On-Policy Distillation (OPD) pipeline: experts covering complementary general-purpose domains are distilled into the search-specialized student to restore and further improve broad capability.
- Framing: specialization and generality are treated as jointly optimizable rather than competing, directly targeting the alignment tax seen when tuning LLMs for narrow search behaviors.
- Reported experiments claim competitive search performance alongside consistent general-capability gains; the report does not provide specific benchmark numbers in the abstract, so magnitude of improvement is unverified here.