🛰️ Daily AI Frontier
‹ back to 2026-07-22

Co-Evolving LLM Evaluators and Policies via DynamicRubric

Research LLMs & Foundation Models

Ranking

Overall 86
Content 95
Popularity 66

Observed public metrics from 1 member.

Merged summary

TL;DR - DynamicRubric co-evolves LLM evaluators and policies by generating candidate-set-specific weighted rubrics, preserving useful score gaps as policy outputs converge in quality. With 8B backbones, it outperforms much larger evaluator baselines and improves reasoning, coding, and production search metrics.

  • Frames relative evaluator score gaps as the exact optimization signal for reallocating probability between responses.
  • Generates weighted binary rubric items conditioned on each candidate response set, then aggregates judgments into response-level scores.
  • Provides stronger policy supervision than baselines using a 70B reward model or 235B static rubric generator.
  • A resulting model serves all WeChat Search AI-answering traffic across tens of millions of daily requests.

Sources (1)

Co-Evolving LLM Evaluators and Policies via DynamicRubric

arXiv cs.LG Beining Wang, Weihang Su, Hongtao Tian, Hao Kong, Tao Yang, Ting Yao, Qingyi Pan, Yueyue Wu, Qingyao Ai, Min Zhang, Yiqun Liu 2026-07-22 arXiv:2607.20083
Public signals Semantic Scholar citations 1 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 1 · Influential citations 0 X · N/A Fetched 2026-08-15 14:32:06.612341 UTC

TL;DR - DynamicRubric co-evolves LLM evaluators and policies by generating candidate-set-specific weighted rubrics, preserving useful score gaps as policy outputs converge in quality. With 8B backbones, it outperforms much larger evaluator baselines and improves reasoning, coding, and production search metrics.

  • Frames relative evaluator score gaps as the exact optimization signal for reallocating probability between responses.
  • Generates weighted binary rubric items conditioned on each candidate response set, then aggregates judgments into response-level scores.
  • Provides stronger policy supervision than baselines using a 70B reward model or 235B static rubric generator.
  • A resulting model serves all WeChat Search AI-answering traffic across tens of millions of daily requests.
item →