Reinforcement Learning for Large Language Model Selective Evidence Adoption from Contaminated Retrieval Results
Merged summary
TL;DR - SelectBench trains LLMs to use valid retrieved evidence while rejecting misleading or harmful content. Reinforcement learning produced modest gains without degrading general capabilities, but did not improve prompt-injection resistance.
- SelectBench-v2 evaluates selective evidence adoption across 325 corrected test examples.
- DAPO post-training raised strict success from 22.46% to 25.54% with rule rewards and 26.46% with a frozen semantic judge.
- Trained models adopted less forbidden content and generated shorter, more focused answers.
- Gains did not survive Holm correction; MMLU and clean HotpotQA performance remained stable.
Sources (1)
Reinforcement Learning for Large Language Model Selective Evidence Adoption from Contaminated Retrieval Results
TL;DR - SelectBench trains LLMs to use valid retrieved evidence while rejecting misleading or harmful content. Reinforcement learning produced modest gains without degrading general capabilities, but did not improve prompt-injection resistance.
- SelectBench-v2 evaluates selective evidence adoption across 325 corrected test examples.
- DAPO post-training raised strict success from 22.46% to 25.54% with rule rewards and 26.46% with a frozen semantic judge.
- Trained models adopted less forbidden content and generated shorter, more focused answers.
- Gains did not survive Holm correction; MMLU and clean HotpotQA performance remained stable.