Record Grouping Controls Evidence Weight in Language Models
TL;DR - This paper shows that how retrieved records are grouped before being presented to a language model can substantially alter their evidential weight and the model’s decisions. It proposes a content-aware representation that deduplicates within groups, aggregates complementary information, and limits each group’s contribution.
- Equal numbers of record groups can represent different evidence states depending on their content and partitioning.
- Across 104,402 trials and six public checkpoints, false splits increased measured effects by 10.27–32.66 percentage points, while false merges reduced them by 9.13–31.79 points.
- A matched six-slot control preserved the positive direction in all 16 tested cells, indicating that the split effect was not solely due to added presentation slots.
- A 48-item controlled campaign panel found partition-induced decision shifts across all four models, with checkpoint-dependent behavior and substantial ordering interactions.