A Geometric Perspective on Stabilizing Value Conflict Resolution
Merged summary
TL;DR - This paper studies how chain-of-thought reasoning can stabilize LLM responses to conflicting human values by smoothing sharp directions in the loss landscape. A value-conflict-focused CoT design also improves moral-reasoning performance, suggesting a path toward more pluralistic alignment.
- Scalar RLHF rewards can create optimization instability when compressing conflicting values into one signal.
- CoT reasoning correlates with greater loss-landscape smoothing along its sharpest direction.
- Value-conflict-focused CoT generalizes across different moral-reasoning benchmarks.
- Explicitly redesigned reasoning dynamics further increase smoothing and moral-reasoning performance.
Sources (1)
A Geometric Perspective on Stabilizing Value Conflict Resolution
TL;DR - This paper studies how chain-of-thought reasoning can stabilize LLM responses to conflicting human values by smoothing sharp directions in the loss landscape. A value-conflict-focused CoT design also improves moral-reasoning performance, suggesting a path toward more pluralistic alignment.
- Scalar RLHF rewards can create optimization instability when compressing conflicting values into one signal.
- CoT reasoning correlates with greater loss-landscape smoothing along its sharpest direction.
- Value-conflict-focused CoT generalizes across different moral-reasoning benchmarks.
- Explicitly redesigned reasoning dynamics further increase smoothing and moral-reasoning performance.