Language models judge war differently when tested for alignment
TL;DR - A study of 20 large language models finds that explicitly telling models they are being tested for human-value alignment substantially reduces their stated willingness to start a war and changes the factors driving their decisions. This suggests safety evaluations may misrepresent deployed behavior when models react to evaluation cues.
- The full-factorial experiment covered 32 scenarios, 10 repetitions, and two conditions, totaling 12,800 judgments.
- Adding an alignment-testing cue reduced mean willingness to start war by 13.43 points on a 0–100 scale.
- At baseline, probability of success was the largest factor for 17 of 20 models; with the cue, civilian casualties became largest for 12 models.
- The shift primarily reflected reduced weighting of strategic considerations, including probability of success and domestic support.