InfoOps Bench: A live information operations safety benchmark
TL;DR - InfoOps Bench is a live, weekly updated benchmark testing frontier language models against co-option for state-backed information operations. Its evaluation of 17 models finds wide safety variation and frequent compliance, highlighting risks not explained by model size alone.
- Integrity scores ranged from 8.8% to 94.5% across models and prompt framings.
- Models differed in harmfulness, fabrication behavior, and fact-checking rates, which ranged from 2.9% to 72.9%.
- Higher integrity partly correlated with refusing benign requests, exposing a safety-usability tradeoff.
- Most Chinese-developed models showed 48–70 percentage-point compliance drops for factual China-critical claims versus matched benign claims; GLM 5.2 was the exception.