🛰️ Daily AI Frontier
‹ back to 2026-08-02

InfoOps Bench: A live information operations safety benchmark

Research LLM Safety

Ranking

Overall 82
Content 100
Popularity 40

Observed public metrics from 1 member.

Representative image for InfoOps Bench: A live information operations safety benchmark

Merged summary

TL;DR - InfoOps Bench is a live, weekly updated benchmark testing frontier language models against co-option for state-backed information operations. Its evaluation of 17 models finds wide safety variation and frequent compliance, highlighting risks not explained by model size alone.

  • Integrity scores ranged from 8.8% to 94.5% across models and prompt framings.
  • Models differed in harmfulness, fabrication behavior, and fact-checking rates, which ranged from 2.9% to 72.9%.
  • Higher integrity partly correlated with refusing benign requests, exposing a safety-usability tradeoff.
  • Most Chinese-developed models showed 48–70 percentage-point compliance drops for factual China-critical claims versus matched benign claims; GLM 5.2 was the exception.

Sources (1)

InfoOps Bench: A live information operations safety benchmark

arXiv cs.AI Dorian Quelle, Lisa-Maria Neudert, Jonathan Bright, John Gallacher 2026-07-30 arXiv:2607.28503
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-08-07 09:43:34.238785 UTC

TL;DR - InfoOps Bench is a live, weekly updated benchmark testing frontier language models against co-option for state-backed information operations. Its evaluation of 17 models finds wide safety variation and frequent compliance, highlighting risks not explained by model size alone.

  • Integrity scores ranged from 8.8% to 94.5% across models and prompt framings.
  • Models differed in harmfulness, fabrication behavior, and fact-checking rates, which ranged from 2.9% to 72.9%.
  • Higher integrity partly correlated with refusing benign requests, exposing a safety-usability tradeoff.
  • Most Chinese-developed models showed 48–70 percentage-point compliance drops for factual China-critical claims versus matched benign claims; GLM 5.2 was the exception.
item →