🛰️ Daily AI Frontier
‹ back to 2026-08-27

PLCBench: Can Autonomous LLM Agents Turn PLC Access into Sustained Physical Impact?

arXiv cs.CR LLM Agents Yitian Zhou, Jingyu Zheng, Qiliang Jiang, Linkang Du, Haoming Liu, Lichao Wu, Shiyi Zhao, Mengxiang Liu, Ruilong Deng 2026-08-27
Representative image for PLCBench: Can Autonomous LLM Agents Turn PLC Access into Sustained Physical Impact?

TL;DR - PLCBench is a hardware-in-the-loop benchmark that measures whether autonomous, tool-using LLM agents can turn access to commercial programmable logic controllers into sustained physical impact. Across 240 trials, agents achieved their physical objectives in 31.3% of episodes, demonstrating a measurable cyber-physical threat while pinpointing common failure stages.

  • The framework combines vendor-native PLC interaction, four commercial PLCs, four closed-loop simulated processes, and independent deterministic outcome verification.
  • Of 240 episodes spanning five LLM families, 75 sustained their assigned physical objectives.
  • Ninety-eight episodes failed before a valid native read; another 62 achieved a process-linked write but failed to sustain the objective.
  • Richer process observations raised conditional success after a process-linked write from 44.2% to 64.0%.

view merged work →