🛰️ Daily AI Frontier
‹ back to 2026-08-27

PLCBench: Can Autonomous LLM Agents Turn PLC Access into Sustained Physical Impact?

Research LLM Agents

Ranking

Overall 83
Content 100
Popularity 42

Observed public metrics from 1 member.

Representative image for PLCBench: Can Autonomous LLM Agents Turn PLC Access into Sustained Physical Impact?

Merged summary

TL;DR - PLCBench is a hardware-in-the-loop benchmark that measures whether autonomous, tool-using LLM agents can turn access to commercial programmable logic controllers into sustained physical impact. Across 240 trials, agents achieved their physical objectives in 31.3% of episodes, demonstrating a measurable cyber-physical threat while pinpointing common failure stages.

  • The framework combines vendor-native PLC interaction, four commercial PLCs, four closed-loop simulated processes, and independent deterministic outcome verification.
  • Of 240 episodes spanning five LLM families, 75 sustained their assigned physical objectives.
  • Ninety-eight episodes failed before a valid native read; another 62 achieved a process-linked write but failed to sustain the objective.
  • Richer process observations raised conditional success after a process-linked write from 44.2% to 64.0%.

Sources (1)

PLCBench: Can Autonomous LLM Agents Turn PLC Access into Sustained Physical Impact?

arXiv cs.CR Yitian Zhou, Jingyu Zheng, Qiliang Jiang, Linkang Du, Haoming Liu, Lichao Wu, Shiyi Zhao, Mengxiang Liu, Ruilong Deng 2026-08-27 arXiv:2608.26882
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-08-28 14:14:00.546110 UTC

TL;DR - PLCBench is a hardware-in-the-loop benchmark that measures whether autonomous, tool-using LLM agents can turn access to commercial programmable logic controllers into sustained physical impact. Across 240 trials, agents achieved their physical objectives in 31.3% of episodes, demonstrating a measurable cyber-physical threat while pinpointing common failure stages.

  • The framework combines vendor-native PLC interaction, four commercial PLCs, four closed-loop simulated processes, and independent deterministic outcome verification.
  • Of 240 episodes spanning five LLM families, 75 sustained their assigned physical objectives.
  • Ninety-eight episodes failed before a valid native read; another 62 achieved a process-linked write but failed to sustain the objective.
  • Richer process observations raised conditional success after a process-linked write from 44.2% to 64.0%.
item →