🛰️ Daily AI Frontier
‹ back to 2026-08-23

Certified Multi-Turn Robustness for LLM Safety via Compositional Bounds and Safety Persistence

Research LLMs & Foundation Models

Ranking

Overall 82
Content 100
Popularity 41

Observed public metrics from 1 member.

Representative image for Certified Multi-Turn Robustness for LLM Safety via Compositional Bounds and Safety Persistence

Merged summary

TL;DR - MTCR is a framework for certifying LLM safety against multi-turn jailbreaks without the rapidly weakening guarantees produced by naive per-turn composition. It provides tighter lower bounds, interpretable safe-horizon estimates, and matching information-theoretic upper bounds.

  • Models conversational safety as a State-Adversarial MDP and defines robustness by the worst-case safety probability over (k) adversarial turns.
  • Uses embedding-space mode decomposition to produce tighter compositional certificates than multiplying independent single-turn bounds.
  • Introduces ((\alpha,\beta))-safety persistence, improving bound degradation from (\underline{p}^{k}) to (\beta^{k}), where (\beta>\underline{p}).
  • Across six LLMs and both (\epsilon)-bounded and Crescendo-style attacks, observed safety remained above the certified lower bounds.

Sources (1)

Certified Multi-Turn Robustness for LLM Safety via Compositional Bounds and Safety Persistence

arXiv cs.AI Yang Liu, Bin Chong, Wenkai Yang, Shuai Zhang, Yancheng Chen, Feiyu Han, GuoZhen, Cheng Zhang, Huaibing Xie, Changze Lv, Shihan Dou, Pluto Zhou 2026-08-21 arXiv:2608.20820
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-22 14:33:07.401959 UTC

TL;DR - MTCR is a framework for certifying LLM safety against multi-turn jailbreaks without the rapidly weakening guarantees produced by naive per-turn composition. It provides tighter lower bounds, interpretable safe-horizon estimates, and matching information-theoretic upper bounds.

  • Models conversational safety as a State-Adversarial MDP and defines robustness by the worst-case safety probability over (k) adversarial turns.
  • Uses embedding-space mode decomposition to produce tighter compositional certificates than multiplying independent single-turn bounds.
  • Introduces ((\alpha,\beta))-safety persistence, improving bound degradation from (\underline{p}^{k}) to (\beta^{k}), where (\beta>\underline{p}).
  • Across six LLMs and both (\epsilon)-bounded and Crescendo-style attacks, observed safety remained above the certified lower bounds.
item →