ClawSentry: A Progressive Multi-Tier Security Monitor for Safeguarding Autonomous LLM Agents
Ranking
Overall
75
Content
90
Popularity
39
Observed public metrics from 1 member.
Merged summary
TL;DR - ClawSentry is an open-source, framework-agnostic security gateway that monitors autonomous LLM agents across skill admission, intent review, execution, and post-action consequences. Its progressive review tiers substantially reduce successful attacks while largely preserving performance on benign tasks.
- Combines deterministic checks, semantic review, and bounded read-only agentic investigation, escalating only ambiguous cases.
- Reviews skill packages before first use and tracks rephrased or tool-switched retries across an entire session.
- Applies a shared policy across Codex, Claude Code, Kimi CLI, and Gemini CLI without modifying agent internals.
- On SkillInject with Codex/GPT-5.4, attack success fell from 39.55% to 2.61%, while task success decreased from 83.78% to 83.05%; across five agents, clean-skill task success remained 98.7%.
Sources (1)
ClawSentry: A Progressive Multi-Tier Security Monitor for Safeguarding Autonomous LLM Agents
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - ClawSentry is an open-source, framework-agnostic security gateway that monitors autonomous LLM agents across skill admission, intent review, execution, and post-action consequences. Its progressive review tiers substantially reduce successful attacks while largely preserving performance on benign tasks.
- Combines deterministic checks, semantic review, and bounded read-only agentic investigation, escalating only ambiguous cases.
- Reviews skill packages before first use and tracks rephrased or tool-switched retries across an entire session.
- Applies a shared policy across Codex, Claude Code, Kimi CLI, and Gemini CLI without modifying agent internals.
- On SkillInject with Codex/GPT-5.4, attack success fell from 39.55% to 2.61%, while task success decreased from 83.78% to 83.05%; across five agents, clean-skill task success remained 98.7%.