Beyond the Payload: How User Invocation Shapes Coding Agent Vulnerability to Repository Poisoning
TL;DR - CIPR is a benchmark showing that coding agents’ vulnerability to poisoned repositories depends strongly on how users invoke them. Task type, prompt style, and supplied rules can alter both attack success and whether agents raise alerts.
- CIPR includes 1,920 instances across 20 real-world repositories, four task types, three prompt styles, and three skill/rule conditions.
- Task type produced up to a 4.5-fold difference in attack success rate.
- Test-execution tasks were a silent attack surface, combining high attack success with low alert rates.
- Underspecified prompts reduced execution depth and attack success, while noisy prompts tended to make malicious content less conspicuous and suppress alerts.