Paranoid - a reality check for AI coding agents
Stops an AI coding agent from calling a broken feature "done." Across a 42-session eval, agents left 75% of sessions on still-broken software - with Paranoid, that went to zero, for 22 cents a session.
A Claude Code plugin that blocks an AI coding agent from declaring a task "done" until a developer-owned check passes against the running application - closing the "green tests, broken feature" gap. Measured with a pre-registered 42-session eval: ungated agents ended 75% of sessions on still-broken software (while honestly reporting it); Paranoid took that to 100% of sessions ending with the check passing, at +$0.22/session, with zero false blocks on healthy code. Hardened over four adversarial AI-vs-AI audit rounds; every session row and both refuted hypotheses are published.
Inside Claude Code, in a project you trust. Then commit a .paranoid.json naming your check.



