CrewScore finds missing written guardrails in agent system prompts — injection defense, human approval, cost limits, stop conditions — offline, no API key, open rules.
It is a checklist of 23 published controls, not a quality ranking and not runtime red-teaming. Low coverage is actionable; high coverage only means the text is present.
We scanned 356 publicly collected agent prompts: 83 production-labeled prompts and 273 general-purpose prompts. Among the production-labeled subset, median coverage was 10 of 100. GPT-Store median: 0 of 100. Numbers → · Shareable card → · Live checker →
Example: control coverage 8/23 written · first gap to review: A human must approve. CrewScore checks whether controls are written down, not whether an agent obeys them.
Try it live, no install: crewscore.ai
Created and maintained by Sarosh Hussain. Pendoah is the company operating context for this project. Technical claims are grounded in the code, tests, and cited validation material.
pip install crewscore
crewscore scan .
# Gate the one control that matters:
crewscore scan . --require human_gate.approval_requiredDeterministic regex over prompt text and SYSTEM_PROMPT / system_prompt
string literals in .py / .ts / .js source. Offline, no API key, no LLM.
CrewScore is a checklist, not a benchmark. The number is the share of 23 published controls your prompt states — nothing about whether they are well specified, mutually consistent, or obeyed at runtime.
| What the prompt does | What it scores |
|---|---|
| Nothing written down | 0 |
| One control in each of the 8 dimensions | 36 |
| All 23 controls | 100 |
| One control restated five different ways | same as stating it once |
So a low score is actionable — you probably have not written down an injection policy, a human gate, or a safe-stop rule, and those are worth writing. A high score means the text is present, not that the agent obeys it. Don't rank prompts, teams, or vendors by this number, and don't treat a threshold as a safety bar. Prefer the findings to the total.
Three dimensions — Cost, Compliance, Audit — ship with known-thin
construct validity and say so. Ruleset 0.6.0 tightened their patterns against
measured corpus false positives; Compliance is still not lawful handling.
📄 The validation study → — including the arithmetic
showing our own scale was broken through 0.1.0, which we published before
fixing.
📊 Measured against 356 publicly collected prompts → — Cliff's δ = 0.614 separating production-labeled agent prompts from general-purpose ones, generated by a committed harness rather than typed by hand.
crewscore scan . # prompts + AGENTS.md + inline SYSTEM_PROMPT=...
crewscore scan . --no-inline # file discovery only
crewscore init . # prompt-free regression baseline + PR workflow
crewscore scan . --fail-on-regression --baseline .crewscore-baseline.json
crewscore scan . --require human_gate.approval_required # one-control CI habit
crewscore test --prompt-file ./prompt.md # coverage N/23 + first gap to review
crewscore fix --prompt-file ./prompt.md --plan # what's missing, no writes
crewscore rules --concepts # the 23 controls, and the rules behind themFull CLI reference → · How scoring works →
Start report-only — it never fails a build:
- uses: shmindmaster/crewscore@v2
with:
scan-path: "."When you know which controls your workflow actually needs, name them — the build then fails only when a named control is missing, never on the coverage average:
- uses: shmindmaster/crewscore@v2
with:
scan-path: "."
required-controls: "human_gate.approval_required,safe_stop.stop_condition"
sarif: "crewscore.sarif"Posts a sticky PR comment with the open rule findings (give the job
permissions: pull-requests: write; without it the action logs a warning and
moves on). Failed gates surface as ::error annotations naming the control.
Guard downstream steps on the scored output, not on score — an empty
score casts to 0.
Action inputs, outputs, and the CLI variant →
CrewScore judges two kinds of file, and tells you which it thinks it is looking at. Detection is by filename and path — never by sniffing content.
| Artifact | Examples | Judged on |
|---|---|---|
| Coding-agent config | AGENTS.md, CLAUDE.md, .cursorrules |
Configuration smells |
| Agent system prompt | system-prompt.md, anything under prompts/ or agents/ |
The 8 governance dimensions |
A file saying "always use pnpm" is telling a coding agent how to work in your repo. It has no reason to contain HIPAA language, and scoring it against that is a category error.
We know the size of that error because we measured it: against the 100
most-starred repos with an AGENTS.md
(arXiv:2606.15828), the governance ruleset
put all 100 in the worst tier. A scale the entire population fails carries
no information. So config files get a smell verdict instead — and in --json,
no governance grade at all.
crewscore test --prompt-file AGENTS.md
# -> CONFIG: NO SMELLS DETECTED (not "0/100 CRITICAL GAPS")Problems in the shape of an instruction file rather than its content, from a published catalog — Configuration Smells in AGENTS.md Files (dos Santos et al., 2026), which found 91 of 100 popular projects carried at least one.
| Smell | Heuristic | Found in |
|---|---|---|
| Context Bloat | ≥ 200 lines | 42% of studied projects |
| Lint Leakage | Style rules a configured linter already enforces | 62% |
| Init Fossilization | Tracked by git with exactly one commit | 24% |
The paper's other three smells need an LLM to detect. We would rather ship three honest detectors than six approximate ones. Lint Leakage is an approximation of the paper's detector and says so in its output; Init Fossilization cannot tell "never needed revising" from "never got revised."
Smells never change the score. Folding them in would silently change what
every existing --threshold means.
git clone https://github.com/shmindmaster/crewscore.git
cd crewscore
pip install -e ".[dev]"
pytestDevelopment guide → · AGENTS.md · CONTRIBUTING.md
| Validation | What the number does and does not measure |
| Corpus validation | Generated result over 356 publicly collected prompts |
| Scoring and controls | Formula, 23 controls, charter, governance |
| CLI | Every command and flag |
| GitHub Action | Action inputs/outputs and CLI-in-CI |
| Policies and SARIF | Regression and required-control CI without score gating |
| Architecture | Modules, data flow, lean target |
| Development | Local setup, rules, packaging, media |
| Live eval handoff | Promptfoo / garak after structural gate |
| Roadmap | Available work and deliberately deferred capabilities |
| Security | Private vulnerability reporting |
| Community discussions | Questions, adoption feedback, and open-ended ideas |
| Comparison | Other tools, and what to use after this one |
| CHANGELOG | Including every scoring change and its measured delta |
Live adversarial red-teaming · runtime tool-gate enforcement · a security or compliance certification · proof the model will obey the text.
Roadmap: framework adapters that extract prompts from LangGraph / CrewAI / AutoGen graphs; optional live adversarial testing (post-traction, not the default path).
MIT licensed.