SentinelAI β Detection Rules
Ruleset version: 2026.09.1 Β· 23 rules across 5 categories.
The classifier is 100% deterministic rules + regex β no AI/LLM in the decision path.
Everything runs locally in the endpoint agent (agent/app/classifier.py + risk.py) in <5 ms.
This document is the human-readable list; the live source of truth is the running agent:
curl http://127.0.0.1:8787/health # β { β¦ "rules_version": "2026.09.1" }
curl http://127.0.0.1:8787/rules # β { rules_version, rules: [ {category, name, severity}, β¦ ] }
Rules live in code (deterministic and versioned by design); changing them is a code change + a
RULES_VERSION bump, not customer config. Org/department tailoring is config (policy.json) β
see docs/TAILORING.md.
Phase 2 adds an optional AI verification layer to reduce false positives (see
docs/AI.md). It is off by default; the rules below are always the baseline.
1. What is detected
PII
| Rule | Pattern (plain English) | Severity | Validation |
|---|---|---|---|
email_address |
name@domain.tld |
20 | β |
phone_number |
US-style phone numbers | 25 | β |
us_ssn |
123-45-6789 |
80 | β |
credit_card |
13β16 digit card numbers | 85 | Luhn checksum (rejects random digits) |
Secrets & credentials (severity 95 β always block-worthy)
| Rule | Detects |
|---|---|
aws_access_key |
AKIA⦠access key IDs |
aws_secret |
aws_secret_access_key = β¦ (40-char) |
openai_key |
sk-β¦ API keys |
github_token |
ghp_/gho_/ghu_/ghs_/ghr_β¦ |
google_api_key |
AIza⦠|
google_oauth_secret |
GOCSPX-β¦ |
stripe_secret_key |
sk_live_β¦ / rk_live_β¦ |
slack_token |
xoxb-/xoxp-/xoxa-/xoxr-/xoxs-β¦ |
slack_webhook |
https://hooks.slack.com/services/β¦ |
azure_storage_key |
AccountKey=β¦ connection strings |
npm_token |
npm_β¦ (36-char) |
private_key |
-----BEGIN β¦ PRIVATE KEY----- |
jwt |
eyJβ¦.β¦.β¦ JSON Web Tokens |
bearer |
Bearer <token> |
password_assignment |
password/passwd/pwd/secret = β¦ |
Source code (severity 55)
Fires when β₯2 code signals appear together (reduces false positives on prose):
def/class β¦, function/const/let/var β¦, import β¦, from β¦ import β¦, #include <β¦>,
Java signatures (public static void β¦), SELECT β¦ FROM β¦, and operators/markers
(=>, ::, console.log(, println!(, System.out.print).
Financial (severity 45, IBAN 70)
financial_termsβ revenue, EBITDA, salary, compensation, invoice, routing/account number, net income, forecast, P&L, gross margin, β¦ibanβ International Bank Account Numbers, validated with the ISO mod-97 checksum.
Contract / legal (severity 45)
contract_language β confidential(ity), NDA, non-disclosure, master service agreement,
indemnifβ¦, "hereinafter", termination clause, governing law, "whereas", β¦
2. How a decision is made (risk.py)
Base score = the single highest-severity finding, plus a volume bump (
+3per additional match, capped at+15).Destination weight multiplies the score by where the prompt is going:
Destination Weight Rationale ChatGPT / Claude / Gemini Γ1.0 consumer, highest exfil risk Microsoft Copilot Γ0.85 M365 Copilot Γ0.70 enterprise tenant, lower default risk Bands map the final 0β100 score to an action:
Score Decision Label 81β100 block + alert SOC Restricted 51β80 block Confidential 26β50 warn (coaching, proceed on ack) Internal 0β25 allow (log only) Public
Worked example: prompt containing an sk-β¦ key sent to ChatGPT β
severity 95, Γ1.0, +volume β 98 β block / Restricted.
Decision flow
Secrets/PII always land in the block bands via rules. The optional AI layer sits after this β it can suppress a fuzzy-category block (β allow) or raise an allow it judges sensitive (β block); see
docs/AI.md.
3. Known limitations (be upfront with stakeholders)
- Structured data only. Regex excels at keys/PII/patterns; it cannot catch semantic secrets with no fixed shape β e.g. "our Q3 acquisition target is Acme for $40M." That is the gap the Phase-2 AI layer addresses.
- No context. A public code sample and proprietary source both trip
source_code. The AI layer is designed to suppress exactly these false positives (never to add detections). - Static thresholds. Bands and destination weights are hardcoded defaults; per-org policy is an open decision (Clarification #6).
- Maintenance. Vendor token formats change; patterns need periodic review.
4. Extending
Add a rule in agent/app/classifier.py (a (name, regex) pair in _SECRETS, or a new detector
in classify()), pick a severity, and add a test in agent/tests/test_classifier.py. Keep
high-false-positive patterns gated behind a checksum (see _luhn_ok, _iban_ok) where possible.