SentinelAI β AI Verification (on-device, optional)
The MVP decision path is deterministic rules (docs/RULES.md). This adds an optional AI layer
that runs a small model on the end-user's laptop to cut false positives β without changing the
low-friction, privacy-preserving story.
Principles
The AI works in both directions, but is bounded so it can never cause a leak:
- SUPPRESS false positives. On a fuzzy category (
source_code,financial,contract) that the rules flagged, the model may downgrade block/warn β allow if it's confident the match is benign/public. It can only relax these categories. - DETECT what rules miss. On a prompt the rules allowed, the model may raise it to block if it spots sensitive/proprietary content that has no fixed pattern (internal plans, codenames, M&A, org changes). Detection can only raise, so a wrong model over-blocks (a UX cost) but never leaks.
- Hard matches are untouchable. Secrets, SSNs, credit cards, and IBANs always block via
rules β the AI is never consulted for them (
SUPPRESSIBLE_CATEGORIESinrisk.py). - Off by default. A toggle in the extension popup (
aiVerify) turns it on. The demo stays pure-rules and deterministic unless you flip it. - On-device + fail-safe. Inference runs locally via Ollama; prompt text never leaves the machine. Any error/timeout/absent model β the deterministic rules decision stands.
Flow
extension popup toggle ββΆ background.js adds {ai_verify:true} to /classify
agent: rules score ββΆ if ai_verify:
β’ decisionβ{warn,block} AND category fuzzy β verify_false_positive β maybe SUPPRESS to allow
β’ decision==allow β detect_sensitive β maybe RAISE to block
thresholds: suppress β₯70 conf, detect β₯75 conf (detection is more conservative)
reasons: ai_suppressed:β¦ / ai_detected:β¦
Try it: ./scripts/demo-ai.sh runs both directions against the local model.
Model β "easily portable to a laptop"
Runtime Ollama (nothing to host; it just runs locally). Default model llama3.2:3b
(Q4, ~2 GB, CPU-friendly, cross-platform). For lighter machines use llama3.2:1b or
qwen2.5:1.5b. The task is a narrow binary judgment, so a 1β3B model is plenty.
Note: "Kimi" (Moonshot) is a large MoE, not an on-device SLM β not a fit for endpoints.
Enable it (per laptop)
# one-time
curl -fsSL https://ollama.com/install.sh | sh # or the Ollama desktop app
ollama pull llama3.2:3b
# the agent auto-detects Ollama on 127.0.0.1:11434; then flip the popup toggle
Configuration (agent env)
| Var | Default | Purpose |
|---|---|---|
OLLAMA_URL |
http://127.0.0.1:11434 |
Ollama endpoint |
OLLAMA_MODEL |
llama3.2:3b |
local model (swap for 1b/qwen to taste) |
OLLAMA_TIMEOUT |
5.0 |
seconds before falling back to rules |
SENTINEL_AI_MIN_CONF |
70 |
min confidence to suppress a false positive |
Tailoring to your organization (Epic H1)
The base model only knows "familiar to the internet," not "proprietary to us." A per-org /
per-department policy (agent/policy.json, keyed by SENTINEL_ORG_ID, default fallback)
makes suppression org-aware. It is edited centrally and pushed to endpoints β laptops run
inference only. Edits are picked up live (mtime reload, no agent restart).
{
"default": {
"context": "Kshetra Studio, a software/AI product company (engineering org).",
"proprietary_markers": ["orion", "vega", "kshetra-internal", "projectx"],
"public_hints": "Standard textbook algorithms and public/OSS code are NOT proprietary."
}
}
Two effects:
context+public_hintsare injected into the judge prompt, so the model reasons for your org ("anything tied to Kshetra Studio is not a false positive").proprietary_markersare a hard guard: any text containing a marker is never suppressed, regardless of what the model says (the model isn't even called). Verified:Input Result public quicksort allowβai_suppressed: public sorting algorithm (conf 90)def orion_deploy(...)blockβpolicy:proprietary_marker 'orion' β not suppressed
Central distribution (agent fetches the bundle from the backend instead of a local file) is Epic H2.
Roadmap
- Cloud option (deferred): route verification through the SaaS Lambda β Amazon Bedrock for
laptops that can't run a local model. Trade-off: prompt text leaves the device, so it's opt-in
per policy. The
verify_false_positiveinterface is provider-agnostic, so this is an added implementation, not a rewrite. - Per-category confidence thresholds; user "was this wrong?" feedback to tune.