The SourceVault Gateway
The Gateway is the governed-access layer every license and the 7-day trial carry: an access policy, secret redaction, a tamper-evident audit chain, signed AI-change provenance, and named agent identities, enforced once on every surface. This page is the operator's reference. The product ships the same material in its own docs (GATEWAY.md, SENTINEL.md, COMPLIANCE.md, PROVENANCE.md) with every install.
On this page: components · turning it on · writing a policy · redaction · the audit chain · provenance · agent identity · the console · the honest-use boundary · configuration
The components
Each component ships in the engine's data path, not beside it. Every surface (the dashboard, the MCP server, the signed machine API, and the Hermes plugin) calls the same tool layer, so one policy evaluation governs them all and every call lands in the same audit chain.
| Component | What it does |
|---|---|
| Sentinel | Allow, deny, or redact rules over repo-relative paths, plus a gitleaks-style DLP pass that scrubs secrets from every file read, search result, and answer before content leaves the engine. |
| Audit log | One hash-chained JSONL record per ask, search, file read, settings change, and policy decision, anchored by a signed head so truncation and deletion are detectable. Verifiable offline. |
| Provenance | Flags AI-authored, security-relevant commits and issues signed DSSE attestations, verifiable against the install's ed25519 key. |
| Agent identity | Named, scoped agent tokens, so the audit chain answers who asked, not just what was read, and policy rules can be written per agent. |
Licensing: the Gateway is granted by a single gateway license feature, which every license carries. The 7-day trial runs it on one repository; when the trial ends, the engine stays free on that repository and the Gateway waits for a key. Keys issued before the bundle existed, naming the components individually, still verify.
Turning it on
Each component has a toggle under Settings → Features: Sentinel, Audit log, and Provenance. They are off by default on a fresh install, and the write is refused without the entitlement. When Sentinel is off every gate is a no-op and content flows as before; the settings reference covers each toggle and its environment variable.
Enabling Sentinel without the audit log still yields the decision log: enforcements are recorded whenever either toggle is on.
Writing a policy
Policy lives in .sourcevault/policy.json at the state root. It is re-read within a few seconds of a change, with no restart.
{
"default": "allow", // "allow" | "deny" | "redact"; applied when no rule matches
"rules": [
{ "match": ["**/*.pem", "**/secrets/**"], "action": "deny", "reason": "PCI scope" },
{ "match": ["infra/prod/**"], "action": "deny", "repos": ["backend"] },
{ "match": ["src/payments/**"], "action": "redact", "reason": "card handling" },
{ "match": ["billing/**"], "action": "deny", "subjects": ["ci-reviewer"], "reason": "CI reads no billing" }
]
}
| Field | Meaning |
|---|---|
match | An array of globs. A rule fires if any glob matches the file. |
action | deny blocks the read and drops the chunk from search results; redact returns the content as [[redacted:policy]]; allow passes it. |
reason | Optional. Surfaced in the policy_denied error and in the decision log. |
repos | Optional. Restricts the rule to the named indexed repositories. |
subjects | Optional. Restricts the rule to named agent identities: minted agent names, operator (the dashboard), or shared-secret (signed HTTP callers using the shared secret). |
Rules are checked in order and the first match wins; if none match, default applies.
Subject scoping fails closed, like repo scoping. A subject-scoped deny or redact also applies to a caller with no identity at all, so a fenced file cannot slip through by arriving without a name; write an explicit allow above it if unattributed traffic should pass. A subject-scoped allow never fires for a caller that cannot prove it is that subject.
Glob syntax
Matching is over POSIX-style, repo-relative paths and is case-insensitive, so secrets/** also blocks SECRETS/… on macOS and Windows.
| Pattern | Matches |
|---|---|
* | Any run of characters except / |
? | One character except / |
** | Anything, crossing / boundaries |
**/ | Zero or more leading directories: src/**/x matches src/x and src/a/b/x |
secrets/ | A trailing slash means the subtree; equivalent to secrets/** |
Fail-safe behaviour
- Repo scoping fails closed. If a chunk's repository cannot be determined (a multi-repo ask, say), a repo-scoped
denyorredactstill applies, but a repo-scopedallowdoes not. - A malformed
policy.jsonkeeps the last good policy and logssentinel_policy_parse_failed. It never silently reverts to allow-all. - A single unparseable glob is dropped on its own (logged
sentinel_policy_rules_ignored) rather than collapsing the rest of the rule. A rule whosesubjectsfield is malformed is dropped entirely rather than loaded unscoped. - No policy file means default-allow with an empty ruleset: Sentinel is on but unconfigured.
Delegated tasks
A task handed from one agent to another through the relay targets a repository rather than a file, so it is evaluated against the repository root, path "". Only a rule whose glob matches the empty path applies: **, optionally scoped by repos or subjects, or the policy default. A deny on ** for a repository refuses every hand-over against it; a redact lets the task run but redacts the free text of what comes back. Path-scoped rules never block a hand-over; they keep governing the reads made during the task, and results are DLP-scrubbed on delivery regardless.
Secret redaction
When a file or chunk is not denied or policy-redacted, its content is scanned for secrets and every match is replaced with a generic [[redacted]] marker. The marker is deliberately typeless so it does not disclose which kind of secret sits where; the rule that fired is recorded only in the decision log. Redaction preserves line structure, paths, and line ranges, so citations keep resolving.
Detected out of the box, all locally and with no network verification: PEM private keys, AWS access keys, GitHub tokens, Stripe keys, Slack tokens, Google API keys, OpenAI and OpenAI project keys, Anthropic keys, JWTs, and a generic rule for name = "value" assignments whose value clears an entropy threshold, so password = "changeme" does not trip it.
Secrets straddling a read cap or a chunk boundary are covered by an overlap scan. A single secret larger than the overlap is an accepted residual, and the product says so.
The pass runs on everything leaving the engine: file reads, search results and previews, answers (including retrieval, context expansion, the symbol and history graph legs, and cached answers), MCP tool results and resources, and runbook exports.
The audit chain
With the audit log on, one JSONL event is appended per ask, search, file read, settings change, token rotation, license install, repository delete, provenance scan, relayed task, and policy decision, to monthly files under .sourcevault/audit/. Events record repository names, paths, question text, and cited files, never file contents, and live in the local state directory with the rest of your data, so retention is yours to set.
Each record carries prev_hash (the previous record's hash, or a genesis zero-hash) and hash = SHA-256(prev_hash + record), and the chain is anchored by a signed head. A deleted, reordered, or edited line breaks the chain; a truncated or missing file is detectable from the head. Verification recomputes the chain and reports the first break:
GET /api/dashboard/audit?month=YYYY-MM # read a month's records GET /api/dashboard/audit/verify?month=YYYY-MM # recompute the chain; verdict plus the first break, if any
Verification works offline and needs nothing from us. Enforcements are logged; plain allows are not. The artifact is the record of what was stopped, and of what the AI read.
Tamper-evident, not tamper-proof. The log is written by a process the operator controls, which is the honest-use boundary described below. Logging fails closed: if the append lock cannot be taken, the engine does not write an unprotected record.
AI-change provenance
With provenance on, commits are classified for AI authorship and security relevance (path rules plus a pass over the diff itself) and recorded as signed DSSE attestations that anyone holding the install's public key can verify. A drift correlator ties index drift back to the commits that caused it, and the MCP server exposes the same records as list_provenance_alerts and get_attestation, so an agent can gate a merge on unresolved AI drift. The product's PROVENANCE.md documents the attestation format.
Agent identity
Every Gateway surface can answer what was accessed; identity makes the audit chain answer who. Named, scoped agent tokens are minted from the console or the CLI:
sourcevault agent mint claude-mcp --surfaces mcp --repos backend,frontend sourcevault agent list sourcevault agent revoke claude-mcp
The plaintext token is shown exactly once; only its SHA-256 is stored, in agent-tokens.json in the state directory with mode 0600. Scopes bound surfaces (http, mcp, relay) and repositories. A request outside its scope is refused at the door, and the refusal is itself an audit record.
| Surface | How the agent presents its identity |
|---|---|
| MCP | Set SOURCEVAULT_AGENT_TOKEN in the MCP client's server environment. The server refuses to start if the token is set but unknown, revoked, or not scoped to mcp; it is never silently unattributed. See the MCP setup page. |
| Signed HTTP and the relay | Add the -agent: <name> variant of the signature header and sign with sha256(token) instead of the shared secret. Naming an agent is a commitment: an unknown name or an out-of-scope surface refuses, with no fallback to the shared secret. |
Audit records gain a subject field (old records verify unchanged). Existing shared-secret setups keep working and attribute as the reserved shared-secret subject; the dashboard operator is operator. Policy rules can be scoped to subjects, closing the loop from who asked to who may.
Minting and revoking are Gateway-licensed operations. Verification of already-minted tokens deliberately is not, so a lapsed license never locks agents out of their own vault.
The operator console
The dashboard's Gateway view reads the records the Gateway writes: agent activity by surface and repository, Sentinel denials and redactions with the rule that fired, and the audit-chain verdict in plain language (what truncated, file_missing, and head_signature_invalid mean). It is a pure aggregation of the log already on disk; viewing it records nothing.
The honest-use boundary
The Gateway governs honest use and produces evidence. It is not containment. Policy and DLP run in a process the operator controls, so they gate agents and tools acting through SourceVault's surfaces, not a hostile local process that can already read the filesystem directly. The MCP server's trust boundary is the same as the CLI's: whoever can start a process on the machine. This boundary is stated wherever the Gateway is described, deliberately, because the audit chain's value to a security review rests on the claims around it staying exact. Position it as policy plus DLP plus audit, not as a sandbox.
Verify it is working
With Sentinel on and a deny rule for secrets/**, the code-read script runs the same file-read tool the MCP server uses:
npm run code-read -- -r demo secrets/prod.env # → error: policy_denied npm run code-read -- -r demo src/config.js # → const key = "[[redacted]]"; grep policy_decision ".sourcevault/audit/audit-$(date +%Y-%m).jsonl"
Configuration
| Environment variable | Default | Purpose |
|---|---|---|
SOURCEVAULT_POLICY_CACHE_MS | 5000 | How long a parsed policy.json is cached. |
SOURCEVAULT_SETTINGS_CACHE_MS | 5000 | TTL of the settings snapshot the gates read; a toggle also invalidates it immediately. |
SOURCEVAULT_DLP_ENTROPY_MIN | 3.5 | Bits per character before a generic assignment is redacted. |
SOURCEVAULT_AUDIT_LOCK_WAIT_MS | 3000 | Maximum wait for the cross-process audit append lock before proceeding, logged. |
SOURCEVAULT_AUDIT_LOCK_STALE_MS | 10000 | Age past which a lock is treated as a crashed holder and taken over. |
The feature toggles and their environment variables are on the settings page; the MCP server and its agent token are on the MCP page; the reviewer's view of all of this is the trust page.