Architecture series: Overview · Native Scan · Agentic Scan · Data & Storage · Server & API
1. Multi-Tenancy: the project_uuid spine
All scan data is partitioned by project, a named container with a UUID, optional config overlay, and optional access-control lists. There is no separate database per project; isolation is a project_uuid column on every major table, filtered on every read and stamped on every write.
- Default project:
00000000-0000-0000-0000-000000000001, created duringvigolium init. Used whenever no project is selected. - Selection precedence:
--project-uuid>--project-name>VIGOLIUM_PROJECT_UUID>VIGOLIUM_PROJECT_NAME> default. On the server, theX-Project-UUIDrequest header plays the same role. The legacyVIGOLIUM_PROJECTenv var was removed in v0.4.3. - Project config is a partial YAML overlay (same shape as a scanning profile) at
~/.vigolium/projects/<uuid>/config.yaml; only the keys it sets are overridden. - Access control: a project is a data boundary, not an auth boundary — the server’s API key authenticates,
X-Project-UUIDselects, and any authenticated caller can address any project. Isolate engagements with separate servers or databases.VIGOLIUM_PROJECT_READONLY=truedisables the mutatingprojectCLI subcommands.
Tables carrying project_uuid
scans · http_records · findings · scopes · source_repos · oast_interactions · scan_logs
Existing databases are migrated in place, the column is added with the default-project UUID as its default, so pre-multi-tenancy data lands in the default project.
2. The database backend
Vigolium uses the repository pattern over Bun ORM, with two interchangeable backends:
The schema is intentionally denormalized, there are no separate hosts or parameters tables; JSONB columns carry structured sub-data. This keeps a single
http_records row self-contained and avoids join fan-out on the hot ingestion path.
Core models (pkg/database/models.go)
HTTPRecord (table http_records), the unit of ingested traffic:
- Identity:
UUID(PK),RequestHash(SHA-256 of the raw request, used for per-source dedup),ParentUUID(chains a followed redirect hop to the record that produced it) - Host:
Scheme,Hostname,Port,IP - Request:
Method,Path,URL,RequestHeaders(JSONB),RawRequest,RequestBody - Response:
StatusCode,ResponseHeaders(JSONB),RawResponse,ResponseBody,ResponseTitle,ResponseWords,ResponseLocation(a 3xx’sLocationheader, verbatim),ResponseTimeMs - Derived:
Parameters(JSONB array ofEmbeddedParam),RiskScore,SurfaceScore,Technology,Remarks - Metadata:
Source,SentAt,ReceivedAt,CreatedAt
host_observations is a per-host history table: a re-probe appends a row rather than overwriting what the last sweep saw, so a host’s posture over time stays readable.
New columns arrive through the existing self-healing DDL path, so an existing database gains them on the next open with its rows untouched.
currentSchemaVersion is deliberately not bumped for a column that needs no backfill - bumping it would re-run the O(rows) backfills on every database in the field.Read commands (traffic, finding, db ls, export) migrate on open rather than failing an old database’s first read with a bare no such column. A read-only handle cannot migrate, so it names the missing columns and how to proceed instead.Opening is not creating (v0.4.7). A read-write open still writes —
mkdir -p on the parent, a journal_mode PRAGMA, a WAL checkpoint — but a pure read against an explicitly pinned --db or $VIGOLIUM_DB_PATH no longer brings a store into existence: a missing path is source_missing and a foreign SQLite file is source_incompatible, both exit 1. replay and fuzz are exempt, because they record the traffic they send, and the built-in default database is still created on first use.--read-only now covers the whole promise: no journal-mode flip, no checkpoint, no PRAGMA optimize on close, and no -wal/-shm siblings left beside the file. The source’s SHA-256 is unchanged after the read.Finding (table findings), a detected issue:
- Identity:
ID(auto-increment),FindingHash(unique constraint → dedup key, set from theResultEventID) - Module:
ModuleID,ModuleName,Description,Severity,Confidence - Evidence:
MatchedAt(JSONB),ExtractedResults,Request,Response,AdditionalEvidence(merged-duplicate request/response pairs, capped at 10) - Relations:
HTTPRecordUUIDs(JSONB),ScanUUID; thefinding_recordsjunction table is the many-to-many link to HTTP records.
findings table, tagged by source so native and AI results coexist and dedup together.
Converters (pkg/database/converters.go)
The in-memory scan types never touch the DB directly. HTTPRecord.FromHttpRequestResponse() and Finding.FromResultEvent() are the only seam, they generate UUIDs, compute hashes, parse URLs, extract titles, and count words, keeping persistence concerns out of the executor and modules.
3. Write paths
Async batched ingestion, RecordWriter
High-throughput ingestion (proxy capture, bulk import, spidering) does not call the repository synchronously. pkg/database/record_writer.go fronts it with a buffered channel:
WriteResult{UUID, Err} back on a per-request result channel, backpressure is the channel capacity, ordering is preserved, and the DB sees large batched transactions instead of a write per request.
SQLite DSN note: the modernc driver needs pragmas in
_pragma=name(value) form; the mattn-style _busy_timeout= is silently ignored. Relevant when tuning concurrent-writer behavior.4. Deduplication
Two layers, by design:- Per-source HTTP-record dedup:
RequestHash(SHA-256 of the raw request) plusDeduplicateRecordsBySourcecollapses re-ingested identical requests within a source. - Finding dedup: the
finding_hashunique constraint prevents exact duplicates at insert time;DeduplicateFindings()runs after a phase to group near-duplicates (same module/severity/URL), folding the extra request/response pairs into the survivor’sAdditionalEvidence(capped). The multi-driverauditcommand runs an additional project-wide findings dedup pass once its drivers exit.
5. Cloud storage (optional)
Storage is disabled by default. When enabled (storage.enabled: true), a single minio-go S3 client talks to GCS (HMAC), S3, or self-hosted MinIO, the driver differs, the rest is identical.
storage.ValidateKey rejects .., backslashes, absolute paths) and project-prefixed server-side, so one bucket safely holds many projects and clients cannot reach outside their own.
vigolium import auto-detects its input: a vigolium-audit folder, a JSONL export, a .tar.gz/.zip archive of either, or another vigolium SQLite database (detected by its magic header). A SQLite input is a lossless, idempotent SQLite→SQLite merge — records, findings, scans, agentic scans, and OAST interactions are deduped on their natural keys and each row keeps its original project, so re-importing the same database is a no-op. The destination is the --db target (or the default database).
gs:// URLs are first-class inputs/outputs: vigolium import gs://… downloads-then-imports, and any export -o gs://… writes locally then uploads on success. The {ts} and {project-uuid} placeholders expand in any -o path. The bundle export format round-trips a full snapshot (JSONL + HTML report + manifest + agent session dirs) that another machine can re-import.
Related
Projects API
Project CLI/API recipes and access-control management.
Storage API
Full
vigolium storage command and gs:// reference.Native Scan
Stage 11 traces a finding from
ResultEvent to row.Configuration
The
storage: and project config YAML blocks.