The single source of truth for KPubData Builder's (package kpubdata-builder) HTTP wire contract is contract/builder-api.yaml.
- Endpoints, request bodies, response bodies, status codes and security schemes follow the OpenAPI document.
info.versionincontract/builder-api.yamlmust matchkpubdata_builder.service.API_CONTRACT_VERSION.tests/unit/test_service_contract.pyverifies the version match, static route/status alignment, and wire-level conformance of actualdispatch()responses.- Consumers like Studio base compatibility on the OpenAPI SSOT and
GET /version, not on this document. - Version bumps, Studio compatibility ranges, and release freeze procedures follow ADR 0013.
- The contract version on
mainis always a stable SemVer; the OpenAPI document and the code use the same value. - Additive wire changes are minor, contract-error fixes that preserve existing semantics are patch, breaking changes are major.
- Same-wire changes (examples, descriptions, internal refactors) do not bump the version.
- The two rules above are CI-enforced (#693). On every PR,
scripts/check_contract_compat.pycompares against the base branch's contract and fails when a normative change (anything other than description, summary, example,x-*) does not raiseinfo.version, or when a breaking change — removing an operation, status code, media type, property, schema or enum value; changing a type or$ref; adding a new required parameter — arrives without a major bump. The explicit approval of an intended breaking change is the major bump itself. - Studio checks the same major plus a per-feature minimum SemVer, not exact equality. It bumps its schema/client and minimum feature version only when it actually consumes a new operation.
- Completing Epic #484 is the point to freeze the final contract and record it in the release manifest and tag, not to defer version changes during development.
These are the rules the reading side follows. They pair with the server-side rules above.
- Minor versions may add optional fields to responses. Unknown fields are ignored. The contract's
additionalProperties: falsedescribes what this version of Builder sends, not an instruction for the client to reject extra fields. A strict response parser breaks on additive changes — Studio'ssilverColumnInfoSchema.strict()did exactly that when 1.30.0 addedlogical_type/wire_encoding(#735). - Required field names and types do not change until the next major. A missing required field or a wrong type is still rejected. Leniency applies to unknown fields only.
- Unknown enum values (e.g. a new
wire_encoding) are treated as opaque text, not guessed as numbers. logical_type: identifier(1.61.0, #702) is a text column the kpubdata spec declares assemantic_kind: code(legal district codes, PNU, postal codes). Itswire_encodingis alwaysstring; a client must not coerce it withNumber()— leading zeros vanish and JOINs misalign. Which columns are identifiers comes from the kpubdata declaration; Builder never guesses from the values.
These rules ship as fixtures so machines can check them. contract/fixtures/responses.json holds three bodies for every 2xx named response example in the contract:
| Key | Meaning | Client expectation |
|---|---|---|
current |
Exactly what this contract version sends | Pass |
with_additive_fields |
future_optional_field added to every object that declares properties (additive_paths marks the location) |
Pass |
required_type_broken |
One top-level required field (broken_path) with its type changed |
Reject |
contract_version in the file header is the contract version the fixtures were generated from. scripts/generate_response_fixtures.py generates this file from the contract; it is not hand-edited. tests/unit/test_response_fixtures.py checks that the committed file matches regeneration, that current conforms to the contract, that additive-field bodies are rejected by a strict parser and accepted by a lenient one, and that required-field type errors are rejected by both. When a contract example changes, re-run the script and commit both. Studio's contract tests read this file to check their own response schemas (companion issue).
Error responses carry the same three bodies separately in error_fixtures (1.72.0, #947, #951), split from fixtures so a client that maps 2xx to a success parser never receives an error body. Every non-2xx named example appears once — operation-specific responses under operation_id/method/path/status, shared responses (components.responses: Unauthorized, SignupNotApproved, etc.) once under response/status. Examples: saveRevision/revertRevision 409 RevisionConflict (with current_revision), SignupNotApproved 403 SignupPending/SignupRejected. SignupNotApproved is not attached per-operation (many already declare their own 403); it is written on bearerAuth and Error.code as something any authenticated operation can return, with the status code in x-status.
This document is the operational guide humans read. It does not restate wire shapes.
Every enum in the contract belongs to Builder (Independence Rule 7). Even when the values currently match kpubdata's, Builder is the owner, and kpubdata enum values are never passed through to the wire as-is. src/kpubdata_builder/service/vocabulary.py explicitly maps each kpubdata value to a Builder value, and any value not in the mapping becomes the declared fallback.
| Contract field | Builder vocabulary | Current value source | kpubdata value not in mapping |
|---|---|---|---|
DatasetStatusAxes.access |
AccessStatus |
kpubdata probe classification + unknown |
unknown |
CatalogDataset.representation |
Representation |
kpubdata Representation |
other |
CatalogDataset.operations[] |
Operation |
kpubdata Operation |
Removed from the list |
CatalogQuerySupport.pagination |
PaginationMode |
kpubdata PaginationMode |
query_support becomes null entirely |
When kpubdata adds values, the wire does not change. Builder decides what to call the new value in this mapping; adding a value to the wire bumps the contract version. tests/unit/test_wire_vocabulary.py verifies the mapping matches the contract enums, covers every value of the installed kpubdata, and that unknown values become the fallback.
The v0.4 Builder service keeps a synchronous execution model.
| Scope | Model | Reference |
|---|---|---|
/validate, /preview, /build, read endpoints |
Request-response, synchronous | ADR 0002 |
Async job model (POST /builds, GET /builds/{run_id}, POST /builds/{run_id}/cancel) |
Accept and return immediately; poll for status | ADR 0008 / #334 |
Principles:
POST /buildruns the pipeline within the current request and returns the success or failure result.POST /builds/GET /builds/{run_id}/POST /builds/{run_id}/cancelfollow ADR 0008 (#334): state machine, idempotency, cooperative cancellation, and partial artifact rules.- When Medallion stage-specific artifact/preview access is needed, a stage-specific endpoint is added to the OpenAPI SSOT first.
Detailed schemas and status codes follow the OpenAPI SSOT. The policy-level semantics are:
| Situation | Policy |
|---|---|
| Successful request | Per-endpoint success response, 200 |
| BuildSpec parse/load failure | Input error the client can fix |
| BuildSpec validation failure | Input error including a problem list |
| Preview source failure | The HTTP request itself can succeed; per-source status/error carries the failure |
| Build source failure | Treated as an upstream/source dependency failure; manifest preserved where possible |
| Per-run BuildSpec retrieval | GET /builds/{run_id}/spec returns the redacted canonical YAML and its bytes' digest |
| Built dataset retrieval | GET /datasets, GET /datasets/{dataset_id}, GET /datasets/{dataset_id}/runs return grouping by BuildSpec.dataset_id, latest run, and run history |
| Stage summary/preview | GET /builds/{run_id}/stages, GET /builds/{run_id}/stages/{stage} return per-source Bronze/Silver/Gold status with safe summaries |
| Structured quality/drift | GET /builds/{run_id}/quality returns per-source quality_results/schema_drift for the run; GET /datasets/{dataset_id}/quality/history returns the dataset's per-run PASS/WARN/FAIL aggregate history |
| Read-only query | POST /query registers the server-resolved Silver/Gold table as the logical dataset and runs within separate capacity/timeout limits |
| Composition (join) | When BuildSpec.composition is present, the POST /build response exposes a separate composition key (the join result) alongside per-source outcomes |
| Async job cancellation | POST /builds/{run_id}/cancel moves a queued job directly to cancelled before execution, and a running job through cancelling to cancelled at a safe stage boundary. Terminal jobs or jobs confirmed to have exited normally return 409 |
| Partial artifacts of a cancelled run | Preserved with a partial manifest (status: cancelled, partial: true). Stages that never ran are not recorded as successes; a cancellation is never labelled a failure, nor a failure a cancellation |
| Authentication failure | 401 means re-authenticate, 403 means request permission, 503 means a temporary JWKS outage. Repeated 401 from the same client becomes 429 (code: "auth_throttled", retry_after_seconds) and is rejected without attempting authentication |
When policy and implementation disagree, do not add implementation footnotes. Fix in this order:
- If the intended contract is different, update
contract/builder-api.yaml. - If it is an implementation bug, fix the service code and conformance tests.
- If Studio is affected, open a separate Studio client/docs PR.
- The
specfield in HTTP requests is a YAML string, as now. The OpenAPIBuildSpeccomponent defines the canonical domain structure that YAML represents, so type generators can read it. metadata,sources[].params,exports[].optionsvalues are standard JSON-compatible.- Source preview returns
source_key,status,error,schema,sample,total_rows,statisticsfor both success and failure. A failed source within HTTP 200 is represented bystatus: failed, empty schema/sample, zero-based statistics, and a stringerror. - The removed
transforms, top-levelnormalization_mode, andsources[].normalization_modeare not contract fields; the parser rejects them explicitly. - A spec that passes validation is atomically saved to
{output_root}/{run_id}/buildspec.yamlbefore entering the pipeline. For legacy runs without a snapshot, the API returns an unavailable404rather than guessing from the manifest. - Snapshot redaction applies only to explicitly-mapped credential keys. Inline secrets are replaced with
<redacted>, so a snapshot alone cannot re-execute a run that needs credentials; credentials must be supplied again from environment/service configuration. spec_digestis the SHA-256 of the redacted canonical snapshot bytes actually stored, not of the original object. Two specs that differ only in credential values intentionally share the same snapshot/digest.
- Identity: a built dataset's identity is solely its
BuildSpec.dataset_id. Directory names and source catalog names are never used to guess. Legacy runs without abuildspec.yamlsnapshot (#487) have no guessabledataset_idand are silently excluded fromGET /datasets*grouping (they still appear inGET /builds). - Latest run: each dataset is summarized by the run with the latest
finished_atamong runs the principal can access. Ties at the samefinished_atare broken deterministically by descendingrun_id. Ownership filtering applies before latest-run selection, so another user's run with the samedataset_idnever appears in latest candidates or metadata. - row_count: a multi-source dataset's row count is not collapsed to a single scalar.
row_counts(a per-source_key map) andtotal_row_count(the sum) are provided together. - quality: the
qualityfield is alwaysnulluntil #486 (structured quality gates) is implemented. Log-only quality warnings are not arbitrarily converted to PASS/WARN/FAIL — unevaluated is not PASS. - stage status: four states —
completed/failed/not_run/unavailable. Success is never guessed from filesystem presence alone; manifest failure records and sidecar completeness are checked together so partial/failed runs still distinguish "Bronze succeeded → Silver failed → Gold never ran". - Secret/path non-exposure: Bronze
fetch_params/provenance.fetch_params, Gold exportoptions/output_path, and every response never expose filesystem paths. Silversampleis returned within the preview limit persisted at build time (default 5 rows) and does not read the full Parquet. dataset_idcan contain characters that cannot be used verbatim in a URL path (slashes, spaces), so clients must percent-encode thedataset_idinGET /datasets/{dataset_id}. No new constraint is added toBuildSpec.dataset_iditself for this API.
- The manifest is the source of truth:
quality_results(per-source_key list ofQualityCheckResult) andschema_drift(per-source_key list ofSchemaDriftFinding) are recorded in manifest.json at build time;GET /builds/{run_id}/qualityexposes them as-is. Nothing is recomputed. availabilitydistinguishes "zero checks" from "never computed" (#514): an emptyquality_results: {}alone cannot tell whether there were no rules to evaluate or whether the quality stage never ran.GET /builds/{run_id}/qualitynow also returnsavailability(available/partial/unavailable) andevaluated_checks(an integer).availablemeans every source the run attempted (manifest.inputs) has quality results, thoughevaluated_checksmay be zero (no rules configured).partialmeans only some sources have results (e.g., one source's Silver failed in a multi-source run).unavailablemeans no results at all — including legacy runs predating #486 (noquality_resultsfield) and new runs where the field exists (as{}) but no attempted source has a result (e.g., every source failed before the quality stage — the manifest writer always records an empty{}even when nothing was computed). This field is additive; the existingquality_results/schema_driftshapes are unchanged.- Full preservation including PASS:
QualityCheckResultcontains only checks that were actually evaluated, including PASS. Rules not configured or not evaluable (missing column, denominator zero) are excluded entirely — never faked as PASS. - Extended rule condition preservation: a
rangeresult'sthresholdpreservesmin/max; acompare_columnsresult preservesoperator/right_column. When the column exists but the dtype is incompatible, the rule is not omitted — a WARN/FAIL at the declared severity is recorded with a safedetail. - WARN/FAIL gate: WARN lets the build continue and records the result. FAIL fails the source before Gold; the already-computed
quality_resultsare still preserved in the manifest. - Preview/Build same evaluation: each
SourcePreview.quality_resultsfromPOST /previewis the result of the same evaluator (quality.evaluate_quality) the build uses. Preview does not include drift (it writes nothing to the workspace, so there is nothing to compare against a previous run). - Schema drift comparison scope:
detect_driftcompares only against the previous successful run of the samedataset_idandsource_key, not "any previous run". It never compares against another dataset/source's silver, which would create fake drift. - Quality History aggregate:
GET /datasets/{dataset_id}/quality/historyreuses #488's dataset-to-run lookup helpers to return per-run pass/warn/fail counts,evaluated_checks,rule_pass_rate(pass_count / evaluated_checks,nullwhenevaluated_checks == 0), andvalidated_rowsfor accessible runs.validated_rowsreuses therow_countssum semantics #488 already defined (per-source Silver row_count), not a sum overQualityCheckResult.evaluated_rowsmultiplied by rule count. - Legacy/partial/failed runs: a legacy run without
quality_resultsin its manifest is represented asevaluated_checks=0, rule_pass_rate=null— unevaluated is not interpreted as "all PASS". Partial and failed runs with structured results are included in history (not excluded by policy). - Ownership: History and detail share the same ownership semantics as
/datasets/{dataset_id}and/datasets/{dataset_id}/runs. Another user's runs never mix in, even with the samedataset_id. - AI interpretation does not affect gates: even when drift cause interpretation (#448, advisory) exists, it does not participate in PASS/WARN/FAIL decisions or dataset quality summaries.
- The
qualityfield inGET /datasetsandGET /datasets/{dataset_id}responses is still alwaysnull. This is intentional: no arbitrary composite quality score is invented; structured results live in the two dedicated endpoints above.
/queryallows only a single SELECT/CTE referencing thedatasetphysical relation at least once. CTEdatasetshadowing, recursive CTEs, external tables/table functions, filesystem/network access, and DML/DDL are rejected.- Queries use a bounded capacity separate from the HTTP worker pool. The query timeout actually terminates the child process; 429/504 errors are distinguished by stable
codevalues.
Gold masks declared PII (kpubdata license.pii_columns + BuildSpec sources[].gold.pii_columns) (#689) but Silver and Bronze preserve the original values (#611). So every service path that reads Silver or Bronze masks or refuses the same columns in the same way as Gold (contract 1.68.0).
| Path | Behaviour |
|---|---|
POST /query stage: silver |
Queries run over a masked copy of Silver. Expressions like upper(col), substr, WHERE col = '…' never see the original values. The copy is read with the same scan_builder_parquet the query engine uses, restoring the Builder dtypes and real column names DuckDB stored in file metadata (#891), so the response's columns/column_meta match an unmasked query (including all-null, Duration, Int128, zone). The response's masked_columns lists the masked columns. stage: gold is already masked at build time and does not change |
POST /preview |
Masks each source's sample, source_sample (original field names — walking back through schema.coalesce and rename), and those columns' diffs; records masked_columns |
GET /builds/{run_id}/stages/silver/{source} |
Masks sample, records masked_columns. Bronze stage detail has no rows |
GET /artifacts/{run_id}/{file_path} |
bronze/{source}/… and silver/{source}/… files for sources with declared columns are all 403 declared_pii_withheld (with column names in columns). Gold files and the manifest are served as-is |
- Masking method is the same as Gold's (#902): text values become
[masked], non-text dtypes become null, null stays null.masked_columnsappears only when columns were actually masked. - Which columns are masked is decided by one function:
stages/gold/pii.py'scolumns_withheld_from_silver. It is thedeclared_pii_columnsdeclaration interpretation Gold uses, minusgold.publish_unmasked— only columns published in plaintext in Gold are plaintext in reads. Columns removed bygold.selectare absent from Gold but present in Silver, so they are masked.pii.allow_columnsis a scan-gate switch and does not unmask declared columns. Declarations are read with the same client factory as the build, and columns recorded in the run manifest'spii_masking.maskedare added — a later catalog removing a declaration does not unmask runs built while it existed. - When the declaration cannot be read, refuse (fail closed): if the kpubdata declaration lookup for a public_api source fails, which columns are personal is unknown (#688 — "unknown is not permission"). The manifest record cannot substitute: it lists only what Gold masked, missing columns excluded by
gold.selector columns from runs that never reached Gold. So that source's/query(stage: silver),/preview, and Bronze/Silver file downloads are refused with 503pii_declaration_unavailable(with the unreadable dataset indataset); Silver stage detail gives metadata but omits rows withsample: [],sample_withheld: pii_declaration_unavailable(the same shape as #892'sredistribution_forbidden). 503 because the failure is a temporary server-side lookup, not a client error — it can be retried. File and url sources and BuildSpec-onlygold.pii_columnsneed no lookup and are unaffected; Gold reads are also unchanged. When the lookup succeeds but the result has changed, the union with the manifest record is used as-is. - Why the original files are refused, not masked in transit: rewriting
raw_records.jsonl,table.parquet, orpreview.jsonon the fly would export files that are not the artifacts under artifact names. When a masked table is needed, take Gold's. Sources without Silver (Bronze-only) are judged against all declared columns available; a stage directory whose source cannot be determined is judged against all sources of the run — unknown is not permission. - The same checkpoints as #892 (redistribution gate): artifact download and stage detail run through
BuilderService.serve_artifact_fileandget_run_stage_detail;/querythroughQueryApiService.query— right where #892 put the terms check, and after it. Whenforbidden, nothing leaves, so there is nothing to mask./previewchecks terms before fetch (from the spec alone) and PII after fetch (needing fetched results and the declarations the client read), so it masks within the sameSpecApiService.preview, right after fetch. The policy modules were not merged (service/redistribution.pyandservice/pii_reads.py): terms refuse reads, declarations mask them — different answers. - Warehouse reads, exports and profiles read the Gold snapshot, so they are not in scope for this change.
BuildSpec.compositionjoins the validated Silver of two sources and produces a separate combined Gold dataset (gold/{composition.name}/) as an addition. Per-source independent Gold is preserved; existing multi-source BuildSpecs withoutcompositionare unaffected.- Sources referenced by
composition.join.left/rightmust declare analias, and within a BuildSpec that usescomposition, the declaredaliasvalues must differ — this rule does not apply to BuildSpecs withoutcomposition. - Structural problems — alias references,
join.type/join.on_duplicate_keyvocabulary — are rejected byvalidate_specimmediately after parsing. Join key column existence and dtype compatibility (exact match only, no automatic casting) require both sources to have passed Silver, so they are the build pipeline's runtime gate. - When both join keys have duplicate values (many-to-many), the output rows can multiply explosively — the default
on_duplicate_key: warnproduces the result and logs a warning;failfails the composition. - The
compositionkey inPOST /buildresponses is{name, status, error}wherestatusis one ofok/failed(the join itself failed)/skipped(a referenced source failed, so the join was never attempted). When only the composition fails and all sources succeed, the top-levelerrorsummary derives from the composition's error. manifest.json'scomposition(CompositionProvenance, additive) records the join conditions and both sides' row counts, distinct key counts, and output row count, so the sources and join conditions are traceable. Legacy manifests lack this field or have null — they are read as "runs without composition".
- Display identity (human-readable labels) and persistent resource ownership identity are separated.
manifest.json'screated_bykeeps the existing (#388) display/legacy label; the newowner_id(additive) is the canonical stable identity used for ownership decisions. owner_idis a domain-separated SHA-256 hash per principal kind. For OIDC,sha256(kind + "\0" + issuer + "\0" + subject)hashes the full issuer/subject (no truncation) — the value does not change when email or display name changes, and never collides with another user whose subject prefix happens to overlap. Rawsub/email and other sensitive claims are never left inowner_idor in logs.- Ownership decisions (
_check_ownership,/query, dataset/quality/stage listings) compareowner_idfirst when both the record and the principal have one. When either side lacks it (e.g., legacy runs predating #505), they fall back to the existingcreated_by/label comparison — existing resources do not become immediately inaccessible. Records with neitherowner_idnorcreated_byare not treated as "accessible to everyone" but are refused (fail-closed). owner_idis internal to ownership decisions. It is stored in the on-diskmanifest.jsonand derivedBuildIndex, but removed from HTTP responses including/buildslistings andGET /builds/{run_id}/manifest. It is therefore not a public property of the OpenAPIBuildManifest, and neither the wire contract nor the API contract version changes.- Subject-prefix collision prevention applies to new resources recorded with
owner_idafter #505. Pre-#505 legacy resources withoutowner_idfall back to the existingcreated_bylabel for compatibility, so already-stored truncated subject prefix collisions cannot be retroactively resolved. - Adopting a specific new IdP or email/password login is outside this section's scope — the stable owner identity computation was settled first, independent of any future IdP decision (#515).
GET /monitoring/summary and GET /monitoring/builds provide the system/aggregate observability Studio's Monitoring view needs. Per-run events (#496) are a separate scope.
- "Unknown" and "zero" are distinguished. Values that were never measured are not faked as
0/healthybut expressed asnullwithavailable/partial/unavailable(reusing the vocabulary quality.py already defines). - Aggregate status (
MonitoringSummaryResponse.status): a deterministic judgment computed from theavailabilityof required subsystems (api/queue/workers/artifact_store). All fouravailable→healthy; anypartial/unavailable→degraded. Latency SLA thresholds (e.g., p95 100ms/500ms/1s) are not used without evidence (ADR/config) —sample_count=0/p95_latency_ms=nullitself, or an actual zero count inqueue/workers, is not a degraded reason (only actualunavailable/partialavailability is). Provider status (#492) is optional and not in this judgment. - Builder API status:
dispatch()execution times are recorded in a bounded ring buffer of the most recent 1000 requests, and p95 is computed nearest-rank (no interpolation).sample_count=0→p95_latency_ms=null. When the collector itself is broken and samples cannot be read (#527):availability=unavailable+sample_count=null+p95_latency_ms=null— "healthy with no samples" and "cannot measure" are not conflated, and this subsystem's failure does not fail the original request or the monitoring response. Healthy/degraded latency threshold judgments are not invented without evidence (ADR/config) — only rawsample_count/p95_latency_msare provided. - Queue/Worker: the async build execution model is implemented with
AsyncBuildExecutor/AsyncBuildJobRegistry(#511/#513);BuilderServicealways creates it and uses it forPOST /builds(async) submissions.queue/workersdirectly reflect this executor's read-only snapshot (AsyncBuildExecutor.stats()), so in a normal runtime they are alwaysavailability: available.waitingis the count of active jobs with status=queued,runningwith status=running, andtotal = waiting + running— terminal (succeeded/failed/cancelled) job history the registry retains is not mixed in.workers.activeequalsrunning(one worker runs one job);workers.capacityis themax_workerspreserved at executor creation, not read fromThreadPoolExecutor's private fields.workers.utilizationisactive / capacity(0.0–1.0).availability: unavailableis kept only as a fallback for (currently unreachable) genuinely async-unsupported configurations, and only then are the remaining fieldsnull— "zero jobs" and "cannot check" are distinguished.BoundedThreadingHTTPServer'sThreadPoolExecutor(#253) is the HTTP connection concurrency limit, unrelated to this. - Artifact store: the mere existence of the
output_rootfolder does not count asavailable— a BuildIndex query must also succeed.last_write_atis obtained only from thefinished_atof the most recent successful (ok) build recorded in BuildIndex; without a success record it isavailablewithnull(distinguishing zero from unknown). Filesystem paths are never exposed. - Build statistics (
/monitoring/builds): timezone is UTC, bucket boundaries are half-open[start, end), and the bucket timestamp isfinished_at(BuildIndex records only completed builds, ADR 0003). Currently onlywindow=24h/bucket=houris supported; other values return 400. Malformed timestamps (parse failure, containing NULL) are excluded from aggregation and reflected inexcluded_count, making the overallavailabilitypartial(a BuildIndex query failure isunavailable; a normal aggregate result of zero isavailable). Bucket count wire fields aretotal/success/failed/cancelled— the internal BuildIndex status valueokis kept as-is (no change); only the external Monitoring API field name maps tosuccess(#527). - Provider status is not included in this version's Monitoring responses because it would trigger real network probes per request (#492).
- Ownership: the system aggregates (
api/queue/workers/artifact_store) contain no individual run's dataset/owner/credential information, so no filtering is needed./monitoring/buildsbucket aggregates and recent runs are filtered with the same policy asprincipal_owns()(#505) whenENFORCE_OWNERSHIP+oidc principal applies, so another user's runs never mix in. Bucket aggregates fetch the whole window first and filter with no loss;recent_runsis a fixed LIMIT-10 query, so the filter is applied in SQL before the LIMIT (BuildIndex.list_recent_owned, #527) — otherwise another principal's recent runs could fill the LIMIT and omit the requester's own.
Public API, file and URL sources share the same canonical source contract and Bronze→Silver→Gold pipeline. BuildSpec.sources[].kind distinguishes public_api (default) / file / url; omitting kind always means public_api, preserving existing behaviour.
- File upload:
POST /uploadsaccepts a raw binary body, not JSON (binary upload instead of multipart).format/encoding/filenameare passed as query parameters; the server validates parseability immediately (corrupt or empty files return 400) for fail-fast. Stored content is isolated by the requesting principal'sowner_id; no API returns the original content —GET /uploads/{upload_id}returns metadata only. An upload referenced byBuildSpec.sources[].upload_idmust have the same owner as the principal requesting the build/preview; otherwise it is treated identically to not-found without distinguishing existence (fail-closed, same ownership pattern as #505).sources[].format/encodingmust exactly match the values validated at upload time. Upload content is stored in SQLite; at 8 MiB or above it spills to a server-named file (#622) — either way, the user-supplied filename/path is never used in the storage path, so there is no path traversal surface, andupload_idis an opaque server-issued identifier (upl_<hex32>) — users cannot reference filenames/paths directly. - URL fetch (P0): a
kind="url"source supports only GET with Auth=None over safe HTTP(S) (Bearer credential integration is P1 after #492). As SSRF defence, schemes other thanhttpsand URLs containing userinfo are refused; hostnames are resolved directly via DNS and connections to non-global addresses (loopback, private, link-local, reserved) are blocked (any non-global address refuses the whole request). The actual TCP connection opens directly to the validated IP, preventing DNS rebinding between validation and connection; each redirect repeats the same validation (max 5). Response size and connect/read timeouts are capped. The BuildSpec contract has no header/POST/PUT/PATCH fields at all, so arbitrary headers or non-GET methods cannot be expressed.urlsources are disabled in multi-user deployments (#685).POST /preview,POST /build,POST /buildsrefuse with403 url_source_forbiddenbefore sending any request — seeBUILD_SPEC.md. - Provenance/manifest non-exposure: a file source's provenance contains only the
upload_id, never the local filesystem path. A url source's provenance/manifest contains only the endpoint with the query string removed, so no incidentally-embedded secret survives. Both kinds reuse the existingSourceProvenanceshape (provider/dataset fields) — file fillsprovider="file", dataset=upload_id; url fillsprovider="url", dataset=<path-safe slug based on host+path>(the human-readable original endpoint stays separately infetch_params.endpoint). - Preview/Build same path:
POST /previewandPOST /buildshare the same source resolver, so file/url sources go through the same schema/sample/quality evaluation flow as public_api sources.
CORS is default-deny for browser clients (Studio, etc.).
- Without
KPUBDATA_BUILDER_ALLOWED_ORIGINSset, cross-origin requests are refused. - Same-origin requests are allowed.
- Preflight allows
GET, POST, PUT, DELETE, OPTIONSandContent-Type, X-API-Key, Authorization, X-Provider-Key, X-Publish-Credentialfor allowed origins.
Authentication supports two paths:
| Method | Header | Use | Environment |
|---|---|---|---|
| API Key | X-API-Key: <secret> |
Service accounts, scheduled workflows | KPUBDATA_BUILDER_API_KEY |
| Bearer (OIDC) | Authorization: Bearer <jwt> |
Human users, Studio | OIDC_ISSUER + OIDC_AUDIENCE |
OIDC is enabled only when OIDC_ISSUER/OIDC_AUDIENCE and allowlists (OIDC_ALLOWED_HD, OIDC_ALLOWED_SUBJECTS, OIDC_ALLOWED_EMAILS) are configured.
Whether a provider key is stored depends on the deployment mode, so "keys are not stored" is true of one mode only:
| Mode | Provider key |
|---|---|
Multi-user (OIDC, or KPUBDATA_BUILDER_ENFORCE_OWNERSHIP) |
Not stored. Sent per request in X-Provider-Key and held in memory for the request or the async job only (#683); PUT /providers/{provider}/credential answers 403 credential_storage_disabled |
| Single-user (API key) | Stored, encrypted with KPUBDATA_BUILDER_CREDENTIAL_MASTER_KEY (AES-GCM, ADR 0012), per principal |
CLI and HTTP service mode share the same domain contract.
| CLI | API operation |
|---|---|
kpubdata-builder validate spec.yaml |
validateSpec |
kpubdata-builder preview spec.yaml |
previewBuild |
kpubdata-builder build spec.yaml |
createBuild |
When CLI and HTTP response semantics diverge, check the OpenAPI and service conformance tests first.
Python code can call the same service logic without HTTP through BuilderService.
from pathlib import Path
from kpubdata_builder.service import BuilderService
service = BuilderService(
output_root=Path("./dist"),
client_factory=lambda: my_kpubdata_client,
)
validate_response = service.validate(spec_yaml_str)
build_response = service.build(spec_yaml_str, run_id="my-run-001")The returned objects' body shapes are also subject to the OpenAPI SSOT and conformance tests.
| Document | Role |
|---|---|
| contract/builder-api.yaml | HTTP wire contract SSOT |
| BUILD_SPEC.md | BuildSpec input contract |
| BUILD_STATE.md | Build state model |
| BOUNDARY.md | Builder–Studio boundary |
| docs/adrs/0002-build-execution-model.md | v0.4 synchronous build model decision |
| docs/adrs/0005-api-contract-single-source.md | OpenAPI SSOT decision |