from kpubdata import Clientclient = Client(
provider_keys={
"datago": "...",
"seoul": "...",
},
timeout=10.0,
cache=True,
cache_ttl_seconds=3600,
)Client(...) supports these transport/cache options:
timeout: float = 30.0max_retries: int = 3cache: bool | ResponseCache = FalseFalse: disable cache (default)True: use default disk cache directoryResponseCache(...): use a caller-supplied cache instance
cache_ttl_seconds: int = 86400env_keys: bool = TrueTrue: a provider missing fromprovider_keysis looked up in the environment (KPUBDATA_<PROVIDER>_API_KEY, then<PROVIDER>_API_KEY)False: explicit-keys-only — onlyprovider_keysis used, for every call (includingprobe), and no credential is read from the environment (includingKPUBDATA_SGIS_CONSUMER_SECRET). Use it when the client holds someone else's key, so a missing entry is never filled with the operator's key (#694)
client = Client.from_env()Client.from_env() additionally honors:
KPUBDATA_CACHE=1KPUBDATA_CACHE_DIR=/custom/cache/pathKPUBDATA_CACHE_TTL=3600
client.datasets.list()
client.datasets.list(provider="datago")
client.datasets.search("예보")
client.datasets.search("forecast", provider="datago")
client.datasets.search("weathr", threshold=0.5) # fuzzy match
client.dataset("datago.apt_trade")search(text, *, provider=None, threshold=0.5) scores each dataset against
text using substring, token-overlap, and fuzzy matching across name,
description, tags, id, dataset_key, and provider. Results are
returned in descending relevance order.
- Exact substring match in any field → score 1.0 (always included).
- Fuzzy match via
difflib.SequenceMatcher→ included if score ≥ threshold. - threshold default is
0.5. Raise it (e.g.0.8) for stricter results, lower it (e.g.0.2) for broader recall.
Each DatasetRef returned by discovery exposes:
id,provider,dataset_key,name,representation,operationsdescription: human-readable summary (may beNone)tags: categorization keywords as a tuple (e.g.("weather", "forecast"))source_url: link to original API documentation (may beNone)license: the spec'slicensesection as parsed (LicenseSpec), orNonewhen the dataset declares none.Nonemeans unknown, never "no restrictions" (#609)
dataset = client.dataset("datago.apt_trade")
result = dataset.list(lawd_code="11680", deal_ym="202503")list() returns exactly one page of results. When another page is available,
the returned RecordBatch.next_page is set.
dataset = client.dataset("datago.apt_trade")
for batch in dataset.list_all(lawd_code="11680", deal_ym="202503"):
for item in batch.items:
print(item)list_all() yields one RecordBatch per page and follows next_page
automatically until pagination is exhausted. page and page_size are
pagination controls, not filters: they are never sent as raw provider
parameters. When more than max_pages pages (default 1000) would be needed,
the pages fetched so far are yielded and then InvalidRequestError is raised.
For spec-backed datasets, list_all() decides column casting across all pages
at once, so it buffers: every page is fetched before the first batch is
yielded, and memory grows with the total result (bounded by max_pages).
page_size is capped to the spec's pagination.max_size, as in list(). Each
batch still carries its own page's raw, meta["provenance"], next_page and
validation; meta["validation_total"] holds the report for the whole result,
which is the one that decided casting.
Dataset metadata may expose provider-specific pagination styles through
DatasetRef.query_support.pagination, including offset, cursor, and
index-window based index pagination.
list() validates canonical query parameters and raises InvalidRequestError for invalid values before provider invocation:
| Parameter | Accepted values |
|---|---|
page |
None or positive integer (>= 1) |
page_size |
None or positive integer (>= 1) |
cursor |
None or non-empty string |
start_date, end_date |
None or non-empty string (provider-specific format) |
fields, sort |
None or list[str] |
Other kwargs are passed as provider-specific filters to the adapter. For example, dataset.list(region="11", custom_param="value") passes {"region": "11", "custom_param": "value"} as provider-specific filters.
Provider-specific rules (enforced by adapters):
- Required provider parameters (e.g., some datasets require
start_date/end_date) - Provider-specific date formats (e.g.,
YYYYMMvsYYYYMMDD) - Provider-specific maximum page size
schema = dataset.schema()raw = dataset.call_raw(operation="list", lawd_code="11680", deal_ym="202503")from kpubdata import PROBE_STATUSES, Client, ProbeResult
client = Client(provider_keys={"datago": user_key}, env_keys=False)
one: ProbeResult | None = client.probe("datago.apt_trade")
every: list[ProbeResult] = client.probe_all(provider="datago")Client.probe(dataset_id) -> ProbeResult | None— at most one call per dataset, classified intoPROBE_STATUSES(theProbeStatusvocabulary of ADR 0005).Nonewhen no spec exists for the id: only spec-defined datasets can be probed.Client.probe_all(*, provider=None) -> list[ProbeResult]— every spec-defined dataset, or those of one provider.- A failed call is the verdict, not an exception.
- Probing uses its own fast-fail transport (15 s timeout, zero retries, no
cache), not the client's
timeout,max_retriesorcache. - When the dataset needs a key the client does not have, no call is made and
the status is
auth_unknown. ProbeResultis a frozen dataclass:dataset_id,service_id(the data.go.kr service the activation request is made for),status,probed_at(ISO 8601, UTC),detail, and the evidence fields the drift step reads (docs/LIVE_PROBE.md):http_status(200 on success, the raised error's status otherwise),result_code(the provider code read off the raised error;Noneon success — a successful query does not parse the envelope),latency_ms(the one call),schema_hash(a fingerprint of the returned field names;Nonewhen the call failed or the page was empty) andclassification(theDriftClassificationthis outcome feeds the dataset status machine;Nonewhen the outcome says nothing about the upstream API). Each evidence field isNonewhere it could not be observed — never guessed. Configured keys never appear indetail, in log records or in exception messages.
Optional convenience aliases may be added for common datasets, but only if they do not obscure the canonical dataset id.
Example:
client.apartment_trades.list(lawd_code="11680", deal_ym="202503")Rules:
- convenience aliases are secondary
- docs should always show the canonical
client.dataset(...)path
Returns RecordBatch.
RecordBatch.validation is a ValidationReport | None (None when the adapter did
not run field-level validation). A report holds issues: tuple[FieldIssue, ...] and
exposes:
| Member | Description |
|---|---|
ok |
True when there are no issues |
issues_of(kind) |
Issues of one IssueKind ("uncastable", "missing", "undeclared"); raises ValueError for an unknown kind |
to_dict() |
JSON-serialisable {"ok": bool, "issues": [...]} |
Counts are taken on the value the cast sees, so numeric null markers ("", "-")
count as nulls. A page with no records produces an empty (ok) report.
Returns a generator of RecordBatch values, one per page.
Returns SchemaDescriptor | None.
SchemaDescriptor.fields contains FieldDescriptor objects. Each FieldDescriptor may carry a constraints: FieldConstraints | None with structured metadata:
| Attribute | Type | Description |
|---|---|---|
max_length |
int | None |
Maximum character length |
min_value |
float | None |
Minimum numeric value |
max_value |
float | None |
Maximum numeric value |
pattern |
str | None |
Regex pattern |
allowed_values |
tuple[str, ...] | None |
Permitted values |
format |
str | None |
Semantic format hint (e.g. "YYYYMM") |
Returns provider-native payload / response object.
KPubData promises stability for:
ClientClient.from_env()Client.probe()/Client.probe_all(),ProbeResult,PROBE_STATUSES- dataset discovery methods
Dataset.list/list_all/schema/call_raw- canonical model classes
- canonical error types
KPubData does not promise stability for:
- internal adapter helper functions
- internal transport implementation details
- provider-specific metadata layout beyond documented keys
Avoid turning the public API into:
- copied endpoint names from providers
- a giant generic
query()with dozens of weakly-defined arguments - a mandatory dataframe API
- a fake SQL-like language
Preferred:
client.dataset("seoul.subway_realtime_arrival").list(station_name="강남")Acceptable advanced usage:
client.dataset("seoul.subway_realtime_arrival").call_raw(operation="list", stationNm="강남")Discouraged as the main public entry point:
client.call_provider_endpoint("seoul", "SearchSTNTimeTableByIDService", ...)설치 후에는 콘솔 스크립트 kpubdata를 사용할 수 있습니다.
- 데이터셋 목록:
kpubdata datasets list - 메타데이터 확인:
kpubdata datasets show <dataset_id> - 정규화 조회:
kpubdata fetch <dataset_id> ... - Raw 호출:
kpubdata raw <dataset_id> <operation> ...
상세 사용법, 출력 형식, 종료 코드, 환경 변수 연동은 docs/cli.md를 참고하세요.
| 문서 | 설명 |
|---|---|
| ARCHITECTURE.md | 시스템 아키텍처 설계 |
| CANONICAL_MODEL.md | 표준 데이터 모델 정의 |
| PROVIDER_ADAPTER_CONTRACT.md | 어댑터 구현 규약 |
| VALIDATION.md | 아키텍처 타당성 검증 |
| 저장소 | 문서 | 설명 |
|---|---|---|
| kpubdata-builder | API_CONTRACT.md | Builder API 규약 |
| kpubdata-studio | API_CONTRACT.md | Studio API 규약 |