Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -30,6 +30,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

- **`list_all()` on a spec dataset holds one page in memory, not the whole result** (#789). Global column casting (#481) needs every page before any is cast, and the fetched pages waited in a list, so memory grew with the row count: 40 pages of 500 rows peaked at 16.5 MB against 1.7 MB for 4 pages. Each page now goes to a temporary file and only its evidence for the casting decision is kept; the pages are read back, cast and yielded one at a time once the last is in — 1.3 MB for 4 pages and for 40. The casting rule, the batches and their reports are unchanged, and `SpecExecutor._finalize_casting` now uses the same decision code. What a failing page does is now documented, not changed: the exception propagates, no batch is yielded and the earlier pages are discarded.
- The package metadata names its repository (#791): `[project.urls]` gives the homepage, documentation, repository, issue tracker and changelog under `kpubdata-lab/kpubdata`, so the PyPI page links back to the source. There was no `[project.urls]` at all.
- `compatibility.json` says what it is (#788). Its description claimed "Builder and Studio CI read this file to check their kpubdata dependency range"; no workflow or script in kpubdata, kpubdata-builder or kpubdata-studio reads it. The sentence is replaced: the file is a record for people and release notes. `tests/unit/test_compatibility_json.py` checks what can be checked without another repository — the file's shape, that exactly one range is `supported` and the package's own version is inside it, and that the range is the pin `docs/compatibility.md` states.
- **HTTP 401 raises `AuthError` and HTTP 503 raises `ServiceUnavailableError`** (#786). Both used to leave the transport as a plain `TransportError`, told apart only by `status_code`; 429 was already `RateLimitError`. **Behaviour change:** `AuthError` is not a `TransportError`, so `except TransportError` no longer catches a 401 — catch `AuthError` (or `PublicDataError`). A 503 is still a `TransportError` (its subclass), still retried, and typed once the retries run out. The `kpubdata` CLI exits `3` instead of `4` on a 401. `Client.probe` reports a 401 as `auth_unknown`, as before. Every error now has a stable `code` (`auth_error`, `service_unavailable`, `rate_limited`, …) and `to_dict()`, a JSON-serialisable dict of `code`, `type`, `message`, `provider`, `dataset_id`, `operation`, `status_code`, `provider_code` and `retryable`, so a consumer maps an error without comparing message text; the table is in `API_SPEC.md` §7.
- **Breaking: `Client` refuses an unknown keyword argument** (#781). `Client(**extra)` accepted any keyword and dropped it without an error or a warning — `Client(timeoutt=5)` built a client with the default timeout, and `env_keys=False` passed to 0.8.0, which does not have it, left the environment keys in use (#780). It is now a `TypeError` naming the argument. `Client.from_env(extra=...)`, whose dict was forwarded into the same channel, is removed with it; `from_env` takes `provider_keys`, `timeout`, `max_retries`, `cache` and `cache_ttl_seconds`. Nothing in the library read the dropped values, so a caller passing only documented options is unaffected.
- The repository moved from `yeongseon/kpubdata` to `kpubdata-lab/kpubdata`. Links, the documentation site (`https://kpubdata-lab.github.io/kpubdata/`) and the shared GitHub Actions references now use the new owner.
Expand Down
2 changes: 1 addition & 1 deletion compatibility.json
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
{
"version": 1,
"owner": "kpubdata",
"description": "Cross-repo compatibility matrix — single source of truth (#466). Builder and Studio CI read this file to check their kpubdata dependency range.",
"description": "Cross-repo compatibility matrix (#466). A record kept for people and release notes: no CI in kpubdata, kpubdata-builder or kpubdata-studio reads it to gate a dependency range (#788). tests/unit/test_compatibility_json.py checks its shape and that the released-range row matches Builder's pin as docs/compatibility.md states it.",
"update_policy": "Update on every kpubdata release. Builder/Studio link here, never copy.",
"matrix": [
{
Expand Down
93 changes: 93 additions & 0 deletions tests/unit/test_compatibility_json.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,93 @@
"""``compatibility.json`` is a record, and the record is kept consistent (#788).

The file said Builder and Studio CI read it to check their kpubdata range; nothing in
any of the three repositories does. It is kept as documentation, so what can be checked
is checked here: its shape, and that its supported range agrees with the version this
package declares and with the pin ``docs/compatibility.md`` states.
"""

from __future__ import annotations

import json
import re
from pathlib import Path
from typing import Any

import pytest

_ROOT = Path(__file__).resolve().parents[2]
_STATUSES = {"supported", "ci-tested", "released"}
_ROW_KEYS = {"kpubdata", "builder", "studio", "status", "note"}


@pytest.fixture(scope="module")
def document() -> dict[str, Any]:
loaded: dict[str, Any] = json.loads((_ROOT / "compatibility.json").read_text("utf-8"))
return loaded


def _version() -> tuple[int, ...]:
text = (_ROOT / "pyproject.toml").read_text("utf-8")
match = re.search(r'^version = "(\d+)\.(\d+)\.(\d+)"', text, re.MULTILINE)
assert match
return tuple(int(part) for part in match.groups())


def in_range(version: tuple[int, ...], spec: str) -> bool:
"""Whether ``version`` satisfies a ``>=a.b.c,<d.e`` range as the matrix writes it."""
for clause in spec.split(","):
match = re.fullmatch(r"\s*(>=|<)\s*([\d.]+)\s*", clause)
assert match, f"unreadable range clause {clause!r}"
bound = tuple(int(part) for part in match.group(2).split("."))
padded = bound + (0,) * (len(version) - len(bound))
if match.group(1) == ">=" and version < padded:
return False
if match.group(1) == "<" and version >= padded:
return False
return True


def test_the_shape_is_what_a_reader_expects(document: dict[str, Any]) -> None:
assert document["version"] == 1
assert document["owner"] == "kpubdata"
assert document["matrix"]
for row in document["matrix"]:
assert set(row) == _ROW_KEYS, row
assert row["status"] in _STATUSES, row
assert all(isinstance(row[key], str) and row[key] for key in _ROW_KEYS), row


def test_it_no_longer_claims_a_ci_reads_it(document: dict[str, Any]) -> None:
"""Only the false claim is pinned; the description may be reworded freely."""
assert "CI read this file" not in document["description"]


def test_exactly_one_range_is_supported_and_this_version_is_in_it(
document: dict[str, Any],
) -> None:
supported = [row for row in document["matrix"] if row["status"] == "supported"]

assert len(supported) == 1
assert in_range(_version(), supported[0]["kpubdata"])


def test_the_supported_range_is_the_pin_the_compatibility_document_states(
document: dict[str, Any],
) -> None:
(supported,) = [row for row in document["matrix"] if row["status"] == "supported"]
text = (_ROOT / "docs" / "compatibility.md").read_text("utf-8")

assert f"`{supported['kpubdata']}`" in text


@pytest.mark.parametrize(
("version", "spec", "expected"),
[
((0, 8, 0), ">=0.8.0,<0.9", True),
((0, 8, 7), ">=0.8.0,<0.9", True),
((0, 9, 0), ">=0.8.0,<0.9", False),
((0, 7, 9), ">=0.8.0,<0.9", False),
],
)
def test_the_range_reader(version: tuple[int, ...], spec: str, expected: bool) -> None:
assert in_range(version, spec) is expected
Loading