diff --git a/docs/standards/coverage.md b/docs/standards/coverage.md index 06356fe..3692d91 100644 --- a/docs/standards/coverage.md +++ b/docs/standards/coverage.md @@ -3,7 +3,7 @@ Generated from [requirements.json](requirements.json) by `python scripts/check_requirements.py --render`. The default command checks this table without modifying it. -49 supported requirements; 5 explicit profile exclusions; 874 distinct collected tests; 1003 requirement-to-test links. +49 supported requirements; 5 explicit profile exclusions; 905 distinct collected tests; 1034 requirement-to-test links. Counts describe collected evidence, not executed or passing tests. One test can support several rows. This is not a percentage of all clauses in the copied specifications or a complete Cartesian test matrix. @@ -42,8 +42,8 @@ harness skips are reported by pytest when executing the suite, not hidden by col | [FORMAT-FRAMING](#format-framing) | policy / supported | Nonempty mapped RDF output has exactly one trailing newline on normal and carrying degraded paths. | 41 | | [FORMAT-EMPTY](#format-empty) | policy / supported | Empty Turtle and N-Triples stay empty; empty JSON-LD is a newline-terminated empty array. | 1 | | [GRAPH-DEFAULT](#graph-default) | normative / supported | Single-graph N-Quads has no graph label; degraded TriG carries explicit default-graph list triples. | 27 | -| [RDF-ISOMORPHISM](#rdf-isomorphism) | normative / supported | Seeded Turtle graph shapes preserve RDF graph identity after serialization and parsing. | 40 | -| [TURTLE-STABILITY](#turtle-stability) | policy / supported | Seeded Turtle output is idempotent and independent of insertion order and source blank-node names. | 120 | +| [RDF-ISOMORPHISM](#rdf-isomorphism) | normative / supported | Seeded Turtle graph shapes and a complete generated-schema document preserve RDF graph identity after serialization and parsing. | 48 | +| [TURTLE-STABILITY](#turtle-stability) | policy / supported | Seeded Turtle output and every rendering of a complete generated-schema document are idempotent and independent of insertion order and source blank-node names. | 141 | | [RDF-LEXICAL](#rdf-lexical) | normative / supported | Checked integer, boolean, dateTime, decimal, signed-zero and non-finite literal spellings retain lexical term identity. | 18 | | [RDF-LANGUAGE](#rdf-language) | normative / supported | Plain strings and xsd:string render equivalently; normal-path language tags use the permitted lowercase spelling. | 4 | | [RDF-IRI-BASE](#rdf-iri-base) | normative / supported | Fragment, path, query and equal-base namespaces preserve direct and datatype IRIs in the checked verified renderings. | 43 | @@ -72,7 +72,7 @@ harness skips are reported by pytest when executing the suite, not hidden by col | [JSON-UNKNOWN](#json-unknown) | policy / supported | Unknown remote, scoped and unsupported local contexts conservatively preserve descendant arrays and never fetch remote contexts. | 5 | | [JSON-OVERRIDES](#json-overrides) | policy / supported | Custom preserved-key sets replace convenience defaults without disabling keyword-based list protection. | 4 | | [WL-IDENTITY](#wl-identity) | policy / supported | WL relabelling returns a fresh isomorphic quad list, preserves blank-node cardinality and leaves input quads untouched. | 10 | -| [WL-COMPONENTS](#wl-components) | policy / supported | Unrelated additions, edits and removals do not advance a disconnected component's default refinement. | 9 | +| [WL-COMPONENTS](#wl-components) | policy / supported | Unrelated additions, edits and removals do not advance a disconnected component's default refinement, including an added class in a complete generated-schema document. | 11 | | [WL-TIES](#wl-ties) | policy / supported | Structurally tied nodes remain injective and collision suffixes follow numeric canonical numbering. | 4 | | [WL-GRAPHS](#wl-graphs) | policy / supported | Named graphs influence labels; blank graph-name identifiers remain consistent when also used as object terms. | 4 | | [WL-ITERATIONS](#wl-iterations) | policy / supported | Default refinement reaches component fixpoints, explicit rounds remain synchronous, zero is valid and negative counts are refused. | 7 | @@ -292,7 +292,7 @@ References: [NQUADS11 §2.1 Simple Statements](references/n-quads.html#simple-tr ### RDF-ISOMORPHISM -Seeded Turtle graph shapes preserve RDF graph identity after serialization and parsing. +Seeded Turtle graph shapes and a complete generated-schema document preserve RDF graph identity after serialization and parsing. Features: `deterministic_turtle`. @@ -307,11 +307,11 @@ References: [RDF11-CONCEPTS §3.6 Graph Comparison](references/rdf11-concepts.ht | boundary | Not applicable: No additional boundary beyond the explicitly listed examples is claimed by this row. | | property | `tests/properties/test_canonicalization_properties.py::test_p1_canonical_output_is_lossless` (40) | | subprocess | Not applicable: This row has no process-state-specific obligation; process determinism is mapped separately. | -| end-to-end | Not applicable: This row isolates a contract dimension; executable composed workflows are mapped in API-EXAMPLES. | +| end-to-end | `tests/integration/test_naturalistic_schema.py::test_every_rendering_preserves_the_document` (7)
`tests/integration/test_naturalistic_schema.py::test_the_document_survives_a_named_node_rename` (1) | ### TURTLE-STABILITY -Seeded Turtle output is idempotent and independent of insertion order and source blank-node names. +Seeded Turtle output and every rendering of a complete generated-schema document are idempotent and independent of insertion order and source blank-node names. Features: `deterministic_turtle`. @@ -326,7 +326,7 @@ Project contract; no normative standard algorithm is claimed. | boundary | Not applicable: No additional boundary beyond the explicitly listed examples is claimed by this row. | | property | `tests/properties/test_canonicalization_properties.py::test_p2_canonicalization_is_idempotent` (40)
`tests/properties/test_canonicalization_properties.py::test_p3_output_does_not_depend_on_blank_node_labels` (40)
`tests/properties/test_canonicalization_properties.py::test_p4_output_does_not_depend_on_insertion_order` (40) | | subprocess | Not applicable: This row has no process-state-specific obligation; process determinism is mapped separately. | -| end-to-end | Not applicable: This row isolates a contract dimension; executable composed workflows are mapped in API-EXAMPLES. | +| end-to-end | `tests/integration/test_naturalistic_schema.py::test_every_rendering_is_idempotent` (7)
`tests/integration/test_naturalistic_schema.py::test_every_rendering_ignores_incoming_blank_node_names` (7)
`tests/integration/test_naturalistic_schema.py::test_every_rendering_ignores_insertion_order` (7) | ### RDF-LEXICAL @@ -873,7 +873,7 @@ References: [RDF11-CONCEPTS §3.6 Graph Comparison](references/rdf11-concepts.ht ### WL-COMPONENTS -Unrelated additions, edits and removals do not advance a disconnected component's default refinement. +Unrelated additions, edits and removals do not advance a disconnected component's default refinement, including an added class in a complete generated-schema document. Features: `wl_blank_node_labels`, `wl_relabel_quads`, `deterministic_turtle`. @@ -890,7 +890,7 @@ Project contract; no normative standard algorithm is claimed. | boundary | `tests/wl/test_wl_components.py::test_shared_blank_subject_object_roles_form_one_component` (1)
`tests/wl/test_wl_components.py::test_graph_membership_does_not_join_signature_independent_components` (2) | | property | `tests/wl/test_wl_components.py::test_component_labels_ignore_input_order_and_blank_node_names` (1) | | subprocess | Not applicable: This row has no process-state-specific obligation; process determinism is mapped separately. | -| end-to-end | Not applicable: This row isolates a contract dimension; executable composed workflows are mapped in API-EXAMPLES. | +| end-to-end | `tests/integration/test_naturalistic_schema.py::test_adding_a_class_rewrites_nothing_rdfc_alone_would_rewrite` (1)
`tests/integration/test_naturalistic_schema.py::test_adding_a_class_keeps_every_untouched_statement_verbatim` (1) | ### WL-TIES diff --git a/docs/standards/requirements.json b/docs/standards/requirements.json index 73fa5ce..a16559e 100644 --- a/docs/standards/requirements.json +++ b/docs/standards/requirements.json @@ -595,7 +595,7 @@ }, { "id": "RDF-ISOMORPHISM", - "description": "Seeded Turtle graph shapes preserve RDF graph identity after serialization and parsing.", + "description": "Seeded Turtle graph shapes and a complete generated-schema document preserve RDF graph identity after serialization and parsing.", "features": [ "deterministic_turtle" ], @@ -623,18 +623,21 @@ ], "positive": [ "tests/properties/test_canonicalization_properties.py::test_p1_canonical_output_is_lossless" + ], + "end-to-end": [ + "tests/integration/test_naturalistic_schema.py::test_every_rendering_preserves_the_document", + "tests/integration/test_naturalistic_schema.py::test_the_document_survives_a_named_node_rename" ] }, "not_applicable": { "negative": "This is a transformation of admitted values, not a syntax parser or a new input-rejection API.", "boundary": "No additional boundary beyond the explicitly listed examples is claimed by this row.", - "subprocess": "This row has no process-state-specific obligation; process determinism is mapped separately.", - "end-to-end": "This row isolates a contract dimension; executable composed workflows are mapped in API-EXAMPLES." + "subprocess": "This row has no process-state-specific obligation; process determinism is mapped separately." } }, { "id": "TURTLE-STABILITY", - "description": "Seeded Turtle output is idempotent and independent of insertion order and source blank-node names.", + "description": "Seeded Turtle output and every rendering of a complete generated-schema document are idempotent and independent of insertion order and source blank-node names.", "features": [ "deterministic_turtle" ], @@ -657,13 +660,17 @@ ], "positive": [ "tests/properties/test_canonicalization_properties.py::test_p2_canonicalization_is_idempotent" + ], + "end-to-end": [ + "tests/integration/test_naturalistic_schema.py::test_every_rendering_is_idempotent", + "tests/integration/test_naturalistic_schema.py::test_every_rendering_ignores_incoming_blank_node_names", + "tests/integration/test_naturalistic_schema.py::test_every_rendering_ignores_insertion_order" ] }, "not_applicable": { "negative": "This is a transformation of admitted values, not a syntax parser or a new input-rejection API.", "boundary": "No additional boundary beyond the explicitly listed examples is claimed by this row.", - "subprocess": "This row has no process-state-specific obligation; process determinism is mapped separately.", - "end-to-end": "This row isolates a contract dimension; executable composed workflows are mapped in API-EXAMPLES." + "subprocess": "This row has no process-state-specific obligation; process determinism is mapped separately." } }, { @@ -2023,7 +2030,7 @@ }, { "id": "WL-COMPONENTS", - "description": "Unrelated additions, edits and removals do not advance a disconnected component's default refinement.", + "description": "Unrelated additions, edits and removals do not advance a disconnected component's default refinement, including an added class in a complete generated-schema document.", "features": [ "wl_blank_node_labels", "wl_relabel_quads", @@ -2062,12 +2069,15 @@ "boundary": [ "tests/wl/test_wl_components.py::test_shared_blank_subject_object_roles_form_one_component", "tests/wl/test_wl_components.py::test_graph_membership_does_not_join_signature_independent_components" + ], + "end-to-end": [ + "tests/integration/test_naturalistic_schema.py::test_adding_a_class_rewrites_nothing_rdfc_alone_would_rewrite", + "tests/integration/test_naturalistic_schema.py::test_adding_a_class_keeps_every_untouched_statement_verbatim" ] }, "not_applicable": { "negative": "This is a transformation of admitted values, not a syntax parser or a new input-rejection API.", - "subprocess": "This row has no process-state-specific obligation; process determinism is mapped separately.", - "end-to-end": "This row isolates a contract dimension; executable composed workflows are mapped in API-EXAMPLES." + "subprocess": "This row has no process-state-specific obligation; process determinism is mapped separately." } }, { diff --git a/tests/integration/test_naturalistic_schema.py b/tests/integration/test_naturalistic_schema.py new file mode 100644 index 0000000..4417e03 --- /dev/null +++ b/tests/integration/test_naturalistic_schema.py @@ -0,0 +1,288 @@ +"""End-to-end checks on a complete, naturalistic generated-schema document. + +The seeded graphs in ``properties/`` cover the shapes that break canonicalization +one at a time. This module covers the other risk: a whole document of the kind a +schema generator actually emits, where those shapes occur together and interact. +It is the regression target for consumers that commit generated RDF, so the +assertions are about the document as a whole rather than an isolated dimension. + +The fixture deliberately combines, in one graph: + +* OWL restrictions hanging off ``rdfs:subClassOf``, several per class +* ``owl:unionOf`` lists, two of which **share a tail cell** +* a list whose member is itself a list +* one blank node referenced from two subjects, and a blank-node cycle +* SHACL property shapes and an ``sh:ignoredProperties`` list +* language-tagged, plain, empty, escaped and typed literals, and ``rdf:nil`` +""" + +from __future__ import annotations + +import difflib +import re + +import pytest +from rdflib import BNode, Graph, Literal, Namespace, URIRef +from rdflib.compare import isomorphic +from rdflib.namespace import OWL, RDF, SH + +from diffable_rdf import canonicalize_rdf_graph, deterministic_turtle + +EX = Namespace("https://example.org/person-schema/") + +#: A generated schema document, in the shape a LinkML-style OWL/SHACL generator emits. +SCHEMA_TURTLE = r""" +@prefix ex: . +@prefix owl: . +@prefix rdf: . +@prefix rdfs: . +@prefix sh: . +@prefix xsd: . + +ex: a owl:Ontology ; + rdfs:label "Person schema"@en ; + rdfs:label "Personenschema"@de ; + owl:versionInfo "1.4.2" ; + ex:sourceFile "https://example.org/person.yaml"^^xsd:anyURI ; + ex:generated "2026-09-18"^^xsd:date ; + ex:experimental false . + +ex:Person a owl:Class ; + rdfs:label "Person"@en ; + rdfs:comment "A person.\nHas a \"name\", a C:\\path and a tab\there." ; + ex:note "" ; + ex:motto "caf\u00e9 na\u00efve \u2615" ; + rdfs:subClassOf [ a owl:Restriction ; owl:onProperty ex:name ; owl:someValuesFrom xsd:string ] , + [ a owl:Restriction ; owl:onProperty ex:knows ; owl:allValuesFrom ex:Person ] , + [ a owl:Restriction ; owl:onProperty ex:age ; + owl:maxCardinality "1"^^xsd:nonNegativeInteger ] ; + ex:contactVia _:address ; + ex:tagged ( "primary" ( "nested" "inner" ) ) ; + ex:noAliases rdf:nil . + +ex:Organization a owl:Class ; + rdfs:label "Organization"@en ; + rdfs:subClassOf [ a owl:Restriction ; owl:onProperty ex:name ; owl:someValuesFrom xsd:string ] ; + ex:contactVia _:address . + +ex:Employee a owl:Class ; + rdfs:subClassOf ex:Person ; + owl:equivalentClass [ owl:unionOf _:employeeUnion ] . + +ex:Contractor a owl:Class ; + owl:equivalentClass [ owl:unionOf _:contractorUnion ] . + +# Two unions sharing one tail cell: an inline "( ... )" would consume the shared +# cell for whichever list is written first and detach it from the other. +_:employeeUnion rdf:first ex:Person ; rdf:rest _:sharedTail . +_:contractorUnion rdf:first ex:Organization ; rdf:rest _:sharedTail . +_:sharedTail rdf:first ex:Customer ; rdf:rest rdf:nil . + +ex:Customer a owl:Class . + +_:address a ex:Address ; + ex:city "Berlin"@de ; + ex:postalCode "10115" ; + ex:latitude "52.53"^^xsd:double ; + ex:floor 3 . + +# A cycle between blank nodes: label assignment must still terminate. +_:provenanceA ex:derivedFrom _:provenanceB ; ex:step 1 . +_:provenanceB ex:derivedFrom _:provenanceA ; ex:step 2 . +ex: ex:provenance _:provenanceA . + +ex:PersonShape a sh:NodeShape ; + sh:targetClass ex:Person ; + sh:closed true ; + sh:ignoredProperties ( rdf:type owl:sameAs ) ; + sh:property [ sh:path ex:name ; sh:datatype xsd:string ; sh:minCount 1 ; sh:maxCount 1 ; sh:order 0 ] , + [ sh:path ex:age ; sh:datatype xsd:integer ; sh:minInclusive 0 ; sh:maxInclusive 150 ] , + [ sh:path ex:height ; sh:datatype xsd:double ; sh:minExclusive 0.0 ] , + [ sh:path ex:knows ; sh:class ex:Person ; sh:nodeKind sh:IRI ] . +""" + +#: An independent class added later, the way a schema grows in version control. +ADDED_CLASS_TURTLE = r""" +@prefix ex: . +@prefix owl: . +@prefix rdfs: . +@prefix sh: . +@prefix xsd: . + +ex:Robot a owl:Class ; + rdfs:label "Robot"@en ; + rdfs:subClassOf [ a owl:Restriction ; owl:onProperty ex:serial ; owl:someValuesFrom xsd:string ] . + +ex:RobotShape a sh:NodeShape ; + sh:targetClass ex:Robot ; + sh:property [ sh:path ex:serial ; sh:datatype xsd:string ; sh:minCount 1 ] . +""" + + +def _schema_graph() -> Graph: + """Parse the generated-schema fixture.""" + return Graph().parse(data=SCHEMA_TURTLE, format="turtle") + + +def _extended_graph() -> Graph: + """The same schema after an unrelated class is added.""" + graph = _schema_graph() + graph.parse(data=ADDED_CLASS_TURTLE, format="turtle") + return graph + + +def _relabelled(graph: Graph) -> Graph: + """The same graph with every blank node given a different identifier.""" + mapping: dict[BNode, BNode] = {} + + def rename(term): + if isinstance(term, BNode): + return mapping.setdefault(term, BNode()) + return term + + renamed = Graph() + for prefix, namespace in graph.namespaces(): + renamed.bind(prefix, namespace) + for subject, predicate, obj in graph: + renamed.add((rename(subject), predicate, rename(obj))) + return renamed + + +def _reordered(graph: Graph) -> Graph: + """The same graph with its triples inserted in the opposite order.""" + reordered = Graph() + for prefix, namespace in graph.namespaces(): + reordered.bind(prefix, namespace) + for triple in sorted(graph, key=str, reverse=True): + reordered.add(triple) + return reordered + + +def _renderings() -> list: + """Every public rendering this document must survive, with its parse format.""" + return [ + pytest.param(deterministic_turtle, "turtle", id="deterministic_turtle"), + pytest.param(lambda g: canonicalize_rdf_graph(g, "turtle"), "turtle", id="canonical_turtle"), + pytest.param(lambda g: canonicalize_rdf_graph(g, "turtle", diff_stable=True), "turtle", id="stable_turtle"), + pytest.param(lambda g: canonicalize_rdf_graph(g, "nt"), "nt", id="canonical_nt"), + pytest.param(lambda g: canonicalize_rdf_graph(g, "nt", diff_stable=True), "nt", id="stable_nt"), + pytest.param(lambda g: canonicalize_rdf_graph(g, "xml"), "xml", id="canonical_xml"), + pytest.param(lambda g: canonicalize_rdf_graph(g, "json-ld"), "json-ld", id="canonical_jsonld"), + ] + + +def test_fixture_contains_the_shapes_the_other_tests_rely_on() -> None: + """Guard the fixture: every hard shape must really be present. + + Without this, a fixture that silently lost its shared list cell or its + language tags would leave the round-trip assertions passing vacuously. + """ + graph = _schema_graph() + + tails = [cell for cell in graph.subjects(RDF.rest, None) if len(list(graph.subjects(RDF.rest, cell))) > 1] + assert tails, "no shared list tail cell" + + shared = [node for node in graph.objects(None, EX.contactVia) if len(list(graph.subjects(EX.contactVia, node))) > 1] + assert shared and isinstance(shared[0], BNode), "no blank node shared between two subjects" + + cycle = set(graph.subjects(EX.derivedFrom, None)) & set(graph.objects(None, EX.derivedFrom)) + assert len(cycle) == 2, "no blank-node cycle" + + restrictions = set(graph.subjects(RDF.type, OWL.Restriction)) + assert len(restrictions) >= 4, "too few OWL restrictions" + assert all(isinstance(node, BNode) for node in restrictions) + + literals = [obj for obj in graph.objects() if isinstance(obj, Literal)] + assert {literal.language for literal in literals} >= {"en", "de"}, "no language-tagged literals" + assert any(literal.datatype is not None for literal in literals), "no typed literals" + assert any(str(literal) == "" for literal in literals), "no empty literal" + assert any("\n" in str(literal) and '"' in str(literal) for literal in literals), "no escaped literal" + assert (EX.Person, EX.noAliases, RDF.nil) in graph, "no rdf:nil object" + assert set(graph.subjects(RDF.type, SH.NodeShape)), "no SHACL shapes" + + +@pytest.mark.parametrize(("render", "parse_format"), _renderings()) +def test_every_rendering_preserves_the_document(render, parse_format: str) -> None: + """Each public rendering round-trips the whole document without losing anything.""" + graph = _schema_graph() + round_trip = Graph().parse(data=render(graph), format=parse_format) + assert len(round_trip) == len(graph) + assert isomorphic(round_trip, graph) + + +@pytest.mark.parametrize(("render", "parse_format"), _renderings()) +def test_every_rendering_is_idempotent(render, parse_format: str) -> None: + """Re-rendering a rendered document reproduces it byte for byte.""" + first = render(_schema_graph()) + second = render(Graph().parse(data=first, format=parse_format)) + assert second == first + + +@pytest.mark.parametrize(("render", "parse_format"), _renderings()) +def test_every_rendering_ignores_incoming_blank_node_names(render, parse_format: str) -> None: + """Renaming every blank node must not change a single byte.""" + assert render(_relabelled(_schema_graph())) == render(_schema_graph()) + + +@pytest.mark.parametrize(("render", "parse_format"), _renderings()) +def test_every_rendering_ignores_insertion_order(render, parse_format: str) -> None: + """Inserting the same triples in another order must not change a single byte.""" + assert render(_reordered(_schema_graph())) == render(_schema_graph()) + + +def test_adding_a_class_rewrites_nothing_rdfc_alone_would_rewrite() -> None: + """An unrelated addition adds lines without rewriting the rest of the file. + + This is the behavior consumers commit generated RDF for, and the assertion + carries its own control: RDFC-1.0 assigns ``c14nN`` labels in a global order, + so a new blank node renumbers every label sorting after it. Checking the + baseline in the same test keeps a fixture that stopped demonstrating the + problem from turning the real assertions into no-ops. + """ + before, after = _schema_graph(), _extended_graph() + + def rewritten_lines(render) -> list[str]: + diff = difflib.unified_diff(render(before).splitlines(), render(after).splitlines(), n=0, lineterm="") + return [line[1:] for line in diff if line.startswith("-") and not line.startswith("---")] + + baseline = rewritten_lines(lambda graph: canonicalize_rdf_graph(graph, "turtle")) + assert len(baseline) > 10, "fixture no longer demonstrates the RDFC-1.0 relabelling it is meant to contrast" + + for render in (deterministic_turtle, lambda graph: canonicalize_rdf_graph(graph, "turtle", diff_stable=True)): + assert rewritten_lines(render) == [], "diff-stable output rewrote lines the edit did not touch" + + assert len(deterministic_turtle(after).splitlines()) > len(deterministic_turtle(before).splitlines()) + + +def test_adding_a_class_keeps_every_untouched_statement_verbatim() -> None: + """Every N-Triples line, blank-node label included, must survive the edit unchanged. + + The comparison is textual on purpose: parsing mints fresh rdflib identifiers, + which would hide exactly the relabelling this option exists to prevent. + """ + before = canonicalize_rdf_graph(_schema_graph(), "nt", diff_stable=True) + after = canonicalize_rdf_graph(_extended_graph(), "nt", diff_stable=True) + + missing = set(before.splitlines()) - set(after.splitlines()) + assert missing == set(), f"the edit rewrote statements it did not touch: {sorted(missing)[:3]}" + + def labels(document: str) -> set[str]: + return set(re.findall(r"_:(\S+)", document)) + + assert labels(before), "fixture produced no blank-node labels" + assert labels(before) <= labels(after) + + +def test_the_document_survives_a_named_node_rename() -> None: + """Renaming one IRI must change that subject's lines and leave the rest alone.""" + graph = _schema_graph() + renamed = Graph() + for prefix, namespace in graph.namespaces(): + renamed.bind(prefix, namespace) + for subject, predicate, obj in graph: + swap = {EX.Customer: URIRef(EX.Client)} + renamed.add((swap.get(subject, subject), predicate, swap.get(obj, obj))) + + assert not isomorphic(renamed, graph) + round_trip = Graph().parse(data=deterministic_turtle(renamed), format="turtle") + assert isomorphic(round_trip, renamed)