Skip to content

chore: feature integration branch for the ENVITED-X pipeline - #3

Draft
jdsika wants to merge 2 commits into
mainfrom
feat/envited-x-pipeline
Draft

jdsika wants to merge 2 commits into
mainfrom
feat/envited-x-pipeline

Conversation

@jdsika

@jdsika jdsika commented Sep 14, 2026

Copy link
Copy Markdown

Integration branch — tracking only. Do not merge.

What this branch is

feat/envited-x-pipeline is upstream RDFLib/rdflib@main with every contribution this
organisation is still carrying cherry-picked on top, and nothing else. It is the branch to
check out when you need an rdflib that behaves the way our pipeline expects before
upstream has merged the relevant PRs.

It is rebased on upstream main, never merged into. An empty cherry-pick list is the goal
state: it would mean upstream has merged everything, and the branch could be deleted.

Same convention as ASCS-eV/linkml#14.

What it currently carries

Commit Upstream PR Change
81fe1fbe RDFLib/rdflib#3543 JSON-LD collection syntax can silently duplicate shared RDF lists
342175bf RDFLib/rdflib#3546 Deterministic N-Triples serialization via an opt-in canon parameter

Both cherry-picked cleanly onto a090f3f9 with no conflicts.

Each is also tracked individually as a mirrored PR: #1 and #2.

What consumes it

Nothing yet, deliberately. asam-openx-standards pins released rdflib==7.6.0
through scripts/requirements.txt, and diffable-rdf depends on released rdflib too.

That pin is the point: rdflib performs the final serialization of every generated OWL and
SHACL artifact, so a change here moves artifact bytes. Consuming this branch would mean
building from a fork, which the pipeline's reproducibility rules do not allow — the lock
records released versions precisely so a third party can obtain them.

The relationship is therefore one-directional: what lands upstream here eventually arrives
as a released rdflib that the pipeline can bump to, with regenerated artifacts in the same
change. The merged #3504 (the Turtle variant
of the RDFLib#3543 defect) is the precedent — diffable-rdf had to work around it until it was
released.

Converter.to_collection walked an rdf:List chain checking only that each cell
carries one rdf:first/rdf:rest pair and that the chain is acyclic. It did not
check how many statements point into the chain.

The @list form inlines the whole chain at a single reference and writes its
cells nowhere else, so when a cell is referenced from more than one place each
referring list rendered the shared cells again. The round trip then gained
triples:

    _:tail rdf:first "b" ; rdf:rest rdf:nil .
    _:head rdf:first "a" ; rdf:rest _:tail .
    ex:s1  ex:p _:head .
    ex:s2  ex:p _:tail .            # second reference into the chain

serialized as two independent @list values, one per referring subject, so the
tail was emitted twice: 6 triples in, 8 out, not isomorphic.

This is the same defect class as RDFLib#3504, which fixed TurtleSerializer and
LongTurtleSerializer by counting statements that point into the chain. The
JSON-LD converter has its own chain walk and was not covered by that fix, so
it gets the same check: a cell that is the object of more than one statement
is not rendered as a list, and the chain is written out with explicit
rdf:first/rdf:rest instead.

An empty chain still renders as "@list": [], which is rdf:nil rather than a
list structure, so nothing is lost there.
…eter

The N-Triples serializer writes statements in the order the store iterates
them, which follows set iteration and so varies between processes. The same
graph therefore serializes to different bytes in different runs, which shows
up as diff noise in version-controlled RDF.

PR RDFLib#3008 solved this for longturtle with an opt-in `canon` parameter that
canonicalizes the graph and sorts, fixing RDFLib#1890. This gives NTSerializer the
same parameter, with the same default of False, so ordinary serialization is
untouched and still streams statement by statement.

`canon=True` canonicalizes with to_canonical_graph and writes the rendered rows
in sorted order. Sorting the rendered rows rather than the triples keeps the
order total: two distinct terms can share a str(), which would leave their
relative order down to the store iteration order that this is meant to remove,
whereas identical rows mean identical statements.

Measured over six PYTHONHASHSEED values on a graph with a hub, a shared list
tail and a deep blank-node chain: 5 distinct documents before, 1 after.
NT11Serializer inherits it by subclassing.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant