An auditable AI agent that turns messy bibliographic references into verified links to the exact published work (DOIs), and sends everything it is not sure about to a human review queue instead of guessing.
Independent project. It uses the public Crossref and OpenAlex APIs and the Anthropic API. The sample document describes a fictional organisation.
Organisations that publish in several languages spend skilled time matching citations to the exact work cited: the right edition of a report series, the right language version, the right DOI. Citations arrive in every style, with truncated titles, wrong years and mistyped DOIs.
- String matching alone breaks on that mess.
- A language model alone is worse: it will confidently invent a plausible DOI.
This project shows the pattern in between. Code handles what code can decide. The model steps in only where the evidence is ambiguous, and only chooses among records that were actually retrieved. Anything still uncertain goes to a person, with the evidence attached. Quality and cost are measured on a gold set, and CI fails when quality drops.
flowchart TD
D["Document"] --> S["Find the bibliography,<br/>split into citations"]
S --> P["Parse fields<br/>(regex, then Claude in batches)"]
P --> A{"DOI in the<br/>citation?"}
A -- "yes, and the registered<br/>title agrees" --> L["Linked"]
A -- "no, or it points<br/>to another work" --> Q["Search Crossref + OpenAlex,<br/>merge, score every candidate"]
Q -- "score ≥ 0.85 and a clear<br/>lead over the runner-up" --> L
Q -- "plausible but<br/>ambiguous" --> J["Claude chooses among<br/>numbered candidates, or none"]
J -- "confident and<br/>corroborated" --> L
J -- "none of them" --> G["Search agent: reformulated queries,<br/>max 6 steps, retrieved ids only"]
Q -- "nothing plausible" --> G
G -- "confident and<br/>corroborated" --> L
G -- "otherwise" --> R["Human review queue<br/>(best candidates + reasons)"]
J -- "unsure" --> R
Every decision records its method (doi_in_text, deterministic, llm_adjudication,
agent_search) and a trace of each step, so any link can be audited afterwards.
| Risk | Control |
|---|---|
| The model invents a DOI | It never writes identifiers. Adjudication returns a candidate number (a schema enum). The agent may only submit an identifier it retrieved: anything else is rejected and logged. |
| The model is confidently wrong | A model-assisted link also needs a minimum deterministic match score. Otherwise it goes to review. |
| A coin-flip between near-identical records | Automatic links need a clear lead over the runner-up (margin rule). |
| Runaway agent cost | A hard step budget per reference, and a model call only where the scores are ambiguous. Every run reports its token use and cost. |
| A mistyped DOI in the citation | The registered title is checked against the citation before the DOI is trusted. |
| Silent drift in evaluations | Recorded replays must match exactly. A missing recording fails the run instead of changing the result. |
| Load on public infrastructure | Caching, a minimum interval per host, bounded retries, and an optional contact email for the "polite pool". |
Requires Python 3.10+.
python -m venv .venv
# Windows (PowerShell): .venv\Scripts\Activate.ps1 macOS/Linux: source .venv/bin/activate
pip install -e ".[dev]"
pytest # 70 offline tests, no API key needed
# Deterministic steps only: no key, no cost
refresolver run examples/aurora_working_paper.md --no-llm --out out/
# Full pipeline with Claude
export ANTHROPIC_API_KEY=sk-ant-... # PowerShell: $env:ANTHROPIC_API_KEY="sk-ant-..."
export REFRESOLVER_CONTACT_EMAIL=you@example.org
refresolver run examples/aurora_working_paper.md --out out/
refresolver cite "Autor, D. (2016), Why are there still so many jobs?, JEP"The model steps also run on any OpenAI-compatible endpoint: a local model through Ollama, vLLM or llama.cpp, or an LLM gateway such as governed-llm-gateway.
ollama pull qwen2.5:7b # or llama3.1:8b; any model with tool calling
export REFRESOLVER_PROVIDER=openai_compatible # PowerShell: $env:REFRESOLVER_PROVIDER="openai_compatible"
export REFRESOLVER_MODEL=qwen2.5:7b
export REFRESOLVER_PRICE_INPUT=0 REFRESOLVER_PRICE_OUTPUT=0
refresolver run examples/aurora_working_paper.md --out out-local/
refresolver eval evals/gold_references.jsonl --cache-dir evals/recordings-local --out evals/results-localSmall local models do not always honour a forced tool call. The adapter then accepts a JSON object from the text reply, and anything else counts as "no decision", which sends the reference to the review queue rather than linking it: a weaker model lowers automation, not precision. Compare the three result folders to see the trade-off.
A run writes three files to the output folder:
report.md: a summary table with links;resolutions.json: everything, including traces and candidate scores;review_queue.csv: one row per uncertain item, with the suggestion, the alternatives and an emptydecisioncolumn for the reviewer.
The resolver is also an MCP server with two read-only tools, resolve_citation and
resolve_bibliography. Claude Desktop configuration, Windows example:
{
"mcpServers": {
"reference-resolver": {
"command": "C:\\path\\to\\reference-resolver-agent\\.venv\\Scripts\\refresolver.exe",
"args": ["serve"],
"env": { "ANTHROPIC_API_KEY": "sk-ant-...", "REFRESOLVER_CONTACT_EMAIL": "you@example.org" }
}
}
}evals/gold_references.jsonl holds 22 citations with known answers. They cover edition traps,
a French edition, formatting noise, truncated titles, a year off by one, a mistyped DOI, a
DOI-only citation, and four items that must not be linked.
refresolver eval evals/gold_references.jsonl --cache-dir evals/recordings --out evals/resultsThis reports:
- precision of automatic links;
- recall;
- wrong links, and false links on items that have no DOI;
- review rate, and how often the right answer is among the review suggestions;
- model cost per reference.
The live run records every API and model response. After that, the same evaluation replays offline, identically and for free, and CI runs it as a quality gate: it fails below 0.95 precision or on any false link. Details are in evals/README.md.
Run it with and without --no-llm to see what the model steps add, and at what cost.
With Claude Sonnet 5 (claude-sonnet-5), search agent on. CI replays this run on every push
and fails below 0.95 precision, on any false link, or if any database search failed.
| Metric | Value |
|---|---|
| References | 22 |
| Precision of automatic links | 1.00 |
| Recall (resolvable, linked correctly) | 1.00 |
| Wrong links | 0 |
| False links on unresolvable items | 0 |
| Sent to review | 2 (9%) |
| Review items with the answer among suggestions | 0 |
| Unresolved | 2 |
| Source errors (searches that failed) | 0 |
| Model cost | $0.0851 ($0.00387/ref) |
Deterministic steps only (--no-llm), same references:
| Metric | Value |
|---|---|
| References | 22 |
| Precision of automatic links | 1.00 |
| Recall (resolvable, linked correctly) | 0.83 |
| Wrong links | 0 |
| False links on unresolvable items | 0 |
| Sent to review | 4 (18%) |
| Review items with the answer among suggestions | 3 |
| Unresolved | 3 |
| Source errors (searches that failed) | 0 |
| Model cost | $0.0000 ($0.00000/ref) |
What the live runs taught, in order:
- The first run measured an outage, not the resolver. Crossref rejected every search (the request asked for a field Crossref does not accept), and OpenAlex rejected titles containing "?". The resolver carried on with what was left, so the numbers looked plausible. The evaluation now counts source errors, and a run with any of them fails the gate.
- A match without a DOI goes to review, not to a link. A university-repository copy of an OECD manual had been linked instead of the published version.
- Two scoring bugs, found by the second run. Names cited as "Wilkinson, M. D." were read with an initial as the surname, so the right paper lost its author evidence; and a short title found inside a long citation counted as a full match, so a short comment in another journal outscored the cited paper. Both are fixed and covered by tests.
- Compare the two tables to see what the model adds: it reads messy citations (capitals, missing quotes, odd orders) into fields the scoring can use, and adjudicates the close calls.
Llama 3.1 8B and Qwen 2.5 7B through Ollama (8k context, temperature 0) on a laptop CPU (Intel
Core i7-13620H, 16 GB, integrated graphics), same 22 references. Results in
evals/results-llama3.1-8b-ctx8k and evals/results-qwen2.5-7b-ctx8k.
| Metric | Claude Sonnet 5 | Llama 3.1 8B, local | Qwen 2.5 7B, local | No model |
|---|---|---|---|---|
| Precision of automatic links | 1.00 | 1.00 | 1.00 (0.95 before the rule below) | 1.00 |
| Recall (resolvable, linked correctly) | 1.00 | 1.00 | 1.00 | 0.83 |
| Wrong or false links | 0 | 0 | 0 (1 before the rule) | 0 |
| Sent to review | 2 (9%) | 1 (5%) | 2 (9%) | 4 (18%) |
| Unresolved | 2 | 3 | 2 | 3 |
| Model cost | $0.0851 | $0.00 | $0.00 | $0.00 |
| Run time | not recorded | about 16 minutes (laptop CPU) | about 16 minutes (laptop CPU) | not recorded |
- Qwen made the first false link of any run, and the design changed because of it. For a web
page with no DOI (
OECD.AI Policy Observatory, Live data on AI), the registries returned three unrelated works sharing words of the title, tied on score, none by a cited author. Qwen picked one, a Zenodo record on benchmarking platforms, with confidence 0.95. A model may now choose among candidates but not against the evidence: when none of a candidate's authors appears in the citation, the case goes to a person (model_link, tested with this case). Replayed from the recorded answers, the same run gives precision 1.00 and no false link; Claude's and Llama's results are unchanged, since none of their model links had that profile. - A failed call is now recorded too. Qwen's first extraction batch timed out and the run fell back to heuristic parsing, as designed, but the failure was not in the recording, so the run could not be replayed. Failures are now recorded and replayed as failures.
- The first Llama run scored 0.89 recall, and the cause was the adapter. Llama returned tool
arguments in the wrong JSON types (the list of references as a string, "1" for 1, "null" for
null), so every extraction and most adjudications were discarded. Even then it made no wrong
link: a weaker model lowered automation, not precision, as designed. Arguments are now
repaired against the tool's schema (
coerce_to_schema, tested with the recorded shapes), and the re-run matched Claude on precision and recall. - Where a local model adds value is adjudication and search. Llama decided three close calls among retrieved candidates; Qwen decided one and found two through the search agent. Batch extraction still mostly falls back to heuristic parsing (2 of 22 references parsed by either model), which this gold set tolerates; harder citation styles would not. Next: extract one reference per call for small models, and measure it.
- Twenty-two references is a small set, and these are one run's numbers per model on one laptop.
| Variable | Default | Purpose |
|---|---|---|
ANTHROPIC_API_KEY |
none | Enables the model steps. Without it, the deterministic pipeline still runs. |
REFRESOLVER_PROVIDER |
anthropic |
or openai_compatible (Ollama, vLLM, llama.cpp, an LLM gateway) |
REFRESOLVER_MODEL |
claude-sonnet-5 |
The model or deployment name. Compare models on the gold set. |
REFRESOLVER_BASE_URL |
http://localhost:11434/v1 |
OpenAI-compatible endpoint (Ollama by default) |
REFRESOLVER_API_KEY |
none | Only if the endpoint needs one (a gateway or cloud); a local model does not |
REFRESOLVER_PRICE_INPUT / _OUTPUT |
2.0 / 10.0 |
USD per million tokens, for the cost report. |
REFRESOLVER_AUTO_ACCEPT |
0.85 |
Deterministic score for a link without the model |
REFRESOLVER_MIN_MARGIN |
0.05 |
Required lead over the runner-up |
REFRESOLVER_REVIEW_FLOOR |
0.55 |
Below this, a candidate is not worth a reviewer's time |
REFRESOLVER_LLM_ACCEPT |
0.80 |
Model confidence required for a model-assisted link |
REFRESOLVER_MAX_AGENT_STEPS |
6 |
Hard budget for the search agent, per reference |
REFRESOLVER_CONTACT_EMAIL |
none | Identifies you to Crossref and OpenAlex (polite pool) |
OPENALEX_API_KEY |
none | Optional OpenAlex key |
REFRESOLVER_CACHE_DIR |
~/.cache/refresolver |
Recorded HTTP and model responses |
REFRESOLVER_OFFLINE |
false |
Replay recordings only |
The thresholds are policy: they trade automation against precision. Change them, re-run the evaluation, and compare.
- Translation linking: given a work, find its official translation in the target language (what translators need most). Registry links between language editions are sparse, so this needs title translation plus search, measured on its own gold set.
- Learning from reviewers: feed decisions from
review_queue.csvback into the gold set and use them to tune the thresholds. - More registries: DataCite (datasets, arXiv), national library catalogues.
- Throughput: concurrent resolution with a shared rate limiter; prompt caching for extraction batches.
- Native structured outputs as an alternative to forced tool calls.
src/refresolver/
resolver.py the pipeline, prompts and guardrails
scoring.py deterministic match scoring (title, year, authors)
sources.py Crossref and OpenAlex clients, candidate merging
llm.py model interface, Anthropic adapter, record/replay, cost meter
fetch.py HTTP with cache, politeness, retries, offline replay
text.py bibliography detection, splitting, DOIs, normalisation
evaluation.py gold-set metrics
report.py report.md, resolutions.json, review_queue.csv
mcp_server.py MCP tools
cli.py run | cite | eval | serve
tests/ offline tests with fakes and fixtures, including the real SDK path
evals/ gold set and (after a live run) recordings and results
examples/ a fictional working paper with a mixed-style bibliography
docs/ design decisions
MIT. See LICENSE.