Skip to content

About

An 8k context window sees 41% of CUAD's answers - and the 41% is almost entirely the four fields on the first page

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

8 Commits

Folders and files

Repository files navigation

contract-reader

A contract reader that truncates to 8,000 characters sees 41% of the answers — and the 41% is almost entirely the four fields printed on the first page.

Status: the model-free half is complete and measured on all 510 contracts. The 14B reader is not written yet; the findings below do not need it, and they change what it should be asked to do. See What is not measured yet.

The corpus

CUAD — the Contract Understanding Atticus Dataset. 510 commercial contracts, labelled clause by clause by attorneys rather than crowdworkers, over 41 categories that matter commercially.

Questions 20,910
Contracts 510
Clause categories 41
Answerable 6,702 (32%)
Unanswerable 14,208 (68%)
Median contract 52,563 chars (33,143 on a 2026-10-03 re-fetch, see Result 2)

Every contract is asked every category, so most questions are about a clause that simply is not in that contract. That 68% is not a quirk of the sampling — it is asserted in tests/test_corpus.py, along with the fact that every gold span really is a substring of its own contract.

pip install -e ".[dev,fetch]"   # stdlib-only package; pyarrow is needed only to fetch
python -m pytest                # fresh clone: 12 pass, 23 corpus tests skip
python demo.py                  # runs on a synthetic toy corpus until CUAD is fetched
python scripts/fetch_data.py    # rebuild data/cuad.json (38 MB, not in git, ~1 min)
python scripts/measure.py       # prints every table below
python -m pytest                # with the corpus: 35 pass (~3 min)
python demo.py --json           # headline numbers as JSON
python demo.py --contract my_contract.txt   # categories the detector flags in your file

Set CONTRACTREADER_DATA=/some/dir to keep cuad.json elsewhere; the fetch, the tests, measure.py and demo.py all honour it. The fetch streams each parquet part to a .part file, resumes an interrupted download with an HTTP Range request, and writes cuad.json atomically.

CUAD's canonical data.zip now 404s behind an organisation rename, so the fetch rebuilds the corpus from HuggingFace's parquet conversion. The two splits are not shaped alike — test carries whole contracts, train is chunked and augmented — so the script deduplicates to one entry per (contract, category) against the longest context. It verifies it produced 510 × 41 = 20,910 questions and refuses to write anything else, because a fetch that quietly produces a different corpus invalidates every number below.

Result 1 — two scores that are not about reading contracts

Strategy Accuracy
Answer nothing, ever 0.679
Answer something, ever 0.321

Refusing to read scores 68% (14,208 / 20,910; an earlier README said 0.681, which did not match its own counts). Any reader reporting plain accuracy on CUAD is reporting mostly this, which is why everything below uses balanced accuracy instead.

Result 2 — what a context window can even see

Window Gold spans kept
4,000 2,115 / 6,702 31.6%
8,000 2,719 / 6,702 40.6%
16,000 3,603 / 6,702 53.8%
32,000 4,776 / 6,702 71.3%
64,000 5,829 / 6,702 87.0%

A 14B model does not read a 50,000-character contract. It reads the front of it.

Re-fetch drift (checked 2026-10-03). The tables in this README were computed on the original fetch. Re-fetching today from the same HuggingFace parquet gives the same 510 contracts, 20,910 questions and 6,702 answerable ones, and finding 4 reproduces to the third decimal, but the deduplicated contract texts differ slightly: the median contract is 33,143 characters (not 52,563) and the window counts are 2,129 / 2,714 / 3,613 / 4,774 / 5,834 (31.8 / 40.5 / 53.9 / 71.2 / 87.0%). Per-category "kept by 8k" moves by at most one point (Effective Date 84%, Anti-Assignment 7%, Insurance 4%). The 94% vs 20% split in finding 3 is unchanged. The originals are kept below; scripts/measure.py prints whatever your fetch produces.

Result 3 — the average is carried by four trivial fields

This is the one worth the repository. Truncation does not lose a random 59% of clauses.

Clause category n Median position Kept by 8k
Document Name 510 0% 100%
Parties 509 1% 98%
Agreement Date 470 1% 93%
Effective Date 390 1% 83%
…
Insurance 166 61% 5%
Cap On Liability 275 67% 5%
Change Of Control 121 68% 6%
Anti-Assignment 374 79% 6%
Governing Law 437 84% 8%
Uncapped Liability 111 69% 2%
n Kept by 8k
The four header fields 1,879 94%
Every other category 4,823 20%

The clauses that survive truncation are the ones you could find with your eye: the title, the parties, the date. The clauses that do not survive are the ones anyone pays a lawyer to find. A reader truncating to 8k scores respectably on average while being blind to every commercial term in the contract.

Result 4 — how much of this is just word counting

A detector with no model in it: per category, learn from the training contracts which terms distinguish contracts that have the clause from those that do not, then score an unseen contract by how many of those terms it contains.

Split by contract, never by question — every contract is asked 41 questions, so splitting questions at random would put the same contract text on both sides of the line.

Split n Accuracy Always-absent Balanced Sens Spec
train 16,728 0.817 0.681 0.795 0.733 0.857
val 3,116 0.730 0.682 0.713 0.664 0.761
test 1,066 0.689 0.653 0.668 0.603 0.734

All three splits are reported on purpose. The 0.129 train-to-test gap is what fitting 41 categories on 408 contracts buys you, and quoting the train number would be quoting the overfit.

On test, counting words beats refusing to read by 0.036 accuracy and 0.168 balanced accuracy over chance. That is the floor a model has to clear before it has earned its hour on the card.

What is not measured yet

The 14B reader. Not written. Findings 2 and 3 say what it should be asked, and it is not "read this contract": at 8k it cannot see the clauses that matter, so the honest comparison is against a retrieval step that picks which 8k to show it. Measuring a truncating reader against these numbers would mostly measure truncation.

Span extraction. Everything here is the presence question — is this clause in this contract — not which words are it. Presence is the half that the 68% base rate makes treacherous, so it is the half worth doing first.

Layout

src/contractreader/corpus.py    CUAD, parsed. 20,910 questions, spans verified
src/contractreader/split.py     80/15/5 by contract, seeded and stable
src/contractreader/reader.py    window arithmetic + the model-free cue detector
scripts/measure.py              every table above
src/contractreader/sample.py    synthetic toy corpus for demo.py and hermetic tests (not CUAD)
demo.py                         end-to-end demo, --json, --contract FILE
tests/                          12 hermetic tests + 23 against the real corpus (skip without it)

Licence

Code: MIT, see LICENSE. CUAD is not committed; scripts/fetch_data.py downloads it, and it stays under its own licence (CC BY 4.0, per The Atticus Project).

About

An 8k context window sees 41% of CUAD's answers - and the 41% is almost entirely the four fields on the first page

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages