Skip to content

docs/benchmarks: record COPY FROM STDIN as tried-and-not-faster on laptop - #94

Merged
tonibergholm merged 1 commit into
mainfrom
docs/benchmarks-copy-result
May 24, 2026
Merged

tonibergholm merged 1 commit into
mainfrom
docs/benchmarks-copy-result

Conversation

@tonibergholm

Copy link
Copy Markdown
Member

The remaining tracked v1.1 follow-up from the restore review chain was `COPY FROM STDIN` as the "next 10× hop" after batched-INSERT. Prototyped 2026-05-24; the measurement on the same Docker-Desktop Postgres did NOT show the expected speedup.

Measured

N batched-INSERT COPY FROM STDIN Δ
10K 102 ms 104 ms +2%
100K 1.07 s 1.06 s flat
1M 10.34 s 11.88 s +15%

Manifest hashes stayed byte-identical — the change was correct, just not faster.

Diagnosis

At localhost-IPC scale, per-row write cost (WAL emission + btree insert) dominates the protocol overhead COPY skips. Multi-row INSERT at 1000 rows/statement already amortises the round-trips; COPY's text-encoding hot loop added CPU work without saving anything Postgres-side.

What the measurement does not rule out: COPY winning on a representative cloud-class instance where network RTT is higher than localhost IPC, WAL fsync is amortised by group commit, or where binary COPY (skipping the text-encoding pass) is used.

Action taken

  • Worktree + branch deleted. No code PR.
  • This doc-only PR updates three sites in `benchmarks.md` so the next person doesn't repeat the experiment on the same hardware expecting a different result. Future-revisit guidance points at binary COPY + managed Postgres.

Test plan

  • Doc-only change; no code touched.
  • `grep "COPY FROM STDIN"` matches three updated sites with the new framing.

…ptop

The remaining tracked v1.1 follow-up from the restore review chain
was `COPY FROM STDIN` as the "next 10× hop" after batched-INSERT.
Prototyped 2026-05-24; measurement on the same Docker-Desktop
Postgres did NOT show the expected speedup:

| N    | batched-INSERT | COPY FROM STDIN | Δ    |
|------|---------------:|----------------:|-----:|
| 10K  | 102 ms        | 104 ms          | +2%  |
| 100K | 1.07 s        | 1.06 s          | flat |
| 1M   | 10.34 s       | 11.88 s         | +15% |

Manifest hashes stayed byte-identical — the change was correct,
just not faster. Diagnosis: at localhost-IPC scale, per-row write
cost (WAL emission + btree insert) dominates protocol overhead,
which is what COPY skips. Multi-row INSERT at 1000 rows/statement
already amortises the round-trips; COPY's text-encoding hot loop
added CPU work without saving anything Postgres-side.

Worktree + branch deleted, no PR opened. Recorded here in
benchmarks.md so the next person doesn't repeat the experiment
expecting a different result on the same hardware. The doc now
points at the next viable path: binary COPY (skips text-encoding
CPU) on a managed-Postgres instance (where RTT + group commit
change the balance). Cloud-class hardware remains the unambiguous
§9-meeting path.

Three sites updated in benchmarks.md:
- Targets-vs-measured headline row (no longer claims COPY is the
  next 10× hop).
- §"the §9-shaped row" section's framing of the remaining gap.
- §"Coverage gaps" — the COPY bullet now records the negative
  result, the diagnosis, and the conditions under which a
  revisit might land differently.
Copilot AI review requested due to automatic review settings May 24, 2026 07:41

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Updates the architecture benchmark documentation to record that a COPY FROM STDIN prototype (as a potential next restore speedup after batched multi-row INSERT) was measured on the same laptop setup and did not improve performance.

Changes:

  • Reframes the rollback/restore benchmark notes to indicate COPY FROM STDIN was tried (2026-05-24) and wasn’t faster on the laptop setup.
  • Adds measured COPY vs batched-INSERT numbers (10K/100K/1M) plus a brief diagnosis and guidance for future revisits (e.g., binary COPY, managed Postgres).
  • Removes prior wording that implied COPY was the expected “next 10× hop” on this hardware.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

| `commit` (memory snapshot of N-row pgvector) | < 2 s @ 1M rows + 100 deltas | **5 ms @ 10K / 5.6 ms @ 100K / 18.81 ms @ 1M** (laptop) | text-only `episodes` schema; no embeddings, no deltas. ✓ comfortably within target. |
| `commit` (prompts-only, 16-byte system prompt, no memory) | n/a — degenerate input | **2.3 ms median** (Criterion) | ✓ no obvious blocker; not a §9 measurement |
| `rollback` (memory restore of N-row pgvector) | < 5 s @ 1M rows + 10 commits | **102 ms @ 10K / 1.07 s @ 100K / 10.34 s @ 1M** (laptop, post-batched-INSERT) | ⚠ 1M still ~2× over §9 on this hardware shape, but linear and predictable. Multi-row INSERT closed the 12× gap from the per-row path; the next 10× hop is `COPY FROM STDIN` (tracked, v1.1). |
| `rollback` (memory restore of N-row pgvector) | < 5 s @ 1M rows + 10 commits | **102 ms @ 10K / 1.07 s @ 100K / 10.34 s @ 1M** (laptop, batched-INSERT) | ⚠ 1M still ~2× over §9 on this hardware shape, but linear and predictable. Multi-row INSERT closed the 12× gap from the per-row path. `COPY FROM STDIN` was prototyped 2026-05-24 and **did not help on this hardware** (see "Coverage gaps" below for measured numbers and reasoning). A representative cloud-class machine remains the unambiguous §9-meeting path. |
Comment on lines +99 to +101
Diagnosis: on Docker-Desktop Postgres at localhost, the bottleneck for restore is per-row write cost (WAL emission + btree insert), not the frontend/backend protocol overhead COPY skips. Multi-row INSERT at 1000 rows/statement already amortises the protocol round-trip; COPY's text-encoding pass added CPU work in the hot loop without saving anything Postgres-side. Manifest hashes stayed byte-identical, so the change was correct — just not faster.

What the measurement does NOT rule out: COPY winning on a representative cloud-class instance where (a) network RTT is higher than localhost IPC, (b) WAL fsync overhead is amortised by group commit, or (c) the binary COPY format (not yet tried) skips the text-encoding CPU cost. If a future revisit happens, start with binary COPY and run against managed Postgres, not Docker-Desktop.
@tonibergholm
tonibergholm merged commit 1a35f8b into main May 24, 2026
10 checks passed
@tonibergholm
tonibergholm deleted the docs/benchmarks-copy-result branch May 24, 2026 07:43
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants