Skip to content

perf(encode): tag Fast slots only where the tag pays - #549

Draft
polaz wants to merge 1 commit into
mainfrom
perf/#548-fast-band-tags
Draft

polaz wants to merge 1 commit into
mainfrom
perf/#548-fast-band-tags

Conversation

@polaz

@polaz polaz commented Oct 1, 2026

Copy link
Copy Markdown
Member

Summary

  • Fast-band hash slots are tagged only where the tag pays. Level 1 on 1 MiB inputs is 4% faster than main and faster than v0.0.57 again, levels -1 and -7 by 10% and 4%, with output back to v0.0.57's bytes.
  • Smaller frames keep the tag wins they had: 4-20% from 16 to 256 KiB.

Changes

  • The branch probe (window_log >= 19, upstream zstd's useCmov boundary) runs on bare slots. There the tag's extra live values pushed the scan's data base, hash shift and table base to the stack, about 13 more instructions per position at the same probe count (callgrind: 46 against 33). The tagged branch-probe kernels are no longer built.
  • A copy-mode dictionary below 64 KiB leaves the table bare. Its fill tagged small frames where bare slots ran 2-4.5% faster; tags pay from 64 KiB. Upstream strips the tags when it copies a CDict table into the context (ZSTD_copyCDictTableIntoCCtx).
  • The cmov probe keeps the fill gate it had.

Performance

Runner (Xeon E5-2697 v4, x86_64), bench profile, five interleaved rounds of perf stat -r 3, ms for the whole loop. main is 16ecdc6. Bytes are one frame.

input level bytes (v0.0.57 / main / this / libzstd) v0.0.57 main this PR libzstd
z000033, 1 MiB 1 571124 / 571138 / 571124 / 571525 1024.9-1042.8 1043.2-1049.4 997.5-1008.9 756.2-771.7
z000033, 1 MiB -1 595173 / 595238 / 595173 / 595456 820.1-821.8 861.9-867.2 779.1-781.7 609.6-616.0
z000033, 1 MiB -7 699235 / 699013 / 699235 / 699611 494.8-499.9 490.7-492.9 470.1-472.4 347.4-350.3
z000033, 256 KiB 1 139684 / 139686 / 139686 / 139684 766.5-770.0 659.5-661.8 655.8-660.4 494.0-500.9
z000033, 32 KiB 1 17399 / 17399 / 17399 / 17399 863.8-873.9 679.7-687.0 678.5-688.2 432.8-438.1
z000033, 32 KiB -7 21869 / 21869 / 21869 / 21869 355.8-360.6 320.3-323.1 319.4-322.3 195.0-198.7
z000033, 10 KiB 1 6917 / 6917 / 6917 / 6917 330.5-335.9 317.0-322.0 317.0-322.0 202.6-207.0
10 KiB random + its dictionary 1 9218 / 9218 / 9218 / 9218 106.4-107.8 111.4-112.4 110.9-112.7 125.0-127.3
z000033 10 KiB + 110 KiB raw dictionary 1 7132 / 7132 / 7132 / 7132 274.2-279.9 267.5-282.7 269.3-275.9 178.1-181.9
z000033, 1 MiB (control, Double Fast) 3 495437 / 495437 / 495437 / 498911 3252.4-3285.0 2978.6-2998.8 2975.7-3038.7 1904.6-1912.2
z000033, 1 MiB (control, Double Fast) 4 496145 / 496145 / 496145 / 498591 3499.0-3509.3 3251.3-3294.6 3261.1-3295.3 1938.0-1981.6

Output is byte-identical to main everywhere the slot format did not change: z000033 at 256, 32 and 10 KiB, the log and low-entropy 1 MiB inputs, at levels -7 to 5. Where tags went off (z000033 512 KiB and 1 MiB on the Fast levels), output is v0.0.57's bytes at 1 MiB and within 52 bytes of main at 512 KiB.

Open: the dictionary row is still above v0.0.57 (111 against 107 ms). It runs fewer instructions than v0.0.57 (119.2 M against 123.1 M per 500 frames), fewer simulated D1 misses and as many branch misses, so it is not work; an aligned-layout comparison is pending.

Testing

fmt, clippy (workspace, all targets; library without default features), rustdoc with warnings denied and the workspace tests pass on macOS aarch64; the wasm module is smaller than main's. Linux i686 and x86_64 crate tests are pending.

Part of #548

- The branch probe (window_log >= 19, upstream zstd's useCmov boundary) runs bare: there the tag's extra live values spilled the scan's data base, hash shift and table base, about 13 more instructions per position at the same probe count, which made level 1 slower than v0.0.57 on frames from 512 KiB. The tagged branch-probe kernels are no longer built.
- A copy-mode dictionary below 64 KiB leaves the table bare: its fill tagged small frames where bare slots ran faster. Upstream strips the tags when it copies a CDict table into the context.
- The cmov probe keeps its tags, which pay 4-20% from 16 to 256 KiB.

Part of #548
@coderabbitai

coderabbitai Bot commented Oct 1, 2026

Copy link
Copy Markdown

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@codecov

codecov Bot commented Oct 1, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 75.00000% with 1 line in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
zstd/src/encoding/simple/fast_matcher.rs 75.00% 1 Missing ⚠️

📢 Thoughts on this report? Let us know!

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant