Conversation
- The branch probe (window_log >= 19, upstream zstd's useCmov boundary) runs bare: there the tag's extra live values spilled the scan's data base, hash shift and table base, about 13 more instructions per position at the same probe count, which made level 1 slower than v0.0.57 on frames from 512 KiB. The tagged branch-probe kernels are no longer built. - A copy-mode dictionary below 64 KiB leaves the table bare: its fill tagged small frames where bare slots ran faster. Upstream strips the tags when it copies a CDict table into the context. - The cmov probe keeps its tags, which pay 4-20% from 16 to 256 KiB. Part of #548
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: true
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Codecov Report❌ Patch coverage is
📢 Thoughts on this report? Let us know! |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Changes
window_log >= 19, upstream zstd'suseCmovboundary) runs on bare slots. There the tag's extra live values pushed the scan's data base, hash shift and table base to the stack, about 13 more instructions per position at the same probe count (callgrind: 46 against 33). The tagged branch-probe kernels are no longer built.ZSTD_copyCDictTableIntoCCtx).Performance
Runner (Xeon E5-2697 v4, x86_64), bench profile, five interleaved rounds of
perf stat -r 3, ms for the whole loop. main is 16ecdc6. Bytes are one frame.Output is byte-identical to main everywhere the slot format did not change: z000033 at 256, 32 and 10 KiB, the log and low-entropy 1 MiB inputs, at levels -7 to 5. Where tags went off (z000033 512 KiB and 1 MiB on the Fast levels), output is v0.0.57's bytes at 1 MiB and within 52 bytes of main at 512 KiB.
Open: the dictionary row is still above v0.0.57 (111 against 107 ms). It runs fewer instructions than v0.0.57 (119.2 M against 123.1 M per 500 frames), fewer simulated D1 misses and as many branch misses, so it is not work; an aligned-layout comparison is pending.
Testing
fmt, clippy (workspace, all targets; library without default features), rustdoc with warnings denied and the workspace tests pass on macOS aarch64; the wasm module is smaller than main's. Linux i686 and x86_64 crate tests are pending.
Part of #548