Repository navigation
Conversation
Keep scratch buffers separate for concurrent GEMMs while preserving reuse on each stream. Add ownership checks for dense, userbuffer and grouped workspaces, plus exact FP32/BF16 numerical coverage for concurrent large reductions. Signed-off-by: Matvey Saprykin <mtvey.s@gmail.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Cache cuBLAS scratch buffers by CUDA stream as well as device and GEMM mode. The current cache returns the same buffer to independent streams, so concurrent GEMMs can overwrite each other's partial sums.
On one H200, two FP32
NTGEMMs with a reduction dimension of 16384 produced incorrect results in 35 of 48 calls with the shared workspace. Separate workspaces produced zero mismatches against serial execution. The regression tests also check BF16 against an exact analytical reference.The workspace factory is identical in current
main(d0b4b32),stable,release_v2.19andrelease_v2.20: the cache key contains device and mode, but no stream.Type of change
Changes
The existing function signature, workspace sizes, GEMM dtypes and algorithm selection are unchanged. Memory use grows by one workspace per additional stream and mode.
Validation
d0b4b32: five failures before the patch, five passes after it.Checklist
Full L0 validation and human author review remain pending; this PR is a draft.