Skip to content

Cache cuBLAS workspace by CUDA stream - #3624

Closed
Vtmpas wants to merge 1 commit into
NVIDIA:mainfrom
Vtmpas:codex/cublas-workspace-per-stream
Closed

Vtmpas wants to merge 1 commit into
NVIDIA:mainfrom
Vtmpas:codex/cublas-workspace-per-stream

Conversation

@Vtmpas

@Vtmpas Vtmpas commented Oct 5, 2026

Copy link
Copy Markdown

Description

Cache cuBLAS scratch buffers by CUDA stream as well as device and GEMM mode. The current cache returns the same buffer to independent streams, so concurrent GEMMs can overwrite each other's partial sums.

On one H200, two FP32 NT GEMMs with a reduction dimension of 16384 produced incorrect results in 35 of 48 calls with the shared workspace. Separate workspaces produced zero mismatches against serial execution. The regression tests also check BF16 against an exact analytical reference.

The workspace factory is identical in current main (d0b4b32), stable, release_v2.19 and release_v2.20: the cache key contains device and mode, but no stream.

Type of change

  • Documentation change
  • Bug fix (non-breaking change which fixes an issue)
  • New feature
  • Breaking change
  • Infra/Build change
  • Code refactoring

Changes

  • Preserve workspace reuse within each CUDA stream and isolate concurrent streams.
  • Cover dense, userbuffer and grouped workspace ownership, plus concurrent FP32/BF16 GEMM numerics.
  • Include the regression in the PyTorch L0 launcher.

The existing function signature, workspace sizes, GEMM dtypes and algorithm selection are unchanged. Memory use grows by one workspace per additional stream and mode.

Validation

  • Five focused regression cases, repeated using the exact workspace functions from current main d0b4b32: five failures before the patch, five passes after it.
  • Runtime integration patch: the same five cases pass; repeated application is idempotent.
  • Hardware/software: one NVIDIA H200, CUDA 13.0, PyTorch 2.13.0+cu130, installed native TE 2.19.0.dev0+b5599209.
  • The changed Python workspace functions were loaded into that installed native TE build; this was not a rebuild of current main.
  • Repository pre-commit checks, changed-file pylint and L0 license check pass on macOS.
  • The full current-main L0 GPU suite has not been run.

Checklist

  • I have read and followed the contributing guidelines
  • The functionality is complete
  • I have commented my code, particularly in hard-to-understand areas
  • I have made corresponding changes to the documentation
  • My changes generate no new warnings
  • I have added tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes

Full L0 validation and human author review remain pending; this PR is a draft.

Keep scratch buffers separate for concurrent GEMMs while preserving reuse on each stream. Add ownership checks for dense, userbuffer and grouped workspaces, plus exact FP32/BF16 numerical coverage for concurrent large reductions.

Signed-off-by: Matvey Saprykin <mtvey.s@gmail.com>
@github-actions github-actions Bot added the community-contribution PRs from external contributor outside the core maintainers, representing community-driven work. label Oct 5, 2026
@Vtmpas Vtmpas closed this Oct 5, 2026
@Vtmpas
Vtmpas deleted the codex/cublas-workspace-per-stream branch October 5, 2026 10:00
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant