Skip to content

[PyTorch] Deprecate grouped linear and MLP environment gates - #3584

Open
zhongbozhu wants to merge 3 commits into
NVIDIA:mainfrom
zhongbozhu:cleanup/grouped-mlp-env-gates
Open

zhongbozhu wants to merge 3 commits into
NVIDIA:mainfrom
zhongbozhu:cleanup/grouped-mlp-env-gates

Conversation

@zhongbozhu

Copy link
Copy Markdown
Collaborator

Description

Please include a brief summary of the changes, relevant motivation and context.

Fixes # (issue)

Type of change

  • Documentation change (change only to the documentation, either a fix or a new content)
  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Infra/Build change
  • Code refactoring

Changes

Please list the changes introduced in this PR:

  • Change A
  • Change B

Checklist:

  • I have read and followed the contributing guidelines
  • The functionality is complete
  • I have commented my code, particularly in hard-to-understand areas
  • I have made corresponding changes to the documentation
  • My changes generate no new warnings
  • I have added tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes

@github-actions github-actions Bot added the community-contribution PRs from external contributor outside the core maintainers, representing community-driven work. label Sep 29, 2026
@greptile-apps

greptile-apps Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 4/5

[Medium risk] Removes environment variable gates for grouped linear and MLP features.

The PR does not appear safe to merge while unrelated CPU-only forwards can still initialize CUDA through the grouped-MLP fusion callbacks.

Findings

  1. P1 Unrelated forwards initialize CUDA ▶

Summary

The PR removes environment gates for grouped-linear parameter layout and grouped-MLP fusion, updates the Mixtral example accordingly, and adjusts tests and QA launchers.

  • Explicit constructor options now select grouped parameters without the former environment variable.
  • Eligible grouped-MLP patterns are registered without the former import-time gate.

Diagram

%%{init: {'theme': 'neutral'}}%%
flowchart LR
  A[GroupedLinear → activation → GroupedLinear] --> B{Recipe and pattern eligible?}
  B -->|No| C[Separate operations]
  B -->|Yes| D{Device and kernels supported?}
  D -->|No| C
  D -->|Yes| E[Fused grouped MLP]
Loading

Reviews (4) · Last reviewed commit: "Merge branch 'main' into cleanup/grouped..."

Comment thread transformer_engine/pytorch/ops/fused/grouped_mlp.py
Comment thread docs/examples/te_mixtral/run_finetune_ep.py
Comment thread tests/pytorch/test_grouped_mlp.py Outdated
@zhongbozhu

Copy link
Copy Markdown
Collaborator Author

/te-ci pytorch L1

1 similar comment
@zhongbozhu

Copy link
Copy Markdown
Collaborator Author

/te-ci pytorch L1

@ptrendx ptrendx added the 2.21 label Sep 30, 2026
Comment thread transformer_engine/pytorch/ops/fused/grouped_mlp.py Outdated
Comment thread transformer_engine/pytorch/ops/fused/grouped_mlp.py Outdated
Comment thread transformer_engine/pytorch/ops/fused/grouped_mlp.py
Comment thread docs/envvars.rst Outdated
Comment thread docs/examples/te_mixtral/tutorial_accelerate_hf_mixtral_with_te.ipynb Outdated
Comment thread tests/pytorch/test_grouped_mlp.py Outdated
Comment thread tests/pytorch/test_grouped_mlp.py Outdated
Comment thread tests/pytorch/test_grouped_mlp.py Outdated
Comment thread transformer_engine/pytorch/ops/fused/grouped_mlp.py Outdated
@zhongbozhu

Copy link
Copy Markdown
Collaborator Author

/te-ci pytorch L1

) -> list[FusibleOperation]:
"""Apply joint GroupedLinear + scaled GLU + GroupedLinear fusion."""

if not torch.cuda.is_available():

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Unrelated forwards initialize CUDA If a machine has a GPU, the first forward of even an unrelated CPU-only Sequential now queries the current CUDA device before checking whether the operations or recipe could use grouped-MLP fusion. The unary fusion callback also checks support before candidacy. This can initialize CUDA in a process that later forks workers, causing a worker to fail when it tries to use CUDA. Check the recipe and operation pattern before probing the device.

Knowledge Base Used: PyTorch tensor and operator abstractions

Signed-off-by: zhongboz <zhongboz@nvidia.com>
Signed-off-by: zhongboz <zhongboz@nvidia.com>
@zhongbozhu
zhongbozhu force-pushed the cleanup/grouped-mlp-env-gates branch from 1ac8a2c to 91032db Compare October 2, 2026 23:57
Signed-off-by: vthumbe1503 <vthumbe@nvidia.com>
@vthumbe1503

Copy link
Copy Markdown
Collaborator

LGTM. Thanks for cleaning these env variables. All @timmoon10's comments seem to be resolved as well. @zhongbozhu I resolved the merge conflicts.

@vthumbe1503

Copy link
Copy Markdown
Collaborator

/te-ci pytorch

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

2.21 community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants