Skip to content

fail fast when FabricManager fails to come up - #908

Merged
tariq1890 merged 1 commit into
mainfrom
fail-fast-on-fm-fail
Aug 10, 2026
Merged

fail fast when FabricManager fails to come up#908
tariq1890 merged 1 commit into
mainfrom
fail-fast-on-fm-fail

Conversation

@tariq1890

Copy link
Copy Markdown
Contributor

Today, the driver containers go into Running mode even if Fabric Manager doesn't come up successfully. Given that Fabric Manager is a must-have for GPUs to be usable in NVSwitch-based nodes, the driver container should fail fast instead. This would make it easier for users to debug any driver container issues relating to Fabric Manager as they get immediate feedback

Signed-off-by: Tariq Ibrahim <tibrahim@nvidia.com>

@rahulait rahulait left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM 🚀

@shivamerla

Copy link
Copy Markdown
Contributor

Looks good to me as well! This needs a good amount of testing to make sure there are no spurious failures with some of these daemons.

@tariq1890
tariq1890 merged commit b14676c into main Aug 10, 2026
81 of 116 checks passed
@tariq1890
tariq1890 deleted the fail-fast-on-fm-fail branch August 10, 2026 18:54
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants