[pull] master from tensorflow:master - #8859
Merged
Merged
Conversation
…topology fingerprints. When different hosts in a multi-host GPU job run different driver versions or have mismatched target configurations, each host builds a local `StreamExecutorGpuTopologyDescription` with a different fingerprint after `ExchangeTopologies`, causing compiled programs to diverge and hang in NCCL collective rendezvous. Add `GpuClientOptions::verify_topology_fingerprint` (enabled by default, overridable via `XLA_PJRT_GPU_VALIDATE_TOPOLOGY`) so that process 0 writes the expected topology fingerprint to `kv_store` and all other processes read and verify that their topology fingerprint matches. PiperOrigin-RevId: 990436201
…r after waiting on stream. In `PjRtStreamExecutorRawLoadedExecutable::Execute`, `BufferSequencingEvent::WaitForEventOnStream()` blocks until the definition event has completed or failed. Previously, `ev->IsPredeterminedError()` was checked before `ev->WaitForEventOnStream()`. When an asynchronous host-to-device transfer (or any in-flight dependency) was still executing or pending, `ev->IsPredeterminedError()` returned false. `ev->WaitForEventOnStream()` then blocked until the event completed with error. But because `ev->IsPredeterminedError()` was not checked after the wait, the error was ignored, and `launch_on_device` proceeded to invoke `RunAsync()` on an uninitialized/poisoned buffer (or asymmetrically skipped `RunAsync` on only a subset of devices, causing cross-node / collective GPU hangs and deadlocks). Move `ev->WaitForEventOnStream()` before checking `ev->IsPredeterminedError()`, matching the behavior of the `else if (event)` branch where `xla::BlockUntilReady(event)` is called before inspecting `event.GetErrorIfPresent()`. Add regression unit tests in `se_gpu_pjrt_client_test.cc` and `se_gpu_pjrt_client_multi_gpu_test.cc`. PiperOrigin-RevId: 9904374
…dows build link error. PiperOrigin-RevId: 990466106
In `TryRemoveTrivialCompare`, for a while loop induction variable `i` starting at initial value `c` (`i >= c`) compared against a constant `x`: - `i < x` (`ComparisonDirection::kLt`) is trivially `false` when `x <= c`. - `i > x` (`ComparisonDirection::kGt`) is trivially `true` only when `x < c` (strict inequality), not `x <= c`. When `x == c`, `i > c` evaluates to `false` on the first iteration (`i == c`) and `true` on subsequent iterations, so folding it to `true` miscompiles the first iteration. PiperOrigin-RevId: 990471976
Log `se_gpu_topology->ToProto()` for each process before computing and comparing topology fingerprints during multi-host GPU initialization to make topology mismatches easier to debug. PiperOrigin-RevId: 990482998
This replaces the generic `ToOrdered`/`FromOrdered` free functions and `if constexpr` logic with an explicit `OrderedTraits` template struct, specialized for `float`, `xla::bfloat16`, and `xla::half`. This maintains functional parity while ensuring strict type safety and compatibility with older toolchains in the OSS CI. PiperOrigin-RevId: 990486882
…he new L4 1GPU runner for benchmark. PiperOrigin-RevId: 990499379
The previous implementation completely only ever looked at the first element of the first tile, and scaled that. It didn't check trailing tiles and assumed the tiling was always of the form (A, B)(32 / bitwidth, 1). The new implementation does not assume anything about the tiling (e.g. it can handle a 16-bit type with a (4, 1) tile instead of (2, 1)) and can handle scaling n-D tiles such as (2, 4, 6)(2, 1) -> (2, 8, 6)(4, 1). PiperOrigin-RevId: 990515869
…ewriter PiperOrigin-RevId: 990520367
…uDNN conv fusion. PiperOrigin-RevId: 990521795
…voking the build script directly. Executing the CI build entrypoint outside of Bazel fails to resolve sibling package imports after recent cross-module refactoring removed the inline shell interpreter directive. Add the repository root to the Python module search path at runtime and invoke the Python interpreter explicitly across remaining workflow scripts. PiperOrigin-RevId: 990523829
Newer versions of CUPTI (up to CUDA 12.8+) introduce additional `CUpti_ActivityOverheadKind` enum values for runtime-triggered module loading, lazy function loading, command buffer full, activity buffer request, and UVM activity initialization overheads. Previously, these overhead kinds were unrecognized by XProf and appeared as `<UNKNOWN>` in traces.
Add `GetExtraActivityOverheadKindString12080` to `cuda_version_variants` (selected at build time via `if_cuda_newer_than("12_8", ...)` without `#if CUDA_VERSION` macros) and use it in `GetActivityOverheadKindString`. Format unknown or unrecognized overhead kinds as `Overhead::UNKNOWN:<int_value>` so the underlying enum value remains visible for diagnosis.
PiperOrigin-RevId: 990524975
PiperOrigin-RevId: 990529446
PiperOrigin-RevId: 990529913
The latest LLVM integration deprecates an MLIR API we use. A prior change migrated most call sites but missed some that were abstracted away by template functions; this change updates those call sites. PiperOrigin-RevId: 990535948
llvm::sys::getDefaultTargetTriple() can return a triple that omits
the vendor component, e.g. "aarch64-linux-gnu". Constructing an
llvm::Triple from it parses the components positionally ($arch,$vendor,$OS),
which in this case results in an empty $OS. Reading $OS from the resulting
triple results in a crash with `LLVM ERROR: unsupported operating system`.
(This can be reproduced with the upcoming full msan support, which
calls `TargetTriple.getOS()`.)
Normalize the triple string consistently so that we always store
canonicalized triples. That is, we don't save "aarch64-linux-gnu"; we save
"aarch64-unknown-linux-gnu". Note that we do preserve empty triples ("")
(vs. normalizing them to "unknown") so that llvm::EngineBuilder::selectTarget
continues to fall back to the host process triple (it checks for "").
Also update the hardcoded host triple for oberon_b200 and oberon_b300 in
gpu_topology and its test from "aarch64-linux-gnu" to the
normalized "aarch64-unknown-linux-gnu". (this was the original, correct
name---see cl/893544830---but it was changed to accommodate XLA:CPU's
TargetMachineOptions. This change restores the original, correct string.
PiperOrigin-RevId: 990537506
…allowing DelinearizeAsync to be overloaded explicitly. Reverts changelist 990024422 PiperOrigin-RevId: 990540455
Ignoring returned future is always an error PiperOrigin-RevId: 990545500
So that all generated code is instrumented. This paves the way for the upcoming msan support. PiperOrigin-RevId: 990569171
…ops. Inside a while loop carrying xla_disable_while_loop_copies, XLA is not free to insert a relayout copy. This change exposes IsWhileLoopCopyDisabled on ComputationLayoutConstraints and threads it to the TPU convolution output layout tie-break to prevent inserting relayout copies inside copy-disabled while loops. PiperOrigin-RevId: 990570717
…Backend PiperOrigin-RevId: 990574854
…ted` kernel.
- Add `FuseA4W2DRQFullyConnectedPass` to collapse blockwise Q/DQ patterns (symmetric 32-element i4/e8m0 dynamic activations + centered per-channel i2 weights) into a single `tfl.fully_connected` carrying `tfl.quant_spec = {spec = "cint2_fp32_int4_e8m0_drq", act_dilation = ...}` and a per-axis `i2` `tfl.pseudo_qconst`.
- Implement `ParseQuantSpec` and `EvalA4W2DRQ` in the `FullyConnected` reference kernel (bumping max version to 15) to evaluate `a4w2_drq_v1` and reject unrecognized `quant_spec` payloads.
PiperOrigin-RevId: 990578615
Based on https://en.cppreference.com/cpp/types/numeric_limits: `::min()` returns "the smallest positive normal value of the given floating-point type" while `::lowest()` returns "the lowest finite value". PiperOrigin-RevId: 990578654
Ignoring returned future is always an error Reverts 543710f PiperOrigin-RevId: 990580556
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to subscribe to this conversation on GitHub.
Already have an account?
Sign in.
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
See Commits and Changes for more details.
Created by
pull[bot] (v2.0.0-alpha.4)
Can you help keep this open source service alive? 💖 Please sponsor : )