Skip to content

[pull] master from tensorflow:master - #8859

Merged
pull[bot] merged 24 commits into
Cache-Cloud:masterfrom
tensorflow:master
Sep 30, 2026
Merged

pull[bot] merged 24 commits into
Cache-Cloud:masterfrom
tensorflow:master

Conversation

@pull

@pull pull Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

See Commits and Changes for more details.


Created by pull[bot] (v2.0.0-alpha.4)

Can you help keep this open source service alive? 💖 Please sponsor : )

junwhanahn and others added 24 commits September 29, 2026 11:46
…topology fingerprints.

When different hosts in a multi-host GPU job run different driver versions or have mismatched target configurations, each host builds a local `StreamExecutorGpuTopologyDescription` with a different fingerprint after `ExchangeTopologies`, causing compiled programs to diverge and hang in NCCL collective rendezvous. Add `GpuClientOptions::verify_topology_fingerprint` (enabled by default, overridable via `XLA_PJRT_GPU_VALIDATE_TOPOLOGY`) so that process 0 writes the expected topology fingerprint to `kv_store` and all other processes read and verify that their topology fingerprint matches.

PiperOrigin-RevId: 990436201
…r after waiting on stream.

In `PjRtStreamExecutorRawLoadedExecutable::Execute`, `BufferSequencingEvent::WaitForEventOnStream()` blocks until the definition event has completed or failed.

Previously, `ev->IsPredeterminedError()` was checked before `ev->WaitForEventOnStream()`. When an asynchronous host-to-device transfer (or any in-flight dependency) was still executing or pending, `ev->IsPredeterminedError()` returned false.

`ev->WaitForEventOnStream()` then blocked until the event completed with error. But because `ev->IsPredeterminedError()` was not checked after the wait, the error was ignored, and `launch_on_device` proceeded to invoke `RunAsync()` on an uninitialized/poisoned buffer (or asymmetrically skipped `RunAsync` on only a subset of devices, causing cross-node / collective GPU hangs and deadlocks).

Move `ev->WaitForEventOnStream()` before checking `ev->IsPredeterminedError()`, matching the behavior of the `else if (event)` branch where `xla::BlockUntilReady(event)` is called before inspecting `event.GetErrorIfPresent()`.

Add regression unit tests in `se_gpu_pjrt_client_test.cc` and `se_gpu_pjrt_client_multi_gpu_test.cc`.

PiperOrigin-RevId: 9904374
…dows build link error.

PiperOrigin-RevId: 990466106
In `TryRemoveTrivialCompare`, for a while loop induction variable `i` starting at initial value `c` (`i >= c`) compared against a constant `x`:
- `i < x` (`ComparisonDirection::kLt`) is trivially `false` when `x <= c`.
- `i > x` (`ComparisonDirection::kGt`) is trivially `true` only when `x < c` (strict inequality), not `x <= c`. When `x == c`, `i > c` evaluates to `false` on the first iteration (`i == c`) and `true` on subsequent iterations, so folding it to `true` miscompiles the first iteration.

PiperOrigin-RevId: 990471976
Log `se_gpu_topology->ToProto()` for each process before computing and comparing topology fingerprints during multi-host GPU initialization to make topology mismatches easier to debug.

PiperOrigin-RevId: 990482998
This replaces the generic `ToOrdered`/`FromOrdered` free functions and `if constexpr` logic with an explicit `OrderedTraits` template struct, specialized for `float`, `xla::bfloat16`, and `xla::half`. This maintains functional parity while ensuring strict type safety and compatibility with older toolchains in the OSS CI.

PiperOrigin-RevId: 990486882
…he new L4 1GPU runner for benchmark.

PiperOrigin-RevId: 990499379
The previous implementation completely only ever looked at the first element of the first tile, and scaled that. It didn't check trailing tiles and assumed the tiling was always of the form (A, B)(32 / bitwidth, 1).

The new implementation does not assume anything about the tiling (e.g. it can handle a 16-bit type with a (4, 1) tile instead of (2, 1)) and can handle scaling n-D tiles such as (2, 4, 6)(2, 1) -> (2, 8, 6)(4, 1).

PiperOrigin-RevId: 990515869
…uDNN conv fusion.

PiperOrigin-RevId: 990521795
…voking the build script directly.

Executing the CI build entrypoint outside of Bazel fails to resolve
sibling package imports after recent cross-module refactoring removed
the inline shell interpreter directive. Add the repository root to the
Python module search path at runtime and invoke the Python interpreter
explicitly across remaining workflow scripts.

PiperOrigin-RevId: 990523829
Newer versions of CUPTI (up to CUDA 12.8+) introduce additional `CUpti_ActivityOverheadKind` enum values for runtime-triggered module loading, lazy function loading, command buffer full, activity buffer request, and UVM activity initialization overheads. Previously, these overhead kinds were unrecognized by XProf and appeared as `<UNKNOWN>` in traces.

Add `GetExtraActivityOverheadKindString12080` to `cuda_version_variants` (selected at build time via `if_cuda_newer_than("12_8", ...)` without `#if CUDA_VERSION` macros) and use it in `GetActivityOverheadKindString`. Format unknown or unrecognized overhead kinds as `Overhead::UNKNOWN:<int_value>` so the underlying enum value remains visible for diagnosis.

PiperOrigin-RevId: 990524975
The latest LLVM integration deprecates an MLIR API we use. A prior change migrated most call sites but missed some that were abstracted away by template functions; this change updates those call sites.

PiperOrigin-RevId: 990535948
llvm::sys::getDefaultTargetTriple() can return a triple that omits
the vendor component, e.g. "aarch64-linux-gnu". Constructing an
llvm::Triple from it parses the components positionally ($arch,$vendor,$OS),
which in this case results in an empty $OS. Reading $OS from the resulting
triple results in a crash with `LLVM ERROR: unsupported operating system`.
(This can be reproduced with the upcoming full msan support, which
calls `TargetTriple.getOS()`.)

Normalize the triple string consistently so that we always store
canonicalized triples. That is, we don't save "aarch64-linux-gnu"; we save
"aarch64-unknown-linux-gnu". Note that we do preserve empty triples ("")
(vs. normalizing them to "unknown") so that llvm::EngineBuilder::selectTarget
continues to fall back to the host process triple (it checks for "").

Also update the hardcoded host triple for oberon_b200 and oberon_b300 in
gpu_topology and its test from "aarch64-linux-gnu" to the
normalized "aarch64-unknown-linux-gnu". (this was the original, correct
name---see cl/893544830---but it was changed to accommodate XLA:CPU's
TargetMachineOptions. This change restores the original, correct string.

PiperOrigin-RevId: 990537506
…allowing DelinearizeAsync

to be overloaded explicitly.

Reverts changelist 990024422

PiperOrigin-RevId: 990540455
Ignoring returned future is always an error

PiperOrigin-RevId: 990545500
So that all generated code is instrumented.

This paves the way for the upcoming msan support.

PiperOrigin-RevId: 990569171
…ops.

Inside a while loop carrying xla_disable_while_loop_copies, XLA is not free to
insert a relayout copy. This change exposes IsWhileLoopCopyDisabled on
ComputationLayoutConstraints and threads it to the TPU convolution output
layout tie-break to prevent inserting relayout copies inside copy-disabled while loops.

PiperOrigin-RevId: 990570717
…ted` kernel.

- Add `FuseA4W2DRQFullyConnectedPass` to collapse blockwise Q/DQ patterns (symmetric 32-element i4/e8m0 dynamic activations + centered per-channel i2 weights) into a single `tfl.fully_connected` carrying `tfl.quant_spec = {spec = "cint2_fp32_int4_e8m0_drq", act_dilation = ...}` and a per-axis `i2` `tfl.pseudo_qconst`.

- Implement `ParseQuantSpec` and `EvalA4W2DRQ` in the `FullyConnected` reference kernel (bumping max version to 15) to evaluate `a4w2_drq_v1` and reject unrecognized `quant_spec` payloads.

PiperOrigin-RevId: 990578615
Based on https://en.cppreference.com/cpp/types/numeric_limits:

`::min()` returns "the smallest positive normal value of the given
floating-point type" while `::lowest()` returns "the lowest finite value".

PiperOrigin-RevId: 990578654
Ignoring returned future is always an error

Reverts 543710f

PiperOrigin-RevId: 990580556
@pull pull Bot locked and limited conversation to collaborators Sep 30, 2026
@pull pull Bot added the ⤵️ pull label Sep 30, 2026
@pull
pull Bot merged commit 39489a6 into Cache-Cloud:master Sep 30, 2026
1 check failed
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.