Skip to content

[pull] master from tensorflow:master - #8860

Merged
pull[bot] merged 22 commits into
Cache-Cloud:masterfrom
tensorflow:master
Sep 30, 2026
Merged

pull[bot] merged 22 commits into
Cache-Cloud:masterfrom
tensorflow:master

Conversation

@pull

@pull pull Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

See Commits and Changes for more details.


Created by pull[bot] (v2.0.0-alpha.4)

Can you help keep this open source service alive? 💖 Please sponsor : )

vishwakt and others added 22 commits September 24, 2026 19:41
The MLIR-generated GPU kernels for these ops, and the TF-dialect
lowering in lower_tf.td that follows the same pattern, select x itself
when x == 0:

  select(x == 0, x, x / y)

That returns -0 for a -0 input. The GPU kernels also flush denormals,
so a subnormal x compares equal to zero and comes back unchanged, which
is how xdivy(1e-40, 1e-40) returns 1e-40 on GPU (#127913). Select zero
instead, as BinaryNoNanPat already does for div_no_nan and as the
tf2xla kernels for these ops already do.

The complex xlog1py kernel also compared x against 1+0j rather than 0,
so xlog1py(1, y) returned 1 for any y, and xlog1py(0, -1) returned nan.
Compare against zero.

The new tests run only on GPU. On CPU these ops run the Eigen functors
in cwise_ops.h, which #127946 covers.
The C++ baselines in gpu_binary_ops_test.cc returned x when x == 0,
which is what the kernels did before this change. For x = -0 that is
-0, which ExpectStrictlyEqual tells apart from the +0 the kernels now
return, so the Half tests failed. Return zero, as the kernels do.
PiperOrigin-RevId: 990590739
`Subgraph::Prepare` and `Subgraph::Invoke` acquire `Delegate::workspace_mutex_` to serialize access to the shared `xnn_workspace` and its intrusive `first_user` linked list of `xnn_runtime` instances. However, `Subgraph::Create` (`xnn_create_runtime_v4`), `Subgraph::~Subgraph` (`xnn_delete_runtime`), and `Delegate::~Delegate` (`xnn_release_workspace`) previously mutated `workspace->first_user` and `workspace->ref_count` or freed the `xnn_runtime` without holding `workspace_mutex_`.

Share `workspace_mutex_` via `std::shared_ptr<std::mutex>` between `Delegate` and `Subgraph` (matching the ref-counted lifetime of `xnn_workspace`), acquire `workspace_mutex_` around `xnn_create_runtime_v4` in `Subgraph::Create`, before resetting `runtime_` in `Subgraph::~Subgraph`, and before resetting `workspace_` in `Delegate::~Delegate`, clean up `runtime_ptr` under `workspace_mutex_` if `StopBuildStep()` fails, and log an error and return `kTfLiteError` if `runtime_` is null in `Subgraph::Prepare` and `Subgraph::Invoke`.

PiperOrigin-RevId: 990599327
Integrate cl/983398077 (2e332453d92) removed the callers mentioned
in the comment. Remove it.

PiperOrigin-RevId: 990604835
…lation

This change introduces ALG_DOT_BF16_BF16_FP8X3 and ALG_DOT_BF16_BF16_FP8X4 to PrecisionConfig in xla_data.proto and plumbs them through StableHLO and MHLO attribute translators.

PiperOrigin-RevId: 990613954
and restore input schedule when we change schduler config.

PiperOrigin-RevId: 990614610
So that sanitizers won't preclude LLVM passes from combining
them into an fma. Sanitizer instrumentation can break basic blocks,
which prevents LLVM from fusing FMA's since FMA's must come from
same-basic-block pairs.

This change is necessary to maintain the same numerics when we
have full msan support, which is upcoming.

PiperOrigin-RevId: 990627586
…SelectOp` with the index tie-breaking condition rather than `MaxOp`. While `SelectOp` produces the same value for normal numbers, it has different IEEE-754 semantics for `NaN`s (dropping `NaN` when it appears on the LHS) and signed zeros, which caused divergent `NaN` propagation behavior and prevented XLA from simplifying unused-index reductions to `kMaximum`. Rely on `MaxOp` for `selected_value` while keeping `SelectOp` for `selected_index`, and add a unit test in `StablehloBuilderTest` verifying the generated reducer body.

PiperOrigin-RevId: 990642043
… by value so it is copied while the lock is held.

PiperOrigin-RevId: 990657976
Adds the llvm_xz (v5.8.3) external archive in WORKSPACE.bazel and provides third_party/xz.BUILD to build the liblzma library.

PiperOrigin-RevId: 990674437
…in YNNPACK.

Instead of folding query heads into the row/sequence dimension for GQA during prefill, split Q into [n_kv, g] and expand K, V, and the mask to 5D so they broadcast across the group dimension. This updates the K and V transposes to 5D, avoids splitting and fusing logits around mask addition, and fuses the [n_kv, g] head dimensions back together after the P @ V matmul. GQA folding is retained for decode paths and sequence-major inputs.

PiperOrigin-RevId: 990686744
PiperOrigin-RevId: 990706829
…r Pow keeps raising InvalidArgumentError for negative exponents.

PiperOrigin-RevId: 990728125
PiperOrigin-RevId: 990730579
…or unusual tilings

It previously made unverified assumptions about tiling (that it is of the form (A, 128)(...)) and simply propagated the tiling. This is clearly incorrect. For example, even for tilings of this form, it can result in an invalid layout if a 2D shape is squeezed into 1D and the tile remains 2D

PiperOrigin-RevId: 990737389
Updates the concurrency group to append the run ID when the workflow is triggered via workflow_dispatch. This allows manual runs (such as autobisect) to execute in parallel, while push events continue to run sequentially to prevent clashes.

PiperOrigin-RevId: 990760148
@pull pull Bot locked and limited conversation to collaborators Sep 30, 2026
@pull pull Bot added the ⤵️ pull label Sep 30, 2026
@pull
pull Bot merged commit 103b097 into Cache-Cloud:master Sep 30, 2026
6 of 11 checks passed
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.