Skip to content

[pull] master from tensorflow:master - #8861

Merged
pull[bot] merged 70 commits into
Cache-Cloud:masterfrom
tensorflow:master
Sep 30, 2026
Merged

pull[bot] merged 70 commits into
Cache-Cloud:masterfrom
tensorflow:master

Conversation

@pull

@pull pull Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

See Commits and Changes for more details.


Created by pull[bot] (v2.0.0-alpha.4)

Can you help keep this open source service alive? 💖 Please sponsor : )

vishwakt and others added 30 commits July 23, 2026 11:40
ParseAndCheckBoxSizes returned early with num_boxes=0 when both boxes
and box_index were empty, skipping all rank validation. A rank-1 empty
boxes tensor (shape [0] instead of [0, 4]) or a rank-2 empty box_index
then reached Tensor::tensor<T, NDIMS>() with the wrong rank, aborting
the process with a fatal CHECK failure instead of raising a catchable
InvalidArgumentError. This crashed CropAndResizeGradImage and
CropAndResizeGradBoxes.

Validate the ranks of boxes and box_index before the empty early
return. Well-formed empty inputs (boxes of shape [0, 4] with box_index
of shape [0]) keep working, and empty tensors of the correct rank are
still accepted.

Fixes #123397
The messages previously rendered as "boxes must be 2-D[0]". They now
read "boxes must be 2-D, got [0]".
Syncs the branch with upstream master so the copybara import runs
against current revisions of the files this change touches. The
array_ops kernel_tests BUILD file changed upstream since this branch
was cut, in a different target from the one this change edits, so the
merge is clean.
…flow_framework

libtensorflow_framework.so is linked with tf_framework_version_script.lds,
which read:

    tensorflow {
      global:
        *;
    };

That declares every symbol global and has no local stanza, so the library
exports its statically linked third-party code as well as its own. Among
those are the LLVM symbols, which then reach the process-wide resolution
scope and collide with any other LLVM loaded into the same process. In the
reported crash an LLVM 15 brought in by llvmlite and numba resolved
llvm::raw_svector_ostream::write_impl against TensorFlow's LLVM 18 and
died on the mismatched layout.

Add a local stanza hiding *llvm* and *mlir*, and add *llvm* to
tf_private_symbols.lds so macOS, which already hides *mlir* through
-unexported_symbols_list, covers the same set.

The change is deliberately narrow. TensorFlow's own symbols stay exported,
because custom op libraries loaded through tf.load_op_library resolve
REGISTER_OP and REGISTER_KERNEL_BUILDER against this library, which is why
the script exports everything today.

Verified with lld, the linker the Linux builds use via -fuse-ld=lld, on a
shared object built from the shipped script: llvm and mlir symbols are
dropped from .dynsym while the tensorflow namespace and the TF_ C API stay
exported. Linking also succeeds, with and without --undefined-version, when
no symbol matches either pattern, so configurations built without LLVM are
unaffected.

Fixes #104038
Per review, the substring patterns were both too broad and too narrow.

Too broad: *mlir* also hides TensorFlow's own symbols whose names happen
to contain the string, such as tensorflow::tf_xla_test_use_mlir, which
would become local and fail to resolve for anything linking against it.

Too narrow: *llvm* is lowercase and never matched the LLVM C API at all,
whose symbols are spelled LLVM*. Those kept leaking into the dynamic
symbol table, which is the same collision class the change is meant to
close.

Match the namespaces through extern "C++" and add an explicit LLVM*
prefix for the C API. Verified with lld on the shipped script: llvm::,
mlir:: and LLVM* symbols are dropped while tensorflow::, including
tf_xla_test_use_mlir, and the TF_ C API stay exported. Linking still
succeeds when no pattern matches.

Also adds *LLVM* to the macOS list, which had the same case gap next to
its existing *mlir* entry.
…ormations

Filter transformation kernels on GPU (TransformFilter and ReverseTransformFilter)
use 32-bit indexing via To32Bit. When filter element count exceeds INT32_MAX,
signed integer overflow caused fatal process crashes in gpu_launch_config.h
(CHECK_GE work_element_count >= 0) and out-of-bounds illegal memory access in
ShuffleInTensor3Simple.

This change adds FastBoundsCheck upfront validation in conv_ops_impl.h,
conv_ops_fused_impl.h, conv_grad_input_ops.cc, and conv_grad_input_ops_3d.cc,
and guards non-positive output sizes in conv_2d_gpu.h.

Fixes #87457
Fixes #87438
Fixes #87454
…comparison, and remove memory-heavy python test
# Conflicts:
#	tensorflow/python/ops/numpy_ops/tests/np_test.py
Hiding the llvm:: and mlir:: C++ namespaces broke the build, because
libtensorflow_cc links against those symbols from
libtensorflow_framework. It imports 5,351 of them in the 2.21.0 Linux
wheel, and 3,816 LLVM symbols in a macOS nightly, where the *llvm* and
*LLVM* patterns would have hidden them just the same.

The LLVM C API is different: no other TensorFlow library imports any of
it, on either platform, so hiding it cannot break a link. Keep only
that. On Linux the pattern is LLVM*. On macOS it is _LLVM*, which
carries the Mach-O leading underscore and, unlike *LLVM*, does not also
match C++ names that mention LLVM types.
In the CUDA builds, libtensorflow_cc links tfcompile's
InitializeTargets(), which calls the LLVMInitialize* functions for each
LLVM target and imports them from libtensorflow_framework, so hiding
all of LLVM* broke that link. They are the only LLVM C API functions
that TensorFlow or XLA code calls. Export them again and keep hiding
the rest of the C API.
With the rank checks first, the early return for empty boxes and
box_index no longer protects anything: well-formed empty inputs pass
the column and row checks on their own. It only let malformed empty
boxes through, such as shape [0, 5] or [2, 0], which every
crop-and-resize kernel then accepted. Remove it, and extend the test
to cover the boxes gradient with a rank-2 box_ind, empty boxes with the
wrong number of columns, and well-formed empty inputs to the boxes
gradient.
Run clang-format on ParseAndCheckBoxSizes and pyink on
testMalformedEmptyBoxesRaisesError, the code this change touches. No
functional change.
…dd OP_REQUIRES on callers

Per maintainer review (dmiltr3):
- Revert if (TF_PREDICT_FALSE(out.size() <= 0)) return; from TransformFilter
  and ReverseTransformFilter in conv_2d_gpu.h. These caused silent skips with
  uninitialized output buffers being passed to cuDNN.
- Add upfront OP_REQUIRES bounds checks in conv_grad_filter_ops_launcher.cc
  and conv_grad_filter_ops_3d.cc before allocate_temp, so invalid filters
  produce a loud InvalidArgument error rather than silent data corruption.

Fixes: #115734
Per maintainer review (dmiltr3):
- Add explicit #include <limits> in conv_grad_filter_ops_3d.cc,
  conv_grad_filter_ops_launcher.cc, conv_grad_input_ops.cc,
  conv_grad_input_ops_3d.cc, and conv_ops_fused_impl.h for IWYU compliance.
- Remove stray newline before kernel_shape_util.h include in
  conv_grad_input_ops_3d.cc.
The overflow guard kept determined_size itself in range, but a large
negative size next to a -1, such as [-1, INT64_MIN], still made
input_size_split_dim - determined_size overflow when computing the -1
size. Reject a negative size before summing it, as the later check
already did for the error message, so that 0 <= determined_size <=
input_size_split_dim and only the upper bound needs a guard.

Extend testSizeSplitsOverflowRaises to int32 split sizes and to an
overflow next to a -1, accept ValueError for graph mode, and call
array_ops.split instead of the generated op, which drops the
array_ops_gen dependency this change had added.
input_shape.dim_size(split_dim) was converted to Tlen implicitly, so
with int32 or int8 size_splits an input size above the maximum of Tlen
was truncated: int32 sizes [1, 2] matched an input of size 2**32 + 3,
and the split silently dropped the rest of the input. Check the size
in int64 before converting it.

Split sizes are now checked for being non-negative before they are
summed, and the -1 size is bounded by the input size, so the later loop
that checked every size again can no longer fail; remove it. Also build
the overflow error with errors::InvalidArgument, as the rest of the
file does.

The int32 case of testSizeSplitsOverflowRaises used an input larger
than INT32_MAX to reach the kernel in graph mode, which the new check
rejects first. Use a small input, where graph mode stops at the shape
function instead, and test the truncation separately.
Fix log fragment extraction and ensure main function is called correctly.
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
…ks with NumPy error precedence, sub-2D tests
tensorflower-gardener and others added 27 commits September 30, 2026 00:24
PiperOrigin-RevId: 990789026
Adds `xla::gpu::PerDeviceState<T>`, which combines pre-allocated `VectorStorage` for device ordinals in `[0, num_devices)` with copy-on-write `CowStorage` fallback for ordinals outside `[0, num_devices)`.

PiperOrigin-RevId: 990801463
PiperOrigin-RevId: 990820587
…avoid race conditions when sharing the stream between multiple threads.

PiperOrigin-RevId: 990824347
Imported from GitHub PR openxla/xla#49717

📝 Summary of Changes
Export control_predecessors() to python

🎯 Justification
JAX has an experimental API to allow python to modify the Thunk scheduling.
To build valid one, we need to know the control depedence.

🚀 Kind of Contribution
✨ New Feature, 🧪 Tests

📊 Benchmark (for Performance Improvements)
No speed-up expected by this exposure.

🧪 Unit Tests:
The new Python API have been tested.

🧪 Execution Tests:
xla/python/xla_hlo_test.py::TestHloModule.testHloInstructionControlPredecessors
Copybara import of the project:

--
9c42acd7cc52413d38d851e914f65af1f5f21135 by Frederic Bastien <fbastien@nvidia.com>:

Export control_predecessors() to python

Merging this change closes #49717

PiperOrigin-RevId: 990831343
Imported from GitHub PR openxla/xla#49741

## Summary
- Remove all `toolkit_version_ < {6,x,0}` runtime checks in `gemm_rewriter.cc` — they are dead code since `rocm_config.h.tpl` requires ROCm >= 7.1 at compile time (`#error` if `TF_ROCM_VERSION < 70100`).
- Simplify the `toolkit_version_ >= {7,0,0}` guard on Swish matching to plain `if (is_rocm)`, since it is always true.
- Remove the now-callerless `TurnF8DotWithUnsupportedOutputTypeIntoF32()` helper.
- Clean up corresponding dead branches in `gemm_rewriter_fp8_test.cc` (ROCm < 6.0 skip, and five if/else blocks selecting CHECK patterns for ROCm < 6.2).

Copybara import of the project:

--
996cd22d95e984681c58ad0187e75c1fb698c666 by Marco Minutoli <marco.minutoli@amd.com>:

[ROCm] Remove dead ROCm version checks from GemmRewriter

rocm_config.h.tpl requires ROCm >= 7.1 (#error if < 70100), making all
runtime checks against ROCm 6.x unreachable. Remove the dead guards:

- toolkit_version_ < {6,2,0} output-type workaround at FP8 dot rewrite
- toolkit_version_ < {6,0,0} FP8 availability check (has_fp8_support remains)
- toolkit_version_ < {6,2,0} output-type restriction in CreateF8CustomCall
- toolkit_version_ >= {7,0,0} always-true guard on Swish matching (simplified)
- TurnF8DotWithUnsupportedOutputTypeIntoF32() helper (no remaining callers)
- Corresponding dead branches in gemm_rewriter_fp8_test.cc

Merging this change closes #49741

PiperOrigin-RevId: 990831715
Imported from GitHub PR openxla/xla#49300

Scan rewrite lowering converts scan operations into call-based expressions. However, the GPU kernel emitter cannot compute indexing maps for Call operations, causing compilation failures for GPU kernels containing scan. Exact error:

```
E0000 00:00:1789415358.746826 1768463 indexing_analysis.cc:1748] ComputeOutputToInputIndexing is not implemented for opcode call
F0000 00:00:1789415358.746906 1768463 computation_partitioner.cc:269] Check failed: operand_maps.size() == 1 (0 vs. 1)
```

This PR inlines the scan computation calls, resolving the indexing analysis failure. CUDA and ROCm backends avoid this issue because they run CallInliner in their device-specific Convolution Canonicalization optimization pipelines that execute before the kernel code generation.
This issue came to light after openxla/xla@33d1a58 stopped converting scan operations to custom calls.
Copybara import of the project:

--
ac1221a33121f1a07207d0aad698d85f43ae3bf0 by Akhil Goel <akhil.goel@intel.com>:

Inline scan calls

--
76a47249786fc94a1b18854444e6b5ba94e72f7b by Akhil Goel <akhil.goel@intel.com>:

Add CallInliner pass

Merging this change closes #49300

PiperOrigin-RevId: 990834717
1. when constructed from a map of interval constraints we have not checked if any of the constraints is unsatisfiable - now we sent it trough ctor that calls AddConstraint and handles that for us.

2. added a check if const constraint is satisfiable, e.g. we can get something like `4 in [0, 3]` that should make the whole construct invalid.

PiperOrigin-RevId: 990839235
… instruction.

`InstructionAnnotation` and `GetInstructionAnnotationMetadata` where being computed twice for the same values.

PiperOrigin-RevId: 990844885
…r-transform-overflow

PiperOrigin-RevId: 990863209
This is needed to support fractional vGPUs (see openxla/xla#49252). Collective fusions are disabled in this case and all collectives should fallback on CPU initiated NCCL without RMA access.

PiperOrigin-RevId: 990867700
Replace host-launched D2D copy to scratch with in-kernel copies to scratch. Each block copies T/R of a tile to the scratch where T is the size of a tile and R is the number of ranks (world_size). Asymmetric block barriers, ensure that consumer blocks wait for all producer blocks to finish before emitting the output copy.

PiperOrigin-RevId: 990876953
Implemented the fix suggested here: openxla/stablehlo#2975.

PiperOrigin-RevId: 990878885
…name across DSOs

Imported from GitHub PR openxla/xla#49719

`AsyncValue` assigns each payload type a 16-bit type id by appending to a process-global `TypeInfoTable` and caching the resulting index in `GetTypeId<T>()`'s function-local static. This is unsafe once XLA is split across dynamically-linked libraries: XLA is built with `-fvisibility=hidden`, which demotes the `GetTypeId<T>()` local static to a per-DSO local. A type registered in more than one DSO therefore appends to the shared table more than once and receives a different id in each
DSO, so `IsType<T>()`/`DynCast<T>()` on an `AsyncValue` that crosses a DSO boundary can silently return the wrong answer.

Make registration idempotent by type name: `CreateTypeInfoAndReturnTypeIdImpl` now keys a `name->id` map on `typeid(T).name()` and returns the existing id when the same type is registered again, instead of allocating a fresh
one. The map lives behind a const-init mutex in the single out-of-line definition of the registration function, so all DSOs that resolve that symbol share one id per type. No caller changes are required.

This mirrors MLIR's `TypeID`, whose fallback resolver (r`egisterImplicitTypeID(getTypeName<T>())`) already keys on the type name rather than on registration order, giving a process-stable identity that survives across DSO boundary.
Copybara import of the project:

--
e94bb1d32f7a06d397a4d33cd4b1a4dd0c07a4e6 by Eugene Zhulenev <ezhulenev@openxla.org>:

[tsl:concurrency] Deduplicate AsyncValue type ids by type name across DSOs

Merging this change closes #49719

PiperOrigin-RevId: 990883982
…implifier and the transpose folding

Imported from GitHub PR openxla/xla#47649

## 📝 Summary of Changes

Two changes that together restore coalesced memory writes for scatters whose
update window dims come before the indexed dims (for example a vmapped
`segment_sum`, where the batch dim is a window dim and the segment ids index
the minor dim):

- `ScatterSimplifier` gets a `reorder_operand_dims_for_coalescing` option
  (enabled in the GPU pipelines only). When a scatter would write strided
  windows under the default layouts, the operand dims are permuted so that
  `scatter_dims_to_operand_dims` becomes the identity mapping and the writes
  become contiguous, with transposes around the scatter restoring the
  original order. This conditionally restores the canonicalization that
  087e1a5960 removed: scatters that are already coalesced keep the current
  no-transpose behavior, and so do variadic scatters, scatters with batching
  dims, and scatters whose written volume is small relative to the operand
  (where the transpose copies would cost more than the coalescing wins).
- `TryFoldTransposeIntoScatter` in the algebraic simplifier now checks the
  same profitability predicate and refuses folds that would make the written
  windows less contiguous. Without this, the fold both undoes the
  ScatterSimplifier rewrite above and defeats user-side workarounds (a
  restoring `.transpose()` gets folded back into the scatter, recreating the
  strided form). Folds that improve or preserve contiguity still fire.

The shared predicate `ScatterSimplifier::WriteRunLength` computes the length
of the contiguous run of operand elements each scatter index writes.

## 🎯 Justification

Fixes #47203 (cross-post of jax-ml/jax#39959): `segment_sum` under `vmap`
regressed ~6x on the reporter's GPU between JAX 0.9.1 and 0.9.2, and its
`out_axes=1` workaround regressed further in 0.10.2. Root cause chain:

- Since 087e1a5960, nothing in the GPU pipeline re-orients a scatter whose
  window dims are major, so every scatter index writes a strided window
  (strided atomics). Isolated on an RTX 2070 with identical indices and
  updates, only the operand orientation differing: 5.11 ms coalesced vs
  91.62 ms strided (18x).
- Since 8a228780ee, the unconditional transpose fold recreates the strided
  form out of the coalesced-scatter-plus-transpose pattern, so no HLO-level
  workaround survives.

With this PR, the issue's repro compiles back to the coalesced scatter plus
one cheap restoring transpose: 85-90 ms -> ~5.5 ms on the RTX 2070 (~16x),
details in the Benchmark section.

## 🚀 Kind of Contribution

⚡️ Performance Improvement / 🐛 Bug Fix

## 📊 Benchmark

Issue repro as a standalone HLO (vmapped segment_sum: scatter-add of
f32[1024,75960] updates into 12123 segments, hashed pseudo-random segment
ids, `hlo_runner_main_gpu --num_repeats=10`, RTX 2070, sm_75):

- before: 85-90 ms per execution, compiled scatter `f32[1024,12123]{1,0}`
  (strided window writes)
- after: 5.4-5.6 ms per execution (~16x), compiled scatter
  `f32[12123,1024]{1,0}` (coalesced) plus a restoring transpose

Isolated scatter kernel with identical indices and updates, only the operand
orientation differing (via JAX on the same GPU): 5.11 ms coalesced vs
91.62 ms strided (18x). A standalone `[12123,1024] -> [1024,12123]`
transpose costs 0.32 ms, so the inserted transposes are ~2% of the win.

## 🧪 Unit Tests:

- `scatter_simplifier_test.cc`: `ReordersOperandDimsForCoalescing`,
  `ReordersOperandDimsWithInsertedWindowDims` (non-monotonic
  `scatter_dims_to_operand_dims`), `DoesNotReorderCoalescedScatter`,
  `DoesNotReorderVariadicScatter`, `DoesNotReorderWhenUpdatesAreSmall`,
  `DoesNotReorderOperandDimsByDefault`.
- `algebraic_simplifier_test.cc`: `FoldTransposeIntoScatter` now uses a
  profitable example (window dims move to minor);
  `DoNotFoldTransposeIntoScatterWhenWritesBecomeStrided` covers the refused
  direction.

## 🧪 Execution Tests:

Verified on a single RTX 2070 (sm_75):

- `//xla/tests:scatter_test_nvgpu_any` (36 cases) and
  `//xla/tests:select_and_scatter_test_nvgpu_any` pass.
- `//xla/service/gpu:gpu_compiler_test_nvgpu_any` and the GPU scatter
  emitter lit tests (`add`, `sorted_indices`, `permuted_indices`,
  `permuted_sorted_indices`) pass.
- `run_hlo_module --platform=gpu --reference_platform=interpreter` on a
  small version of the repro: results match.
Copybara import of the project:

--
bd08ecb02fa0c0e2b705d15d56bdebfd1ac3551f by Stanislav Bardyuk <sbardyuk@google.com>:

[XLA:GPU] Keep scatter window writes coalesced in ScatterSimplifier and the transpose folding

Since 087e1a5960, nothing re-orients a scatter whose update window dims are
major, so each scatter index writes a strided window (strided atomics,
measured 18x slower than coalesced on an RTX 2070). Since 8a228780ee, the
unconditional transpose-into-scatter fold recreates the strided form from
coalesced-scatter-plus-transpose patterns, defeating workarounds.

Add ScatterSimplifier::WriteRunLength (contiguous write-run length under
default layouts), use it to gate TryFoldTransposeIntoScatter, and add a
reorder_operand_dims_for_coalescing option to ScatterSimplifier (enabled on
the GPU pipelines) that restores the pre-087e1a5960 operand permutation
when it grows the write run and the written volume is large enough to pay
for the transposes. Already-coalesced, variadic, batched, and small-update
scatters keep the current no-transpose behavior.

On the issue's vmapped segment_sum repro (RTX 2070), execution goes from
85-90 ms (strided f32[1024,12123] scatter) to 5.4-5.6 ms (coalesced
f32[12123,1024] scatter plus a restoring transpose).

Fixes #47203.

Merging this change closes #47649

PiperOrigin-RevId: 990896312
…ds in bzl

Imported from GitHub PR openxla/xla#49767

📝 Summary of Changes
Restrict rocm rbe builds to use only env variables passed by the config

🎯 Justification
To achieve fully reproducible builds and better cache hits
we would need a control over the env variables used during the build
hence we restrict any externally set env variables.

🚀 Kind of Contribution
Please remove what does not apply: ♻️ Cleanup,

📊 Benchmark (for Performance Improvements)
Not relevant

🧪 Unit Tests:
CI

🧪 Execution Tests:
CI

Copybara import of the project:

--
93e63d21e10002a3547a67deb4851bab0ef9d710 by Alexandros Theodoridis <atheodor@amd.com>:

Restrict usage of system env variables for rbe builds in bzl

Merging this change closes #49767

PiperOrigin-RevId: 990899121
On load, `GpuExecutable` eagerly constructs `ModuleAnnotations` for both XProf and NVTX. When no NVTX profiler is attached (`DefaultProfilerDomain() == nullptr`), `RegisterString` returns a null handle and discards its input.

Guard the NVTX-only work behind `DefaultProfilerDomain() != nullptr`:
- Module-level stack-frame prefix extraction (`GetLongestSourceLocationPrefix`).
- Per-instruction `Basic` payload formatting (`InstructionAsString`, `FormatSourceLocations`, and `CalledInstructionsAsString`).

PiperOrigin-RevId: 990900227
PiperOrigin-RevId: 990918017
@pull pull Bot locked and limited conversation to collaborators Sep 30, 2026
@pull pull Bot added the ⤵️ pull label Sep 30, 2026
@pull
pull Bot merged commit a3165c2 into Cache-Cloud:master Sep 30, 2026
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.