[pull] master from tensorflow:master - #8861
Merged
Merged
Conversation
ParseAndCheckBoxSizes returned early with num_boxes=0 when both boxes and box_index were empty, skipping all rank validation. A rank-1 empty boxes tensor (shape [0] instead of [0, 4]) or a rank-2 empty box_index then reached Tensor::tensor<T, NDIMS>() with the wrong rank, aborting the process with a fatal CHECK failure instead of raising a catchable InvalidArgumentError. This crashed CropAndResizeGradImage and CropAndResizeGradBoxes. Validate the ranks of boxes and box_index before the empty early return. Well-formed empty inputs (boxes of shape [0, 4] with box_index of shape [0]) keep working, and empty tensors of the correct rank are still accepted. Fixes #123397
The messages previously rendered as "boxes must be 2-D[0]". They now read "boxes must be 2-D, got [0]".
Syncs the branch with upstream master so the copybara import runs against current revisions of the files this change touches. The array_ops kernel_tests BUILD file changed upstream since this branch was cut, in a different target from the one this change edits, so the merge is clean.
…flow_framework
libtensorflow_framework.so is linked with tf_framework_version_script.lds,
which read:
tensorflow {
global:
*;
};
That declares every symbol global and has no local stanza, so the library
exports its statically linked third-party code as well as its own. Among
those are the LLVM symbols, which then reach the process-wide resolution
scope and collide with any other LLVM loaded into the same process. In the
reported crash an LLVM 15 brought in by llvmlite and numba resolved
llvm::raw_svector_ostream::write_impl against TensorFlow's LLVM 18 and
died on the mismatched layout.
Add a local stanza hiding *llvm* and *mlir*, and add *llvm* to
tf_private_symbols.lds so macOS, which already hides *mlir* through
-unexported_symbols_list, covers the same set.
The change is deliberately narrow. TensorFlow's own symbols stay exported,
because custom op libraries loaded through tf.load_op_library resolve
REGISTER_OP and REGISTER_KERNEL_BUILDER against this library, which is why
the script exports everything today.
Verified with lld, the linker the Linux builds use via -fuse-ld=lld, on a
shared object built from the shipped script: llvm and mlir symbols are
dropped from .dynsym while the tensorflow namespace and the TF_ C API stay
exported. Linking also succeeds, with and without --undefined-version, when
no symbol matches either pattern, so configurations built without LLVM are
unaffected.
Fixes #104038
Per review, the substring patterns were both too broad and too narrow. Too broad: *mlir* also hides TensorFlow's own symbols whose names happen to contain the string, such as tensorflow::tf_xla_test_use_mlir, which would become local and fail to resolve for anything linking against it. Too narrow: *llvm* is lowercase and never matched the LLVM C API at all, whose symbols are spelled LLVM*. Those kept leaking into the dynamic symbol table, which is the same collision class the change is meant to close. Match the namespaces through extern "C++" and add an explicit LLVM* prefix for the C API. Verified with lld on the shipped script: llvm::, mlir:: and LLVM* symbols are dropped while tensorflow::, including tf_xla_test_use_mlir, and the TF_ C API stay exported. Linking still succeeds when no pattern matches. Also adds *LLVM* to the macOS list, which had the same case gap next to its existing *mlir* entry.
…ormations Filter transformation kernels on GPU (TransformFilter and ReverseTransformFilter) use 32-bit indexing via To32Bit. When filter element count exceeds INT32_MAX, signed integer overflow caused fatal process crashes in gpu_launch_config.h (CHECK_GE work_element_count >= 0) and out-of-bounds illegal memory access in ShuffleInTensor3Simple. This change adds FastBoundsCheck upfront validation in conv_ops_impl.h, conv_ops_fused_impl.h, conv_grad_input_ops.cc, and conv_grad_input_ops_3d.cc, and guards non-positive output sizes in conv_2d_gpu.h. Fixes #87457 Fixes #87438 Fixes #87454
…comparison, and remove memory-heavy python test
# Conflicts: # tensorflow/python/ops/numpy_ops/tests/np_test.py
Hiding the llvm:: and mlir:: C++ namespaces broke the build, because libtensorflow_cc links against those symbols from libtensorflow_framework. It imports 5,351 of them in the 2.21.0 Linux wheel, and 3,816 LLVM symbols in a macOS nightly, where the *llvm* and *LLVM* patterns would have hidden them just the same. The LLVM C API is different: no other TensorFlow library imports any of it, on either platform, so hiding it cannot break a link. Keep only that. On Linux the pattern is LLVM*. On macOS it is _LLVM*, which carries the Mach-O leading underscore and, unlike *LLVM*, does not also match C++ names that mention LLVM types.
In the CUDA builds, libtensorflow_cc links tfcompile's InitializeTargets(), which calls the LLVMInitialize* functions for each LLVM target and imports them from libtensorflow_framework, so hiding all of LLVM* broke that link. They are the only LLVM C API functions that TensorFlow or XLA code calls. Export them again and keep hiding the rest of the C API.
With the rank checks first, the early return for empty boxes and box_index no longer protects anything: well-formed empty inputs pass the column and row checks on their own. It only let malformed empty boxes through, such as shape [0, 5] or [2, 0], which every crop-and-resize kernel then accepted. Remove it, and extend the test to cover the boxes gradient with a rank-2 box_ind, empty boxes with the wrong number of columns, and well-formed empty inputs to the boxes gradient.
Run clang-format on ParseAndCheckBoxSizes and pyink on testMalformedEmptyBoxesRaisesError, the code this change touches. No functional change.
…s, add rank-0/1 tests, quote style
…dd OP_REQUIRES on callers Per maintainer review (dmiltr3): - Revert if (TF_PREDICT_FALSE(out.size() <= 0)) return; from TransformFilter and ReverseTransformFilter in conv_2d_gpu.h. These caused silent skips with uninitialized output buffers being passed to cuDNN. - Add upfront OP_REQUIRES bounds checks in conv_grad_filter_ops_launcher.cc and conv_grad_filter_ops_3d.cc before allocate_temp, so invalid filters produce a loud InvalidArgument error rather than silent data corruption. Fixes: #115734
Per maintainer review (dmiltr3): - Add explicit #include <limits> in conv_grad_filter_ops_3d.cc, conv_grad_filter_ops_launcher.cc, conv_grad_input_ops.cc, conv_grad_input_ops_3d.cc, and conv_ops_fused_impl.h for IWYU compliance. - Remove stray newline before kernel_shape_util.h include in conv_grad_input_ops_3d.cc.
The overflow guard kept determined_size itself in range, but a large negative size next to a -1, such as [-1, INT64_MIN], still made input_size_split_dim - determined_size overflow when computing the -1 size. Reject a negative size before summing it, as the later check already did for the error message, so that 0 <= determined_size <= input_size_split_dim and only the upper bound needs a guard. Extend testSizeSplitsOverflowRaises to int32 split sizes and to an overflow next to a -1, accept ValueError for graph mode, and call array_ops.split instead of the generated op, which drops the array_ops_gen dependency this change had added.
input_shape.dim_size(split_dim) was converted to Tlen implicitly, so with int32 or int8 size_splits an input size above the maximum of Tlen was truncated: int32 sizes [1, 2] matched an input of size 2**32 + 3, and the split silently dropped the rest of the input. Check the size in int64 before converting it. Split sizes are now checked for being non-negative before they are summed, and the -1 size is bounded by the input size, so the later loop that checked every size again can no longer fail; remove it. Also build the overflow error with errors::InvalidArgument, as the rest of the file does. The int32 case of testSizeSplitsOverflowRaises used an input larger than INT32_MAX to reach the kernel in graph mode, which the new check rejects first. Use a small input, where graph mode stops at the shape function instead, and test the truncation separately.
Fix log fragment extraction and ensure main function is called correctly.
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
…ks with NumPy error precedence, sub-2D tests
PiperOrigin-RevId: 990789026
…ectures. PiperOrigin-RevId: 990801444
Adds `xla::gpu::PerDeviceState<T>`, which combines pre-allocated `VectorStorage` for device ordinals in `[0, num_devices)` with copy-on-write `CowStorage` fallback for ordinals outside `[0, num_devices)`. PiperOrigin-RevId: 990801463
…sibility PiperOrigin-RevId: 990807826
PiperOrigin-RevId: 990809173
…validation PiperOrigin-RevId: 990811849
PiperOrigin-RevId: 990820587
…avoid race conditions when sharing the stream between multiple threads. PiperOrigin-RevId: 990824347
Imported from GitHub PR openxla/xla#49717 📝 Summary of Changes Export control_predecessors() to python 🎯 Justification JAX has an experimental API to allow python to modify the Thunk scheduling. To build valid one, we need to know the control depedence. 🚀 Kind of Contribution ✨ New Feature, 🧪 Tests 📊 Benchmark (for Performance Improvements) No speed-up expected by this exposure. 🧪 Unit Tests: The new Python API have been tested. 🧪 Execution Tests: xla/python/xla_hlo_test.py::TestHloModule.testHloInstructionControlPredecessors Copybara import of the project: -- 9c42acd7cc52413d38d851e914f65af1f5f21135 by Frederic Bastien <fbastien@nvidia.com>: Export control_predecessors() to python Merging this change closes #49717 PiperOrigin-RevId: 990831343
Imported from GitHub PR openxla/xla#49741 ## Summary - Remove all `toolkit_version_ < {6,x,0}` runtime checks in `gemm_rewriter.cc` — they are dead code since `rocm_config.h.tpl` requires ROCm >= 7.1 at compile time (`#error` if `TF_ROCM_VERSION < 70100`). - Simplify the `toolkit_version_ >= {7,0,0}` guard on Swish matching to plain `if (is_rocm)`, since it is always true. - Remove the now-callerless `TurnF8DotWithUnsupportedOutputTypeIntoF32()` helper. - Clean up corresponding dead branches in `gemm_rewriter_fp8_test.cc` (ROCm < 6.0 skip, and five if/else blocks selecting CHECK patterns for ROCm < 6.2). Copybara import of the project: -- 996cd22d95e984681c58ad0187e75c1fb698c666 by Marco Minutoli <marco.minutoli@amd.com>: [ROCm] Remove dead ROCm version checks from GemmRewriter rocm_config.h.tpl requires ROCm >= 7.1 (#error if < 70100), making all runtime checks against ROCm 6.x unreachable. Remove the dead guards: - toolkit_version_ < {6,2,0} output-type workaround at FP8 dot rewrite - toolkit_version_ < {6,0,0} FP8 availability check (has_fp8_support remains) - toolkit_version_ < {6,2,0} output-type restriction in CreateF8CustomCall - toolkit_version_ >= {7,0,0} always-true guard on Swish matching (simplified) - TurnF8DotWithUnsupportedOutputTypeIntoF32() helper (no remaining callers) - Corresponding dead branches in gemm_rewriter_fp8_test.cc Merging this change closes #49741 PiperOrigin-RevId: 990831715
…ernels. PiperOrigin-RevId: 990832583
Imported from GitHub PR openxla/xla#49300 Scan rewrite lowering converts scan operations into call-based expressions. However, the GPU kernel emitter cannot compute indexing maps for Call operations, causing compilation failures for GPU kernels containing scan. Exact error: ``` E0000 00:00:1789415358.746826 1768463 indexing_analysis.cc:1748] ComputeOutputToInputIndexing is not implemented for opcode call F0000 00:00:1789415358.746906 1768463 computation_partitioner.cc:269] Check failed: operand_maps.size() == 1 (0 vs. 1) ``` This PR inlines the scan computation calls, resolving the indexing analysis failure. CUDA and ROCm backends avoid this issue because they run CallInliner in their device-specific Convolution Canonicalization optimization pipelines that execute before the kernel code generation. This issue came to light after openxla/xla@33d1a58 stopped converting scan operations to custom calls. Copybara import of the project: -- ac1221a33121f1a07207d0aad698d85f43ae3bf0 by Akhil Goel <akhil.goel@intel.com>: Inline scan calls -- 76a47249786fc94a1b18854444e6b5ba94e72f7b by Akhil Goel <akhil.goel@intel.com>: Add CallInliner pass Merging this change closes #49300 PiperOrigin-RevId: 990834717
1. when constructed from a map of interval constraints we have not checked if any of the constraints is unsatisfiable - now we sent it trough ctor that calls AddConstraint and handles that for us. 2. added a check if const constraint is satisfiable, e.g. we can get something like `4 in [0, 3]` that should make the whole construct invalid. PiperOrigin-RevId: 990839235
… instruction. `InstructionAnnotation` and `GetInstructionAnnotationMetadata` where being computed twice for the same values. PiperOrigin-RevId: 990844885
PiperOrigin-RevId: 990862199
…r-transform-overflow PiperOrigin-RevId: 990863209
PiperOrigin-RevId: 990863212
This is needed to support fractional vGPUs (see openxla/xla#49252). Collective fusions are disabled in this case and all collectives should fallback on CPU initiated NCCL without RMA access. PiperOrigin-RevId: 990867700
Replace host-launched D2D copy to scratch with in-kernel copies to scratch. Each block copies T/R of a tile to the scratch where T is the size of a tile and R is the number of ranks (world_size). Asymmetric block barriers, ensure that consumer blocks wait for all producer blocks to finish before emitting the output copy. PiperOrigin-RevId: 990876953
Implemented the fix suggested here: openxla/stablehlo#2975. PiperOrigin-RevId: 990878885
…name across DSOs Imported from GitHub PR openxla/xla#49719 `AsyncValue` assigns each payload type a 16-bit type id by appending to a process-global `TypeInfoTable` and caching the resulting index in `GetTypeId<T>()`'s function-local static. This is unsafe once XLA is split across dynamically-linked libraries: XLA is built with `-fvisibility=hidden`, which demotes the `GetTypeId<T>()` local static to a per-DSO local. A type registered in more than one DSO therefore appends to the shared table more than once and receives a different id in each DSO, so `IsType<T>()`/`DynCast<T>()` on an `AsyncValue` that crosses a DSO boundary can silently return the wrong answer. Make registration idempotent by type name: `CreateTypeInfoAndReturnTypeIdImpl` now keys a `name->id` map on `typeid(T).name()` and returns the existing id when the same type is registered again, instead of allocating a fresh one. The map lives behind a const-init mutex in the single out-of-line definition of the registration function, so all DSOs that resolve that symbol share one id per type. No caller changes are required. This mirrors MLIR's `TypeID`, whose fallback resolver (r`egisterImplicitTypeID(getTypeName<T>())`) already keys on the type name rather than on registration order, giving a process-stable identity that survives across DSO boundary. Copybara import of the project: -- e94bb1d32f7a06d397a4d33cd4b1a4dd0c07a4e6 by Eugene Zhulenev <ezhulenev@openxla.org>: [tsl:concurrency] Deduplicate AsyncValue type ids by type name across DSOs Merging this change closes #49719 PiperOrigin-RevId: 990883982
…implifier and the transpose folding Imported from GitHub PR openxla/xla#47649 ## 📝 Summary of Changes Two changes that together restore coalesced memory writes for scatters whose update window dims come before the indexed dims (for example a vmapped `segment_sum`, where the batch dim is a window dim and the segment ids index the minor dim): - `ScatterSimplifier` gets a `reorder_operand_dims_for_coalescing` option (enabled in the GPU pipelines only). When a scatter would write strided windows under the default layouts, the operand dims are permuted so that `scatter_dims_to_operand_dims` becomes the identity mapping and the writes become contiguous, with transposes around the scatter restoring the original order. This conditionally restores the canonicalization that 087e1a5960 removed: scatters that are already coalesced keep the current no-transpose behavior, and so do variadic scatters, scatters with batching dims, and scatters whose written volume is small relative to the operand (where the transpose copies would cost more than the coalescing wins). - `TryFoldTransposeIntoScatter` in the algebraic simplifier now checks the same profitability predicate and refuses folds that would make the written windows less contiguous. Without this, the fold both undoes the ScatterSimplifier rewrite above and defeats user-side workarounds (a restoring `.transpose()` gets folded back into the scatter, recreating the strided form). Folds that improve or preserve contiguity still fire. The shared predicate `ScatterSimplifier::WriteRunLength` computes the length of the contiguous run of operand elements each scatter index writes. ## 🎯 Justification Fixes #47203 (cross-post of jax-ml/jax#39959): `segment_sum` under `vmap` regressed ~6x on the reporter's GPU between JAX 0.9.1 and 0.9.2, and its `out_axes=1` workaround regressed further in 0.10.2. Root cause chain: - Since 087e1a5960, nothing in the GPU pipeline re-orients a scatter whose window dims are major, so every scatter index writes a strided window (strided atomics). Isolated on an RTX 2070 with identical indices and updates, only the operand orientation differing: 5.11 ms coalesced vs 91.62 ms strided (18x). - Since 8a228780ee, the unconditional transpose fold recreates the strided form out of the coalesced-scatter-plus-transpose pattern, so no HLO-level workaround survives. With this PR, the issue's repro compiles back to the coalesced scatter plus one cheap restoring transpose: 85-90 ms -> ~5.5 ms on the RTX 2070 (~16x), details in the Benchmark section. ## 🚀 Kind of Contribution ⚡️ Performance Improvement / 🐛 Bug Fix ## 📊 Benchmark Issue repro as a standalone HLO (vmapped segment_sum: scatter-add of f32[1024,75960] updates into 12123 segments, hashed pseudo-random segment ids, `hlo_runner_main_gpu --num_repeats=10`, RTX 2070, sm_75): - before: 85-90 ms per execution, compiled scatter `f32[1024,12123]{1,0}` (strided window writes) - after: 5.4-5.6 ms per execution (~16x), compiled scatter `f32[12123,1024]{1,0}` (coalesced) plus a restoring transpose Isolated scatter kernel with identical indices and updates, only the operand orientation differing (via JAX on the same GPU): 5.11 ms coalesced vs 91.62 ms strided (18x). A standalone `[12123,1024] -> [1024,12123]` transpose costs 0.32 ms, so the inserted transposes are ~2% of the win. ## 🧪 Unit Tests: - `scatter_simplifier_test.cc`: `ReordersOperandDimsForCoalescing`, `ReordersOperandDimsWithInsertedWindowDims` (non-monotonic `scatter_dims_to_operand_dims`), `DoesNotReorderCoalescedScatter`, `DoesNotReorderVariadicScatter`, `DoesNotReorderWhenUpdatesAreSmall`, `DoesNotReorderOperandDimsByDefault`. - `algebraic_simplifier_test.cc`: `FoldTransposeIntoScatter` now uses a profitable example (window dims move to minor); `DoNotFoldTransposeIntoScatterWhenWritesBecomeStrided` covers the refused direction. ## 🧪 Execution Tests: Verified on a single RTX 2070 (sm_75): - `//xla/tests:scatter_test_nvgpu_any` (36 cases) and `//xla/tests:select_and_scatter_test_nvgpu_any` pass. - `//xla/service/gpu:gpu_compiler_test_nvgpu_any` and the GPU scatter emitter lit tests (`add`, `sorted_indices`, `permuted_indices`, `permuted_sorted_indices`) pass. - `run_hlo_module --platform=gpu --reference_platform=interpreter` on a small version of the repro: results match. Copybara import of the project: -- bd08ecb02fa0c0e2b705d15d56bdebfd1ac3551f by Stanislav Bardyuk <sbardyuk@google.com>: [XLA:GPU] Keep scatter window writes coalesced in ScatterSimplifier and the transpose folding Since 087e1a5960, nothing re-orients a scatter whose update window dims are major, so each scatter index writes a strided window (strided atomics, measured 18x slower than coalesced on an RTX 2070). Since 8a228780ee, the unconditional transpose-into-scatter fold recreates the strided form from coalesced-scatter-plus-transpose patterns, defeating workarounds. Add ScatterSimplifier::WriteRunLength (contiguous write-run length under default layouts), use it to gate TryFoldTransposeIntoScatter, and add a reorder_operand_dims_for_coalescing option to ScatterSimplifier (enabled on the GPU pipelines) that restores the pre-087e1a5960 operand permutation when it grows the write run and the written volume is large enough to pay for the transposes. Already-coalesced, variadic, batched, and small-update scatters keep the current no-transpose behavior. On the issue's vmapped segment_sum repro (RTX 2070), execution goes from 85-90 ms (strided f32[1024,12123] scatter) to 5.4-5.6 ms (coalesced f32[12123,1024] scatter plus a restoring transpose). Fixes #47203. Merging this change closes #47649 PiperOrigin-RevId: 990896312
…ds in bzl Imported from GitHub PR openxla/xla#49767 📝 Summary of Changes Restrict rocm rbe builds to use only env variables passed by the config 🎯 Justification To achieve fully reproducible builds and better cache hits we would need a control over the env variables used during the build hence we restrict any externally set env variables. 🚀 Kind of Contribution Please remove what does not apply: ♻️ Cleanup, 📊 Benchmark (for Performance Improvements) Not relevant 🧪 Unit Tests: CI 🧪 Execution Tests: CI Copybara import of the project: -- 93e63d21e10002a3547a67deb4851bab0ef9d710 by Alexandros Theodoridis <atheodor@amd.com>: Restrict usage of system env variables for rbe builds in bzl Merging this change closes #49767 PiperOrigin-RevId: 990899121
On load, `GpuExecutable` eagerly constructs `ModuleAnnotations` for both XProf and NVTX. When no NVTX profiler is attached (`DefaultProfilerDomain() == nullptr`), `RegisterString` returns a null handle and discards its input. Guard the NVTX-only work behind `DefaultProfilerDomain() != nullptr`: - Module-level stack-frame prefix extraction (`GetLongestSourceLocationPrefix`). - Per-instruction `Basic` payload formatting (`InstructionAsString`, `FormatSourceLocations`, and `CalledInstructionsAsString`). PiperOrigin-RevId: 990900227
PiperOrigin-RevId: 990914610
PiperOrigin-RevId: 990917009
PiperOrigin-RevId: 990918017
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to subscribe to this conversation on GitHub.
Already have an account?
Sign in.
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
See Commits and Changes for more details.
Created by
pull[bot] (v2.0.0-alpha.4)
Can you help keep this open source service alive? 💖 Please sponsor : )