Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
70 commits
Select commit Hold shift + click to select a range
a55459e
Reject malformed empty boxes and box_index in crop-and-resize kernels
vishwakt Jul 23, 2026
b7d372e
Add separator before shape in boxes and box_index rank error messages
vishwakt Jul 23, 2026
4e66734
Add missing strict dep for image_grad_test_base
vishwakt Aug 23, 2026
77637cc
Reject SplitV size_splits overflow instead of aborting
vishwakt Aug 26, 2026
86e8344
Skip the SplitV overflow test when XLA is enabled
vishwakt Aug 26, 2026
b47c91e
Use disable_xla decorator for the SplitV overflow test
vishwakt Aug 26, 2026
be5e051
Merge master into fix-splitv-size-overflow
vishwakt Aug 29, 2026
14f3ab7
Merge master into fix-splitv-size-overflow
vishwakt Sep 11, 2026
2ca8a3a
Stop exporting statically linked LLVM and MLIR symbols from libtensor…
vishwakt Sep 21, 2026
f2c41de
Match LLVM and MLIR by namespace instead of by substring
vishwakt Sep 21, 2026
1b5868f
Validate axes in np.rot90 for bounds, duplicates, and length
kaivalya-cyber Sep 19, 2026
1d3b095
Skip rot90 axes validation for non-sequence axes
kaivalya-cyber Sep 22, 2026
c64ede3
Wrap nonempty_nonscalar_array_shapes to satisfy 80-char limit
kaivalya-cyber Sep 22, 2026
779544a
Validate filter element count against 32-bit limit in GPU conv transf…
RohithPariki Sep 23, 2026
b21b554
Address PR feedback: Move bounds check before allocation, use direct …
RohithPariki Sep 24, 2026
9b06f31
Merge remote-tracking branch 'origin/master' into fix/np-rot90-axes
kaivalya-cyber Sep 25, 2026
8d78e08
Hide only the LLVM C API from libtensorflow_framework
vishwakt Sep 25, 2026
38dcd10
Pass check_dtypes=False to assertAllClose in testRot90InvalidAxes
kaivalya-cyber Sep 25, 2026
10a810f
Keep the LLVM target initializers exported
vishwakt Sep 25, 2026
3122d28
Drop the empty-input early return in ParseAndCheckBoxSizes
vishwakt Sep 26, 2026
668a864
Format the crop-and-resize changes
vishwakt Sep 26, 2026
7113623
Validate axes in experimental.numpy transpose
kaivalya-cyber Sep 27, 2026
2d8688d
Address review: allocation-free duplicate check, fix mixed-type bypas…
kaivalya-cyber Sep 28, 2026
0500793
Address silent early return review: revert TF_PREDICT_FALSE guards, a…
RohithPariki Sep 28, 2026
d2eec6a
Explicitly include <limits> for IWYU and fix formatting
RohithPariki Sep 28, 2026
c4bec43
Reject negative split sizes before summing them in SplitV
vishwakt Sep 28, 2026
52cf0df
Check the SplitV input size against Tlen, and drop the redundant loop
vishwakt Sep 28, 2026
c28d047
Correct log fragment extraction and main function call
MaddipatlaChetan24 Sep 29, 2026
ff78dee
Update ci/official/utilities/extract_resultstore_links.py
MaddipatlaChetan24 Sep 29, 2026
19cd995
Address review: length check outside rank guard, allocation-free chec…
kaivalya-cyber Sep 29, 2026
ded5b26
Fix scalar transpose: convert axes with explicit int32 dtype so Tperm…
kaivalya-cyber Sep 29, 2026
b58dac4
Consolidate rot90 green-path checks under a single isinstance block
kaivalya-cyber Sep 29, 2026
9092f25
Drop redundant post-loop duplicate check; in-loop bitmask rejection i…
kaivalya-cyber Sep 29, 2026
63a87fb
Remove unused seen_int_axes counter
kaivalya-cyber Sep 29, 2026
73cd721
Convert rot90 axes to Python int before arithmetic to avoid unsigned …
kaivalya-cyber Sep 29, 2026
8e5c138
Fix NameError in rot90 unsigned-axis tests: use onp.uint32 instead of…
kaivalya-cyber Sep 30, 2026
eb512d3
Merge master into fix/np-rot90-axes
kaivalya-cyber Sep 30, 2026
398deb9
Use plain abs() instead of builtins.abs in rot90 duplicate check
kaivalya-cyber Sep 30, 2026
f83995b
Merge master into fix/np-rot90-axes
kaivalya-cyber Sep 30, 2026
6dedc2c
Disable libnvjitlink disk caching via the -no-cache option.
beckerhe Sep 30, 2026
b1e0513
[IFRT Proxy] Avoid client error masking in client disconnection
hyeontaek Sep 30, 2026
3a6e837
Merge pull request #127683 from kaivalya-cyber:fix/np-rot90-axes
tensorflower-gardener Sep 30, 2026
4b89a61
[XLA:GPU] fix number of symbols when simplifying tiles
metaflow Sep 30, 2026
863f92a
Automated Code Change
tensorflower-gardener Sep 30, 2026
4e9a743
[XLA:GPU] Skip collective symmetric buffer tests on non-Hopper archit…
sohaibiftikhar Sep 30, 2026
29b9c9f
[XLA:GPU] Add PerDeviceState container for per-device runtime state.
sohaibiftikhar Sep 30, 2026
e1d06e3
Merge pull request #127780 from vishwakt:fix-framework-llvm-symbol-vi…
tensorflower-gardener Sep 30, 2026
9402391
Merge pull request #128227 from MaddipatlaChetan24:patch-43
tensorflower-gardener Sep 30, 2026
9d72633
Merge pull request #128143 from kaivalya-cyber:fix/np-transpose-axes-…
tensorflower-gardener Sep 30, 2026
40bfb8d
Automated Code Change
tensorflower-gardener Sep 30, 2026
c3f7240
[XLA:GPU] Run the cuDNN frontend graph warmup on a private stream to …
derdrdirk Sep 30, 2026
2816dca
PR #49717: Export control_predecessors() to python
dlcompilers-infra-bot Sep 30, 2026
bbdf033
PR #49741: [ROCm] Remove dead ROCm version checks from GemmRewriter
mminutoli Sep 30, 2026
17bf4fb
[XLA:GPU] Add REDUCE_SCATTER to xla_gpu_experimental_use_collective_k…
PatriosTheGreat Sep 30, 2026
0d920ad
PR #49300: [XLA:GPU] Inline scan computation to_apply calls
akhilgoe Sep 30, 2026
e1d6d57
[XLA:GPU] tighter constraint checks in indexing map
metaflow Sep 30, 2026
25acadd
[NFC] Compute GPU instruction annotation titles and metadata once per…
EusebioDM Sep 30, 2026
8ebd54a
Merge pull request #123864 from vishwakt:fix-crop-and-resize-empty-crash
tensorflower-gardener Sep 30, 2026
480682d
Merge pull request #127968 from RohithPariki:fix/87457-gpu-conv-filte…
tensorflower-gardener Sep 30, 2026
c443ce9
Merge pull request #126161 from vishwakt:fix-splitv-size-overflow
tensorflower-gardener Sep 30, 2026
9fe191d
[XLA:GPU] Support disabled VMM API case.
PatriosTheGreat Sep 30, 2026
645e3ca
[XLA:GPU]: Move memcpy inside all-gather
sohaibiftikhar Sep 30, 2026
5a1f0fd
[StableHLO] Fix windows link error in evalRunParallel
Sep 30, 2026
270a245
PR #49719: [tsl:concurrency] Deduplicate AsyncValue type ids by type …
ezhulenev Sep 30, 2026
b02620e
PR #47649: [XLA:GPU] Keep scatter window writes coalesced in ScatterS…
kodlan Sep 30, 2026
717f958
PR #49767: [ROCm] Restrict usage of system env variables for rbe buil…
alekstheod Sep 30, 2026
54a6c3a
Skip NVTX-only annotation work when no profiler domain is attached.
EusebioDM Sep 30, 2026
4706a67
Removing deprecated macros in XLA:GPU E2E tests and autotuner
Moerafaat Sep 30, 2026
e1e8f86
Migrate deprecated TF assertion macros in XLA:GPU codegen and Triton
Moerafaat Sep 30, 2026
a3165c2
Automated Code Change
tensorflower-gardener Sep 30, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 1 addition & 2 deletions ci/official/utilities/extract_resultstore_links.py
Original file line number Diff line number Diff line change
Expand Up @@ -114,8 +114,7 @@ def parse_log(file_path: str,
else:
tests_failed = re.search(TESTS_FAILED_RE, backtrack_line)
if build_failed or tests_failed:
log_fragment = '\n'.join(
log_lines[max(k - 20, 0):min(end_line + 1, len(log_lines) - 1)])
log_fragment = '\n'.join(log_lines[max(k - 20, 0) : end_line + 1])
lines['log_fragment'] = log_fragment
lines['status'] = (InvokeStatus.build_failed if build_failed
else InvokeStatus.tests_failed)
Expand Down
9 changes: 9 additions & 0 deletions tensorflow/core/kernels/conv_grad_filter_ops_3d.cc
Original file line number Diff line number Diff line change
Expand Up @@ -17,6 +17,7 @@ limitations under the License.
#define EIGEN_USE_THREADS

#include <algorithm>
#include <limits>
#include <string>
#include <utility>
#include <vector>
Expand Down Expand Up @@ -880,6 +881,14 @@ void LaunchConvBackpropFilterOpImpl(
: TensorShape({filter_shape.dim_size(4), dims.filter_size(0),
dims.filter_size(1), dims.filter_size(2),
filter_shape.dim_size(3)});
// Validate filter element count before allocation to prevent OOM on invalid
// inputs. GPU transformation uses 32-bit indexing via To32Bit().
OP_REQUIRES(
context,
filter_backprop->NumElements() <= std::numeric_limits<int32>::max(),
errors::InvalidArgument("Filter tensor num elements (",
filter_backprop->NumElements(),
") exceeds 32-bit limit for GPU transformation"));
OP_REQUIRES_OK(context,
context->allocate_temp(DataTypeToEnum<T>::value, dst_shape,
&pre_transformed_filter_backprop));
Expand Down
9 changes: 9 additions & 0 deletions tensorflow/core/kernels/conv_grad_filter_ops_launcher.cc
Original file line number Diff line number Diff line change
Expand Up @@ -17,6 +17,7 @@ limitations under the License.
#define EIGEN_USE_THREADS

#include <algorithm>
#include <limits>
#include <utility>
#include <vector>

Expand Down Expand Up @@ -399,6 +400,14 @@ void LaunchConv2DBackpropFilterOpImpl(
// We compute filter backprop into temporary tensor, and then convert it to
// the HWIO data format at the end.

// Validate filter element count before allocation to prevent OOM on invalid
// inputs. GPU transformation uses 32-bit indexing via To32Bit().
OP_REQUIRES(
ctx, filter_backprop->NumElements() <= std::numeric_limits<int32>::max(),
errors::InvalidArgument("Filter tensor num elements (",
filter_backprop->NumElements(),
") exceeds 32-bit limit for GPU transformation"));

Tensor pre_transformed_filter_backprop;
OP_REQUIRES_OK(
ctx,
Expand Down
6 changes: 6 additions & 0 deletions tensorflow/core/kernels/conv_grad_input_ops.cc
Original file line number Diff line number Diff line change
Expand Up @@ -17,6 +17,7 @@ limitations under the License.

#include "tensorflow/core/kernels/conv_grad_input_ops.h"

#include <limits>
#include <utility>

#include "tensorflow/core/profiler/lib/scoped_annotation.h"
Expand Down Expand Up @@ -291,6 +292,11 @@ void LaunchConv2DBackpropInputOpGpuImpl(
: TensorShape({filter.dim_size(3), filter.dim_size(0),
filter.dim_size(1), filter.dim_size(2)});

if (filter.NumElements() > std::numeric_limits<int32>::max()) {
return errors::InvalidArgument(
"Filter tensor num elements (", filter.NumElements(),
") exceeds 32-bit limit for GPU transformation");
}
TF_RETURN_IF_ERROR(ctx->allocate_temp(DataTypeToEnum<T>::value, dst_shape,
&transformed_filter));
functor::TransformFilter<GPUDevice, T, int, 4>()(
Expand Down
7 changes: 7 additions & 0 deletions tensorflow/core/kernels/conv_grad_input_ops_3d.cc
Original file line number Diff line number Diff line change
Expand Up @@ -17,6 +17,7 @@ limitations under the License.
#define EIGEN_USE_THREADS

#include <algorithm>
#include <limits>
#include <string>
#include <utility>
#include <vector>
Expand Down Expand Up @@ -873,6 +874,12 @@ void LaunchConvBackpropInputOpImpl(
: TensorShape({filter_shape.dim_size(4), dims.filter_size(0),
dims.filter_size(1), dims.filter_size(2),
filter_shape.dim_size(3)});
OP_REQUIRES(context,
filter.NumElements() <= std::numeric_limits<int32>::max(),
errors::InvalidArgument(
"Filter tensor num elements (", filter.NumElements(),
") exceeds 32-bit limit for GPU transformation"));

OP_REQUIRES_OK(context,
context->allocate_temp(DataTypeToEnum<T>::value, dst_shape,
&transformed_filter));
Expand Down
7 changes: 7 additions & 0 deletions tensorflow/core/kernels/conv_ops_fused_impl.h
Original file line number Diff line number Diff line change
Expand Up @@ -38,6 +38,7 @@ limitations under the License.
#define EIGEN_USE_GPU
#endif // GOOGLE_CUDA

#include <limits>
#include <string>
#include <type_traits>
#include <utility>
Expand Down Expand Up @@ -558,8 +559,14 @@ struct LaunchFusedConv2DOp<GPUDevice, T> {
: TensorShape({filter.dim_size(3), filter.dim_size(0),
filter.dim_size(1), filter.dim_size(2)});

if (filter.NumElements() > std::numeric_limits<int32>::max()) {
return errors::InvalidArgument(
"Filter tensor num elements (", filter.NumElements(),
") exceeds 32-bit limit for GPU transformation");
}
TF_RETURN_IF_ERROR(context->allocate_temp(
DataTypeToEnum<T>::value, dst_shape, &transformed_filter));

functor::TransformFilter<GPUDevice, T, int, 4>()(
context->eigen_device<GPUDevice>(), dst_format,
To32Bit(filter.tensor<T, 4>()),
Expand Down
6 changes: 6 additions & 0 deletions tensorflow/core/kernels/conv_ops_impl.h
Original file line number Diff line number Diff line change
Expand Up @@ -1133,6 +1133,12 @@ void LaunchConvOpImpl(OpKernelContext* context, bool cudnn_use_autotune,
}
}
TensorShape dst_shape(dst_shape_vec);
OP_REQUIRES(context,
filter.NumElements() <= std::numeric_limits<int32>::max(),
errors::InvalidArgument(
"Filter tensor num elements (", filter.NumElements(),
") exceeds 32-bit limit for GPU transformation"));

OP_REQUIRES_OK(context,
context->allocate_temp(DataTypeToEnum<T>::value, dst_shape,
&transformed_filter));
Expand Down
2 changes: 1 addition & 1 deletion tensorflow/core/kernels/deserialize_sparse_string_op.cc
Original file line number Diff line number Diff line change
Expand Up @@ -203,7 +203,7 @@ class DeserializeSparseOp : public OpKernel {
target_shape.vec<int64_t>()(i) = serialized_sparse.shape().dim_size(i);
}
for (int i = 0; i < output.dims() - 1; ++i) {
target_shape.vec<int64_t>()(i + ndims - 1) = output.shape().data()[i + 1];
target_shape.vec<int64_t>()(i + ndims - 1) = output.shape()[i + 1];
}

ReshapeSparseTensor<CPUDevice>(context, output.indices(), input_shape,
Expand Down
19 changes: 8 additions & 11 deletions tensorflow/core/kernels/image/crop_and_resize_op.cc
Original file line number Diff line number Diff line change
Expand Up @@ -55,24 +55,21 @@ using Callback = std::function<void()>;
static inline absl::Status ParseAndCheckBoxSizes(const Tensor& boxes,
const Tensor& box_index,
int* num_boxes) {
if (boxes.NumElements() == 0 && box_index.NumElements() == 0) {
*num_boxes = 0;
return absl::OkStatus();
}
// The shape of 'boxes' is [num_boxes, 4].
// The shape of 'boxes' is [num_boxes, 4] and the shape of 'box_index' is
// [num_boxes]. The ranks must be validated even when both tensors are
// empty, since the kernels later access them as rank-2 and rank-1 tensors.
if (boxes.dims() != 2) {
return absl::InvalidArgumentError(
absl::StrCat("boxes must be 2-D", boxes.shape().DebugString()));
absl::StrCat("boxes must be 2-D, got ", boxes.shape().DebugString()));
}
if (box_index.dims() != 1) {
return absl::InvalidArgumentError(absl::StrCat(
"box_index must be 1-D, got ", box_index.shape().DebugString()));
}
*num_boxes = boxes.dim_size(0);
if (boxes.dim_size(1) != 4) {
return absl::InvalidArgumentError("boxes must have 4 columns");
}
// The shape of 'box_index' is [num_boxes].
if (box_index.dims() != 1) {
return absl::InvalidArgumentError(
absl::StrCat("box_index must be 1-D", box_index.shape().DebugString()));
}
if (box_index.dim_size(0) != *num_boxes) {
return absl::InvalidArgumentError("box_index has incompatible shape");
}
Expand Down
36 changes: 28 additions & 8 deletions tensorflow/core/kernels/split_v_op.cc
Original file line number Diff line number Diff line change
Expand Up @@ -27,6 +27,7 @@ limitations under the License.
#define PLUGGABLE_DEVICE_SUPPORTED_MACOS 1
#endif

#include <limits>
#include <numeric>

#include "unsupported/Eigen/CXX11/Tensor" // from @eigen_archive
Expand Down Expand Up @@ -100,7 +101,15 @@ class SplitVOpBase : public OpKernel {
"-input rank(-", input.dims(), ") <= split_dim < input rank (",
input.dims(), "), but got ", split_dim_orig)));

Tlen input_size_split_dim = input_shape.dim_size(split_dim);
// Check that the input size fits in Tlen before converting it. Otherwise
// an int32 or int8 Tlen truncates it, and split sizes that sum to the
// truncated size silently drop the rest of the input.
const int64_t actual_input_size = input_shape.dim_size(split_dim);
OP_REQUIRES(context, actual_input_size <= std::numeric_limits<Tlen>::max(),
errors::InvalidArgument(
"Input size along split_dim must be <= max(Tlen). Got: ",
actual_input_size));
Tlen input_size_split_dim = static_cast<Tlen>(actual_input_size);

// Special case 1: num_split == 1. Nothing to do.
if (num_split == 1) {
Expand Down Expand Up @@ -128,6 +137,24 @@ class SplitVOpBase : public OpKernel {
"input."));
neg_one_dim = d;
} else {
// Reject a negative size before summing it, so that
// 0 <= determined_size <= input_size_split_dim below. Otherwise a
// large negative size, such as the minimum of Tlen next to a -1, makes
// `input_size_split_dim - determined_size` overflow.
OP_REQUIRES(context, size >= 0,
errors::InvalidArgument("Split size at index ", d,
" must be >= 0. Got: ", size));
// Accumulate with an explicit overflow guard. `determined_size += size`
// wraps for large `size_splits`, and a wrapped total can equal
// `input_size_split_dim` and pass the check below, letting the aligned
// slicing path compute an endpoint that reaches a fatal `Tensor::Slice`
// invariant. Rejecting overflow here also keeps that later path safe,
// since the accepted total then bounds every partial sum.
OP_REQUIRES(context,
determined_size <= std::numeric_limits<Tlen>::max() - size,
errors::InvalidArgument(
"Sum of size_splits overflows the index type at index ",
d, "."));
determined_size += size;
}
}
Expand All @@ -147,13 +174,6 @@ class SplitVOpBase : public OpKernel {
(*split_sizes_vec)[neg_one_dim] = input_size_split_dim - determined_size;
}

for (int i = 0; i < split_sizes_vec->size(); ++i) {
const Tlen& split_size = (*split_sizes_vec)[i];
OP_REQUIRES(context, split_size >= Tlen(0),
errors::InvalidArgument("Split size at index ", i,
" must be >= 0. Got: ", split_size));
}

// Special case 2: split along the 1st dimension. The requirements are that
// either we are splitting the outer dimension of two or more such that
// every outer subpart is aligned or that the split sizes mean that they are
Expand Down
2 changes: 1 addition & 1 deletion tensorflow/core/lib/db/sqlite_test.cc
Original file line number Diff line number Diff line change
Expand Up @@ -170,7 +170,7 @@ TEST_F(SqliteTest, UnsafeColumn) {
stmt = db_->PrepareOrDie("SELECT b FROM T ORDER BY a");
TF_ASSERT_OK(stmt.Step(&is_done_));
absl::string_view p = stmt.ColumnStringUnsafe(0);
EXPECT_EQ('h', *p.data());
EXPECT_EQ('h', p[0]);
TF_ASSERT_OK(stmt.Step(&is_done_));
// This will actually happen, but it's not safe to test this behavior.
// EXPECT_EQ('t', *p.data());
Expand Down
57 changes: 57 additions & 0 deletions tensorflow/python/kernel_tests/array_ops/split_op_test.py
Original file line number Diff line number Diff line change
Expand Up @@ -120,6 +120,63 @@ def testExplicitNum(self):
self.assertAllEqual(r[1], value[2:4])
self.assertAllEqual(r[2], value[4:])

@test_util.run_in_graph_and_eager_modes
@test_util.disable_xla(
"XLA shape inference rejects the reshape to an INT64_MAX dimension, so "
"the test cannot reach the SplitV kernel under XLA"
)
def testSizeSplitsOverflowRaises(self):
# Regression test for GitHub issue 126126. The cumulative sum of
# size_splits was computed with unchecked signed addition, so a wrapped
# total could equal the input dimension, pass validation, and reach a
# fatal `Tensor::Slice` invariant in the aligned slicing path. It must
# raise instead.
i64_max = (1 << 63) - 1
i32_max = (1 << 31) - 1
for input_size, size_splits, dtype, message in (
(i64_max, [i64_max, i64_max, i64_max, 2], dtypes.int64, "overflow"),
# A -1 does not keep the other sizes from overflowing.
(i64_max, [-1, i64_max, i64_max, 2], dtypes.int64, "overflow"),
# int32 sizes overflow at their own width in the kernel. In graph
# mode the shape function, which sums in int64, rejects the mismatch
# before the kernel runs.
(5, [i32_max, i32_max, 5], dtypes.int32, "overflow|can't split axis"),
):
with self.subTest(size_splits=size_splits, dtype=dtype.name):
value = array_ops.reshape(
constant_op.constant([], dtype=dtypes.float32),
constant_op.constant([input_size, 0], dtype=dtypes.int64),
)
with self.assertRaisesRegex(
(ValueError, errors_impl.InvalidArgumentError), message
):
self.evaluate(
array_ops.split(
value, constant_op.constant(size_splits, dtype=dtype), axis=0
)
)

@test_util.run_in_graph_and_eager_modes
@test_util.disable_xla("Checks the SplitV kernel, which XLA replaces")
def testInputSizeAboveSizeSplitsTypeRaises(self):
# With int32 size_splits, an input size above INT32_MAX was truncated to
# int32, so [1, 2] matched an input of size 2**32 + 3 and the split
# silently dropped the rest of the input. In graph mode the shape
# function, which compares in int64, rejects the mismatch first.
value = array_ops.reshape(
constant_op.constant([], dtype=dtypes.float32),
constant_op.constant([(1 << 32) + 3, 0], dtype=dtypes.int64),
)
with self.assertRaisesRegex(
(ValueError, errors_impl.InvalidArgumentError),
r"must be <= max\(Tlen\)|can't split axis",
):
self.evaluate(
array_ops.split(
value, constant_op.constant([1, 2], dtype=dtypes.int32), axis=0
)
)

@test_util.run_in_graph_and_eager_modes
def testListOfScalarTensors(self):
a = math_ops.cast(5, dtypes.int32)
Expand Down
Loading
Loading