Repository navigation
[RFC]: add APIs for array equality to a scalar #985
Description
Activity
- changed the title
[-]Improved API for array equality to a scalar[/-][+][RFC]: add APIs for array equality to a scalar[/+]on Dec 20, 2025 - addedRFCRequest for comments. Feature requests and proposed changes.Request for comments. Feature requests and proposed changes.Needs DiscussionNeeds further discussion.Needs further discussion.
on Dec 20, 2025 Thanks for the ideas @hmaarrfk
However, this code is terribly inefficient for out of memory array.
It can force them to go through the entire memory.
count_nonzero(arr)andany(arr)(andnot any(arr)) can both short-circuit, so should be more efficient thannp.all(arr == 0).Second, I would like to propose an API for "likely not equal", where the strick inequality is not guaranteed.
That doesn't seem feasible to standardize, because (a) it doesn't exist in current array libraries, and (b) it probably doesn't make sense for most libraries (e.g., why would
numpywant to add this?).I could likely check numpy metadata like
stridesand create my own fastpath for this:That doesn't seem worth doing. Serializing a just-allocated array with all zeros is a very special case, and doesn't really occur in well-written code.
count_nonzero(arr) and any(arr) (and not any(arr)) can both short-circuit, so should be more efficient than np.all(arr == 0).
These are actually pretty good tips... I should have thought of these, or tried to copy this message to AI to suggest them to the zarr project.
In trying to implement it, I realize that zarr doesn't only only check for "equality to zero" but rather equality to the fill value, which can be non-zero.
To do so, they somewhat defined their own function
all_equal.I'm happy to discuss the inclusion of an API where we check equality to a scalar, or an "other", and not consider the other "proposals" i made (including the "likely not").
Note that that's pretty unlikely to fly. Functions like that may in some cases make sense for eager libraries, but we already have
allandequalin the standard, and JIT compilers can transparently optimizeall(equal(..)toall_equal(...)so there's no benefit to a separate function. And there are an endless amount of composed operations that are faster in eager mode (thinkadd_mul, etc.) - NumPy et al don't have those because the tradeoffs of increased API surface vs. performance gain on the operation in quesiton aren't good.My understanding was that JITs were still complicated to use in Windows.
Personally, i tend to avoid them due to the "warmup costs".
I haven't used them in a while, maybe numba got alot of better, and the "caching" of the compiled results is better today, but I think a real concern is startup time for short lived scripts.
AOT compilation is still valuable in 2026 IMO....
You're all correct there, there's an overhead in using JITs. But we still have to deal with design that work for JIT/AOT/eager, so we're very reluctant to add eager-only APIs. There has to be a really compelling case, and performance of a composed operation isn't that.
Perhaps, i'm missing something, is there a path for numpy to optimize operations like
np.all(arr == 0)without allocating an entire other array of size
arr.sizebytes?Specifically, my understanding is that
np.all(arr == 7)would decompose into:
- A memory allocation of
O(arr.size)for the result of arr == 7 - Memory read from RAM (all cache misses for "large" arrays) for the contents of
arr - Memory writes, to RAM (all causing cache evictions) for the contents of
arr == 7 - A memory read of the order of
arr.sizeto establish that all the elements are in fact 7.
In the above decomposition, 1, 3, and 4 are all unnecessary for this kind of "is my array really a constant" kind of question.
- A memory allocation of
is there a path for numpy to optimize operations like
There isn't for
numpy, your analysis looks correct.Thanks or the discussion:
For me the discussion has led to two options:
- Option 1: No new API: Numpy must figure out how to optimize
array_equal. JITs must "inspect" things likestridesto optimize things.
a = np.zeros((100, 1024, 1024), dtype='float64') def array_equal_to_scalar(array, scalar): s = np.broadcast_to(np.array(scalar, dtype=array.dtype), array.shape) return np.array_equal(array, s) array_equal_to_scalar(a, 7)
- Option 2: Create add to the API to help guide people to this optimization
a = np.zeros((100, 1024, 1024), dtype='float64') np.array_equal_to_scalar(array, 7)
I personally think that the whole
broadcast_totrick is so niche it often looks like bad code.Just reading it, I feel compelled to, in 2 years to write code like:
# This code will help trigger a fastpath in numpy that # detects that the strides of the "seven" array array # are zero, and just "compare to a scalar" seven = np.broadcast_to(np.array(7, dtype=array.dtype), array.shape)
IMO: These 3 lines of comments, just feel "wrong" to me.
Reacted by Neil Girdhar- Option 1: No new API: Numpy must figure out how to optimize
Metadata
Metadata
Assignees
Labels
Type
Projects
- StatusShow more project fieldsStage 0
Today, there are a few cases where people might want to check equality to a scalar.
However, this code is terribly inefficient for out of memory array.
It can force them to go through the entire memory.
I think that we can do much better to define an operation that will check for
and those can be implemented in more "streamed" fashion, that allow the underlying implementation to not create a full boolean array for
arr == 0.Second, I would like to propose an API for "likely not equal", where the strick inequality is not guaranteed.
For data compression, we might just be intrested in learning if the dataset is worth compressing or not:
Zarr for example does this:
zarr-developers/zarr-python#3627
For example, consider the task to compress a 800MB array.
Zarr today:
But if you have an out of memory data array, it may be hard to guarantee that things are not all zero, but the check likely doesn't matter, since with modern compression algorithm all zeros can be efficiently compressed.
On my computer, the equality check is roughly the same order of magnitude as the compression itself:
so if zarr misses one of my images, because I think it has non-zero element, its likely not the end of the world in terms of system performance, but if I have to "guarantee that the images are non zero" that can be a costly operation that can't be done through my implementation easily "without checking every single element".
For not, without these APIs, today, zarr does something like:
Alternative:
I could likely check numpy metadata like
stridesand create my own fastpath for this:but this seems more like a hack and would be implicitely redefining "equality operation" to "likely equal" which isn't correct.