NumPy Basics: Arrays, Broadcasting, Views vs Copies, and dtype Pitfalls

Key takeaways

NumPy is fast because an ndarray is one typed block of memory processed by compiled loops. The same design explains its sharp edges: views that alias data, silent integer overflow, and broadcasting errors.

What NumPy actually is

NumPy’s central object is the ndarray: one contiguous block of memory holding values of a single type (the dtype), plus metadata — shape, and strides that say how many bytes to jump to reach the next element along each axis. A Python list, by contrast, is an array of pointers to separate Python objects, each with its own type header and reference count.

That design explains both halves of NumPy’s reputation. Operations like arr * 2 run as a compiled loop over raw numbers, which is why they are fast. And because many operations just create new metadata over the same buffer, and because the dtype is fixed, you get views that alias data and integers that overflow silently. This article covers the basics with those consequences in mind. All outputs below were produced with NumPy 1.26 on CPython 3.11; where NumPy 2 behaves differently, it is noted.

pip install numpy

Creating arrays and choosing a dtype

import numpy as np

arr = np.array([1, 2, 3, 4, 5])
arr2d = np.array([[1, 2, 3], [4, 5, 6]])

np.zeros((3, 4))          # float64 zeros
np.ones((2, 3), dtype=np.int64)
np.empty((2, 2))          # uninitialized: contains whatever was in memory
np.arange(0, 10, 2)       # [0 2 4 6 8]
np.linspace(0, 1, 5)      # [0.   0.25 0.5  0.75 1.  ]

np.empty is not “an array of zeros that is slightly faster”. Its contents are arbitrary leftovers; use it only when you are about to overwrite every element.

NumPy infers the dtype from the data, and the inference is worth knowing:

np.array([1, 2.5]).dtype     # float64 - ints upcast to float
np.array([1, "a"]).dtype     # <U11    - everything became a string
np.array([1, None]).dtype    # object  - boxed Python objects, no speedup

An object array has lost everything that makes NumPy fast; it is a list in disguise. If you see dtype=object after loading data, something (usually a missing value or a stray string) needs cleaning.

Two version notes. np.float, np.int, and np.bool were aliases for built-ins, deprecated in NumPy 1.20 and removed in 1.24, so old tutorials now fail with AttributeError: module 'numpy' has no attribute 'float'. Use float or np.float64. And the default integer dtype was platform-dependent before NumPy 2: np.array([1, 2, 3]).dtype is int32 on Windows with NumPy 1.x and int64 on Linux/macOS. NumPy 2 made it int64 on 64-bit Windows too. If you need a specific width, say so with dtype=.

For random numbers, prefer the Generator API over the legacy np.random.randint family:

rng = np.random.default_rng(seed=0)
rng.integers(0, 10, 5)
rng.normal(size=(5, 3))

A Generator object carries its own state, so two components seeded independently do not interfere with each other through global state.


Vectorization and why loops are still slow

arr = np.array([1, 2, 3, 4, 5])
arr + 10      # [11 12 13 14 15]
arr ** 2      # [ 1  4  9 16 25]
arr * np.array([10, 20, 30, 40, 50])   # [ 10  40  90 160 250]

The operators apply element-wise in C. The trap is thinking that “using NumPy” is enough. On my machine, with one million elements, I measured:

CodeTime per run (measured, min of repeats)
arr ** 2 on an ndarrayabout 0.7 ms
[x ** 2 for x in lst] on a Python listabout 50 ms
[x ** 2 for x in arr] on an ndarrayabout 84 ms

These are single-machine numbers from timeit, only useful for the ratio. The interesting row is the last one: iterating over an ndarray in Python is slower than iterating over a list, because every element has to be boxed into a new NumPy scalar object. Vectorization means expressing the whole computation as array operations, not putting the data into an array and then looping.


Indexing: views vs copies

arr2d = np.array([[1, 2, 3], [4, 5, 6], [7, 8, 9]])
arr2d[0, 0]    # 1
arr2d[1, :]    # [4 5 6]  row 1
arr2d[:, 1]    # [2 5 8]  column 1
arr[arr > 3]   # boolean mask -> [4 5]

The rule you need to memorize:

  • Basic slicing (a[1:4], a[:, 0], a[::2]) returns a view — new metadata over the same buffer.
  • Boolean masks and integer-array indexing (a[a > 2], a[[0, 1]]) return a copy.
a = np.arange(6)
s = a[2:5]
s[0] = 100
print(a)        # [  0   1 100   3   4   5]  -- original changed

a = np.arange(6)
f = a[a > 2]
f[0] = -1
print(a)        # [0 1 2 3 4 5]  -- original untouched

Views are why slicing a huge array is essentially free, and also why a function that “normalizes a slice” can corrupt the caller’s data. The same split exists in reshaping: reshape and ravel return views when the memory layout allows it, while flatten always copies. Whether a view is possible depends on strides — m.ravel() shares memory with a C-contiguous m, but m.T.ravel() has to copy because the transposed data is not contiguous in that order. When in doubt, check with np.shares_memory(a, b) instead of guessing.

The failure mode I see most often looks like this: a preprocessing function takes a slice of a shared array (say, the training rows), subtracts the mean in place with x -= x.mean(axis=0), and returns it. The caller’s full dataset is now partially normalized, and nothing errors — the model just trains on slightly wrong data. My habit now is that functions either document that they mutate their input or start with x = x.copy(), and in-place operators on arguments get extra scrutiny in review.

A related trap from the pandas side: df.values / df.to_numpy() may return a view or a copy depending on the dtypes involved, so do not rely on writes to it flowing back into the DataFrame.


Broadcasting, including when it fails

Broadcasting lets arrays of different shapes combine without copying data. NumPy compares shapes from the last dimension backwards; each pair of sizes must be equal, or one of them must be 1 (a missing dimension counts as 1).

matrix = np.array([[1, 2, 3], [4, 5, 6]])   # (2, 3)
vector = np.array([10, 20, 30])             # (3,)
matrix + vector
# [[11 22 33]
#  [14 25 36]]

(2, 3) vs (3,): last dimensions are 3 and 3, match; the vector is stretched across rows. Now try to add a per-row value:

x = np.ones((2, 3))
y = np.array([1, 2])        # one value per row
x + y
# ValueError: operands could not be broadcast together with shapes (2,3) (2,)

Aligned from the right, 3 and 2 conflict. You have to make the intent explicit by giving y shape (2, 1):

x + y.reshape(2, 1)         # or y[:, None]
# [[2. 2. 2.]
#  [3. 3. 3.]]

This comes up constantly with reductions. Subtracting each row’s mean fails unless you keep the reduced axis:

arr2d = np.array([[1, 2, 3], [4, 5, 6]])
arr2d - arr2d.mean(axis=1)
# ValueError: operands could not be broadcast together with shapes (2,3) (2,)
arr2d - arr2d.mean(axis=1, keepdims=True)
# [[-1.  0.  1.]
#  [-1.  0.  1.]]

The dangerous case is the opposite one: broadcasting that succeeds when you did not mean it to. A (3,) array and a (3, 1) array broadcast to (3, 3), so predictions - targets with one of them accidentally shaped as a column gives you a 3x3 matrix of differences, and .mean() of that produces a plausible-looking number that is simply wrong. When I compute losses or metrics, I assert shapes (assert pred.shape == target.shape) before the arithmetic; it is one line and it has caught this more than once.

Also note that a 1-D array has no orientation: v.T on shape (3,) is still (3,). If you need a column vector, use v[:, None] or v.reshape(-1, 1).


Reductions and the meaning of axis

arr2d = np.array([[1, 2, 3], [4, 5, 6]])
arr2d.sum(axis=0)   # [5 7 9]   collapse axis 0 -> one value per column
arr2d.sum(axis=1)   # [ 6 15]   collapse axis 1 -> one value per row

The mental model that works: axis=k is the axis that disappears. For an image of shape (height, width, channels), mean(axis=2) averages over channels and leaves (height, width).

Watch the statistics defaults. np.std([1, 2, 3, 4, 5]) is 1.414..., the population standard deviation (ddof=0). pandas’ .std() defaults to the sample standard deviation (ddof=1), which gives 1.581... for the same data. If NumPy and pandas numbers disagree slightly, this is usually why.


dtype overflow and precision

Because the dtype is fixed, NumPy integers do not grow like Python integers. They wrap around silently:

img = np.array([200, 250], dtype=np.uint8)
img + 50
# [250  44]     250 + 50 = 300 -> 300 - 256 = 44

np.array([2**31 - 1], dtype=np.int32) + 1
# [-2147483648]

This is the classic image-processing bug: “brighten by 50” turns the brightest pixels nearly black. A common first attempt is np.clip(image + 50, 0, 255), which does not help — the overflow already happened before clip ran. Widen first, then clip, then narrow:

np.clip(img.astype(np.int16) + 50, 0, 255).astype(np.uint8)
# [250 255]

(NumPy 2 changed the promotion rules for Python scalars, NEP 50: img + 50 still wraps because 50 fits in uint8, but img + 300 now raises an OverflowError instead of quietly upcasting.)

Reductions are a bit friendlier: sum() on small integer types accumulates in the platform integer (on Windows with NumPy 1.x that is int32, or uint32 for unsigned input; int64/uint64 elsewhere), so summing a thousand uint8 values of 200 correctly gives 200000. But that only moves the ceiling; summing large int32 counters on Windows can still overflow.

Floats have the opposite problem, precision. Summing ten million float32 values of 0.1 gave me 999989.44 instead of 1000000, while the same sum in float64 gave 999999.99999998. Use float64 for accumulations unless memory forces otherwise, and never compare floats with == — use np.isclose / np.allclose (0.1 + 0.2 == 0.3 is False).

One more comparison trap: if a == b: with arrays raises ValueError: The truth value of an array with more than one element is ambiguous. Use a.any() or a.all(), because == is element-wise. Use np.array_equal(a, b) or np.allclose(a, b).


Linear algebra

A = np.array([[1, 2], [3, 4]])
B = np.array([[5, 6], [7, 8]])

A @ B        # matrix product
# [[19 22]
#  [43 50]]
A * B        # element-wise product, NOT matrix multiplication
# [[ 5 12]
#  [21 32]]
A.T          # transpose
np.linalg.eig(A)[0]   # eigenvalues: [-0.37228132  5.37228132]

Confusing * with @ is the most common linear-algebra bug for people coming from MATLAB, where * is the matrix product. For 2-D arrays np.dot and @ agree; for higher dimensions they differ, and @ (batched matrix multiply over leading axes) is usually what you want.

To solve Ax = b, do not compute the inverse:

b = np.array([5, 6])
np.linalg.solve(A, b)          # [-4.   4.5]
# rather than np.linalg.inv(A) @ b

solve uses an LU factorization directly, which is faster and numerically more stable than forming the inverse explicitly, especially for ill-conditioned matrices. np.linalg.inv is rarely the right tool outside of teaching examples. Also note eig can return complex eigenvalues for general matrices; for symmetric matrices use np.linalg.eigh, which is faster and guarantees real results.


Putting it together: two small, realistic examples

Grayscale conversion without the overflow bug

rng = np.random.default_rng(0)
image = rng.integers(0, 256, (100, 100, 3), dtype=np.uint8)   # (H, W, C)

# Weighted luminance (ITU-R BT.601 weights), computed in float64
weights = np.array([0.299, 0.587, 0.114])
gray = (image @ weights).astype(np.uint8)    # (100, 100)

image @ weights contracts the last axis (channels) against the weight vector. Because weights is float64, the arithmetic happens in float and cannot wrap; the conversion back to uint8 happens only at the end, when the values are known to be in range.

Row-wise L2 normalization

X = rng.normal(size=(5, 3))
norms = np.linalg.norm(X, axis=1, keepdims=True)    # (5, 1)
Xn = X / np.clip(norms, 1e-12, None)                # avoid division by zero
np.linalg.norm(Xn, axis=1)                          # [1. 1. 1. 1. 1.]

keepdims=True makes the (5, 1) shape broadcast against (5, 3), and the clip guards against all-zero rows, which would otherwise produce nan and a RuntimeWarning rather than an exception.


When NumPy is not the right tool

NeedBetter fit
Labeled columns, mixed types, joins, group-bypandas
A loop that cannot be vectorized (data-dependent control flow)Numba @njit, or Cython
GPU execution or automatic differentiationJAX, PyTorch, CuPy
Arrays larger than memorymemory-mapped arrays (np.memmap), Dask, Zarr

Further reading

The NumPy user guide is the reference for everything above, and its broadcasting chapter has the diagrams worth looking at the first time a shape mismatch surprises you.