HOWTO · NumPy

Calculate Euclidean Distance With NumPy

Calculate Euclidean distance with NumPy, compare equivalent Python methods, compute row-wise distances, and diagnose incompatible shapes.

On this page

For two NumPy arrays that represent points with the same number of coordinates, calculate the Euclidean distance with np.linalg.norm(point_a - point_b). Subtraction creates the coordinate differences, and the vector’s L2 norm reduces those differences to one nonnegative distance.

Understand the Euclidean Distance Formula

For points (a=(a_1,\ldots,a_n)) and (b=(b_1,\ldots,b_n)), Euclidean distance is the square root of the sum of the squared coordinate differences:

Euclidean distance formula: the square root of the sum of squared coordinate differences.

In NumPy, np.linalg.norm(a - b) expresses that definition directly. With a one-dimensional difference array and no ord argument, np.linalg.norm computes its 2-norm. Both inputs must describe points in compatible dimensions so that subtraction has the intended meaning.

Prepare Input Coordinates

Coordinate positions must describe the same axes in the same order: compare the first coordinate of one point with the first coordinate of the other, and so on. Represent one point as a one-dimensional array shaped (coordinates,). A pair of arrays shaped (3,), for example, represents two points in the same three-dimensional coordinate system; it does not represent three independent distances.

Tuples and lists can be converted with np.asarray(values, dtype=float) before subtraction. The floating type accepts integer or decimal coordinates and avoids unsigned-integer wraparound when one coordinate is smaller than its counterpart. If the inputs are already suitable floating-point NumPy arrays, conversion is unnecessary. Avoid flattening arbitrary multidimensional data merely to make shapes match, because that can hide an input-structure error rather than correct it.

The examples below were executed with Python 3.14.7, NumPy 2.5.3, and SciPy 1.18.1. The demonstrated APIs are established interfaces, but an application should still use versions compatible with its own supported Python environment. SciPy is optional unless the SciPy method or pairwise-distance tools are used.

Calculate the Distance Between Two Points

The following verified example compares the recommended norm expression with the explicit formula, a dot-product formulation, the standard-library math.dist, and SciPy’s scipy.spatial.distance.euclidean. All five calculations produce the same result for this point pair.

"""Verify equivalent Euclidean-distance APIs for one pair of points."""

import math

import numpy as np
from scipy.spatial import distance

point_a = np.array([1.0, 2.0, 3.0])
point_b = np.array([4.0, 5.0, 6.0])

print(f"np.linalg.norm: {np.linalg.norm(point_a - point_b)}")
print(f"formula: {np.sqrt(np.sum((point_a - point_b) ** 2))}")
delta = point_a - point_b
print(f"dot product: {np.sqrt(np.dot(delta, delta))}")
print(f"math.dist: {math.dist(point_a, point_b)}")
print(f"distance.euclidean: {distance.euclidean(point_a, point_b)}")
np.linalg.norm: 5.196152422706632
formula: 5.196152422706632
dot product: 5.196152422706632
math.dist: 5.196152422706632
distance.euclidean: 5.196152422706632

The explicit formula uses element-wise squaring with NumPy conceptually, followed by a sum and a square root. The exponent form in the example keeps the complete calculation in one expression. The dot-product variant computes the same sum of squares because np.dot(delta, delta) multiplies corresponding differences and adds them.

Use np.linalg.norm when the points are already arrays or are part of a larger NumPy calculation. It communicates the vector operation clearly without spelling out the reduction.

Calculate Row-Wise Distances With axis=1

For several points stored as the rows of a two-dimensional array, subtract one reference point and set axis=1. NumPy broadcasts the reference array across the rows, while axis=1 tells np.linalg.norm to return one norm per row rather than one norm for the whole matrix.

"""Verify row-wise distances from several points to one reference point."""

import numpy as np

points = np.array([[1.0, 2.0, 3.0], [4.0, 5.0, 6.0]])
reference = np.array([1.0, 2.0, 3.0])

print(np.linalg.norm(points - reference, axis=1))
[0.         5.19615242]

The first row equals the reference, so its distance is zero. The second result is the distance calculated in the one-pair example. This pattern covers one-to-many distances; for all pairwise distances between two collections, use a dedicated routine such as scipy.spatial.distance.cdist instead of creating a large broadcasted intermediate array without considering memory use.

Choose Between NumPy, math.dist, and SciPy

The best method depends on the surrounding code, not a universal speed claim:

  • Use np.linalg.norm(a - b) for NumPy arrays, vectorized workflows, and row-wise calculations with axis.
  • Use math.dist(a, b) for one pair of ordinary Python coordinate iterables when NumPy is otherwise unnecessary.
  • Use scipy.spatial.distance.euclidean(a, b) when SciPy is already a dependency or when the calculation belongs with SciPy’s broader distance tools.
  • Use the explicit square-sum formula or dot product when teaching, auditing, or adapting the underlying calculation. They do not make the result more Euclidean than the norm expression.

Converting lists with np.asarray(..., dtype=float) is useful before NumPy subtraction because Python lists do not support element-wise subtraction. math.dist accepts equal-length coordinate iterables directly, while the NumPy and SciPy choices require their respective installed packages.

Handle Incompatible Shapes and Numeric Data Types

Two individual points must have the same number of coordinates. Incompatible one-dimensional shapes cannot be broadcast together, so NumPy raises a ValueError before it can calculate a distance.

"""Capture the diagnostic for points with incompatible dimensions."""

import numpy as np

point_a = np.array([1.0, 2.0])
point_b = np.array([3.0, 4.0, 5.0])

try:
    np.linalg.norm(point_a - point_b)
except ValueError as error:
    print(f"ValueError: {error}".rstrip())
ValueError: operands could not be broadcast together with shapes (2,) (3,)

Check the shapes before subtraction when dimensions come from user input or external data. For row-wise distances, a points array shaped (rows, coordinates) is compatible with a reference shaped (coordinates,); another shape may broadcast in an unintended way or fail.

Also cast unsigned integer coordinates to a signed or floating type before subtracting. Unsigned subtraction can wrap around instead of representing a negative coordinate difference, which makes the subsequent norm incorrect. Using floating-point arrays in the examples avoids that issue and supports non-integer coordinates.

For one pair of equal-dimensional NumPy points, np.linalg.norm(point_a - point_b) is the concise default. Add axis=1 for row-wise distances, and validate shapes and data types whenever the input structure is not already guaranteed.