Loading Python for AI Engineering...
Einstein summation notation provides a unified, concise declarative syntax (np.einsum / torch.einsum) to express transpositions, dot products, batch matrix multiplications, and multi-head attention without intermediate tensor reshaping.
“einsum is like a universal CNC machine: instead of passing a workpiece through 4 separate cutting, drilling, and turning machines, one instruction matrix does the whole contraction.”
Writing convoluted reshape-permute-matmul pipelines that allocate intermediate tensors instead of a single einsum operation.
import numpy as np
B, H, S, D = 2, 4, 16, 64 # Batch, Heads, SeqLen, Dim
Q = np.random.randn(B, H, S, D)
K = np.random.randn(B, H, S, D)
# Multi-head attention raw affinities: Q @ K.T per head
# 'bhsd,bhmd->bhsm' -> contracts across hidden dimension 'd'
scores = np.einsum('bhsd,bhmd->bhsm', Q, K)
print("Attention scores shape:", scores.shape) # (2, 4, 16, 16)