Loading Python for AI Engineering...
AI systems trade numerical precision for throughput and memory efficiency. Floating-point types vary in exponent and mantissa allocation: FP32 (8 exp, 23 mantissa), FP16 (5 exp, 10 mantissa), and BF16 (8 exp, 7 mantissa).
“FP32 is a high-resolution 4K architectural blueprint. INT8 is a rough sketch on a napkin: less precise, but 100x lighter to carry and fast enough to build the frame.”
Casting raw probabilities to FP16 and taking logarithms, resulting in -inf when probabilities round down to true zero.
import numpy as np
# Precision comparisons
fp32_val = np.float32(1e-6)
fp16_val = np.float16(1e-6) # Precision lost / subnormal!
print("FP32:", fp32_val)
print("FP16:", fp16_val)
# Memory impact: 7 Billion parameter model
# FP32: 7B * 4 bytes = 28 GB VRAM
# FP16: 7B * 2 bytes = 14 GB VRAM
# INT8: 7B * 1 byte = 7 GB VRAM
# INT4: 7B * 0.5 byte = 3.5 GB VRAM