Loading Python for AI Engineering...
Constructing Python context managers (__enter__ and __exit__ / @contextmanager) to automate GPU VRAM cache clearing, OpenTelemetry trace spans, and microsecond-precision latency logging.
“A context manager is an automatic door closer: no matter how hurried you are or what happens inside the room, the door is guaranteed to close and lock behind you.”
Forgetting to handle exceptions in __exit__, either swallowing critical errors unintentionally or leaving unclosed GPU resources.
import time
from contextlib import contextmanager
@contextmanager
def inference_timer(stage_name: str):
start = time.perf_counter()
try:
yield
finally:
elapsed = (time.perf_counter() - start) * 1000
print(f"[{stage_name}] Latency: {elapsed:.2f}ms")
# Usage:
with inference_timer("LLM Forward Pass"):
time.sleep(0.12) # Simulating forward computation