Loading Python for AI Engineering...
Utilizing httpx.AsyncClient with persistent TLS connection reuse and HTTP/2 multiplexing to eliminate repeated TCP 3-way handshakes and TLS negotiations on high-throughput LLM API calls.
“Connection pooling is a permanent dedicated phone line between two offices, instead of dialing the operator, negotiating numbers, and verifying identity on every single word.”
Instantiating a new httpx.AsyncClient() inside every function call, closing sockets immediately and suffering connection exhaustion (TIME_WAIT).
import httpx
limits = httpx.Limits(max_keepalive_connections=20, max_connections=100)
timeout = httpx.Timeout(connect=5.0, read=30.0, write=5.0, pool=5.0)
# Singleton client pattern
async def get_client() -> httpx.AsyncClient:
return httpx.AsyncClient(limits=limits, timeout=timeout, http2=True)
# Usage across entire server lifecycle:
# client = await get_client()
# response = await client.post("https://api.openai.com/v1/...", ...)