Loading Python for AI Engineering...
Techniques for executing dozens or hundreds of LLM calls in parallel using asyncio.gather, TaskGroups (Python 3.11+), and asyncio.as_completed for low-latency map-reduce agentic workflows.
“Fan-out is like dispatching 10 delivery drivers simultaneously rather than sending one driver on 10 sequential round trips.”
Failing to handle exceptions inside asyncio.gather without return_exceptions=True, causing a single failure to abort all 99 successful results.
import asyncio
async def process_chunk(chunk_id: int):
if chunk_id == 2:
raise ValueError("Token budget exceeded")
await asyncio.sleep(0.1)
return f"Chunk {chunk_id} embeddings"
async def safe_fanout():
tasks = [process_chunk(i) for i in range(4)]
# return_exceptions=True prevents cascade abortion
results = await asyncio.gather(*tasks, return_exceptions=True)
for r in results:
if isinstance(r, Exception):
print("Handled error:", r)
else:
print("Success:", r)