Loading Python for AI Engineering...
Controlling concurrency using asyncio.Semaphore to enforce vendor TPM (Tokens Per Minute) and RPM (Requests Per Minute) boundaries without getting rejected with HTTP 429 Too Many Requests.
“A semaphore is a nightclub bouncer: only 5 patrons are allowed on the dance floor at once. When one leaves, the next in line enters.”
Launching 500 parallel API calls without a semaphore, instantly exhausting OpenAI/Anthropic rate limits and triggering 429 backoff penalties.
import asyncio
sem = asyncio.Semaphore(3) # Max 3 concurrent requests
async def safe_llm_call(task_id: int):
async with sem: # Waits if 3 tasks are already inside
print(f"Task {task_id} acquired permit")
await asyncio.sleep(0.5) # Network I/O
print(f"Task {task_id} released permit")
return f"Result {task_id}"
async def main():
await asyncio.gather(*(safe_llm_call(i) for i in range(9)))