Loading Python for AI Engineering...
Consuming and emitting real-time token streams via async generators (yield in async def). Powers modern Server-Sent Events (SSE) and WebSocket endpoints to achieve Time-To-First-Token (TTFT) under 300ms.
“Streaming is drinking water from a tap as it flows; non-streaming is waiting for a 5-gallon bucket to fill before taking a single sip.”
Buffering all stream chunks into a string inside the generator before yielding, completely defeating the purpose of streaming.
import asyncio
from typing import AsyncGenerator
async def mock_token_stream(prompt: str) -> AsyncGenerator[str, None]:
tokens = ["The", " future", " of", " AI", " is", " streaming."]
for token in tokens:
await asyncio.sleep(0.08) # Simulate token generation latency
yield token
async def client_consumer():
print("Beginning stream: ", end="", flush=True)
async for chunk in mock_token_stream("hello"):
print(chunk, end="", flush=True)
print("
[Stream complete]")