Loading Python for AI Engineering...
Using generator functions (yield) and itertools to stream, filter, and batch massive token datasets line-by-line in constant O(1) memory, rather than loading multi-gigabyte JSONL files into RAM.
“A generator is an assembly line conveyor belt delivering one part at a time, rather than a dump truck unloading 100,000 parts onto your living room floor all at once.”
Calling list(generator) on a 10-million row pretraining dataset, immediately triggering an Out-Of-Memory (OOM) kernel kill.
from typing import Iterator
def stream_jsonl(filepath: str) -> Iterator[str]:
with open(filepath, 'r', encoding='utf-8') as f:
for line in f:
yield line.strip() # O(1) memory footprint
def batch_stream(stream: Iterator[str], batch_size=32) -> Iterator[list[str]]:
batch = []
for item in stream:
batch.append(item)
if len(batch) == batch_size:
yield batch
batch = []
if batch: yield batch