TIL / Token-aware batching beats fixed-size batching for embedding calls
Token-aware batching beats fixed-size batching for embedding calls
The problem
An embedding endpoint enforces a tokens-per-minute limit, not just a requests-per-minute one.
Batching by a fixed item count (say, 100 texts per call) works fine until a run of unusually
long chunks pushes one batch over the token limit and the call comes back 429. Under peak
indexing load, that turned into a steady trickle of failed batches.
The fix
Track a running token estimate per batch and flush it when either a text-count cap or a token
cap is hit, whichever comes first. On a 429, honor the Retry-After header instead of a fixed
backoff.
def batch_texts(texts, max_items=100, max_tokens=15_000, estimate_tokens=len_tokens):
batch, batch_tokens = [], 0
for text in texts:
t = estimate_tokens(text)
if batch and (len(batch) >= max_items or batch_tokens + t > max_tokens):
yield batch
batch, batch_tokens = [], 0
batch.append(text)
batch_tokens += t
if batch:
yield batch
Gotcha
A token estimate (even a rough one, like len(text) // 4) is enough to keep batches under
budget in practice - don’t reach for the real tokenizer on the hot path just to save a rare
retry. The bigger win is honoring Retry-After exactly rather than guessing a sleep duration;
a fixed backoff that’s shorter than the server’s own window just produces the same 429 again.