Retrying Requests

Definition

Retrying requests means sending a failed luv13 call again after a growing wait, and only for errors that can succeed on a second try.

Key takeaways

  • luv13 runs on single-provider capacity. When it's saturated, requests queue or fail, and luv13.ai/docs asks you to retry with backoff rather than hammering.
  • Retry timeouts, 429 and 5xx errors (including Cloudflare's 522). Don't retry 401, 404 or 405; they'll fail the same way.
  • Wait longer after each failure (for example 1, 2, 4, 8 seconds) and add a little random jitter.
  • Failed calls aren't charged, so retrying a failure doesn't cost extra.
  • If a model is unavailable, the error names it. Switching to another id can beat waiting. See Model Fallback.

What to retry

ResultRetry?
Timeout or dropped connectionYes
429Yes, and honor Retry-After if it's sent
500, 502, 503, 504, 522Yes
401, 404, 405No, fix the request

With the OpenAI SDKs

Both official SDKs already retry 408, 409, 429 and 5xx responses with backoff (checked in openai Python 3.22.1 and Node.js 7.25.0). Python defaults to 2 retries. Raise it if you'd rather wait than fail:

import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.luv13.ai/v1",
    api_key=os.environ["LUV13_API_KEY"],
    max_retries=5,
    timeout=120,
)

In Node.js, pass maxRetries: 5 to new OpenAI({...}).

With curl

curl --retry 5 --retry-delay 0 --retry-max-time 120 \
  https://api.luv13.ai/v1/chat/completions \
  -H "Authorization: Bearer $LUV13_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "luv13/glm-5.3-flash", "messages": [{"role": "user", "content": "ping"}]}'

--retry-delay 0 keeps curl's own doubling backoff. curl retries timeouts, 408, 429, 500, 502, 503, 504 and a few other codes, but not 522. For 522, use a loop or the SDKs.