Picking a value
How long a reply takes depends on the model, the prompt length, and how many tokens it writes. Reasoning Models can take much longer than fast models. As an example, 60 seconds is a reasonable start for short chat replies, and several minutes may be needed for long outputs. Measure your real requests and set the limit a bit above the slow ones.
Types of timeout
- Connect timeout: how long to wait to open the connection. Keep this short, a few seconds.
- Read timeout: how long to wait for data once connected. For a non-streamed reply, that means waiting for the whole answer.
- Total timeout: a cap on the whole request.
Defaults in common tools
- Python
requestshas no timeout unless you passtimeout=. - The official OpenAI SDKs default to 10 minutes and accept a
timeoutoption. See Using the OpenAI SDKs. - Node's
fetchhas no timeout by default. UseAbortSignal.timeout(). See Node.js Fetch.
Example
curl's --max-time sets a total limit in seconds:
curl --max-time 60 https://api.luv13.ai/v1/chat/completions \
-H "Authorization: Bearer $LUV13_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "luv13/glm-5.3-flash",
"messages": [{"role": "user", "content": "Write a haiku about rain."}],
"max_tokens": 100
}'