Where the time goes
- Network: your request reaching the API and the reply coming back.
- Queueing: waiting for capacity if the service is busy.
- Reading the prompt: the model processes every input token. Longer prompts take longer.
- Writing the reply: tokens come out one after another, so a 1,000-token answer takes much longer than a 50-token one.
The last step is often the biggest. That's why output length matters so much. See Throughput for the speed of that step.
Ways to cut it
- Stream the reply so users see progress.
- Ask for shorter answers, and set
max_tokens. - Trim the prompt: drop old chat turns and unneeded context.
- Pick a faster model for simple tasks. luv13's list includes ids like
luv13/glm-5.3-flashandluv13/kimi-k3-fast, but test the speed yourself. Names are a hint, not a promise. - Send independent requests in parallel instead of one after another, within your rate limits.
Measuring it
curl can report timing for a request:
curl -s -o /dev/null \
-w "first byte: %{time_starttransfer}s, total: %{time_total}s\n" \
https://api.luv13.ai/v1/chat/completions \
-H "Authorization: Bearer $LUV13_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "luv13/glm-5.3-flash", "messages": [{"role": "user", "content": "Hi"}], "max_tokens": 20}'For a non-streamed request, "first byte" is close to the full time. Add "stream": true and -N to measure time to first token instead.
