Throughput vs. latency
Latency is how long you wait. Throughput is how fast work gets done once it's moving. A model can start quickly (low time to first token) but write slowly, or the other way around. For short replies, latency matters most. For long ones, throughput does.
As an example, at 50 output tokens per second, a 500-token answer takes about 10 seconds to write, plus the time to the first token.
Measuring it
- Send a request and note the time.
- Note when the first token arrives and when the last one does.
- Divide
usage.completion_tokensby the time between first and last token.
Run it a few times and at different times of day. Numbers vary with load.
Raising app throughput
- Run several requests at once instead of one after another.
- Keep prompts short so each request uses fewer resources.
- Use a faster model for bulk or simple work.
- Batch small tasks into one prompt when the answers are short and independent.
Example
This streams a longer reply so you can watch the pace:
curl -N https://api.luv13.ai/v1/chat/completions \
-H "Authorization: Bearer $LUV13_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "luv13/glm-5.3-flash",
"messages": [{"role": "user", "content": "Write 200 words about rivers."}],
"stream": true
}'