How it works
A model writes its reply one token at a time. A normal request holds all of those tokens on the server and sends them together when the reply is done. A streaming request sends each new bit of text as soon as it exists.
The total time to finish is about the same either way. What changes is how soon you see something. For a chat window or a coding tool, that early first word makes a big difference to how the app feels.
What you get back
In the OpenAI format, a streamed reply is a series of chunks. Each chunk has a choices[0].delta object with the new piece of text in delta.content. You join the pieces in order to rebuild the full reply. The last chunk carries a finish_reason, such as stop or length.
The wire format for these chunks is covered in Server-Sent Events.
When to use it
- Use it for chat interfaces, editors and anything a person watches in real time.
- Skip it for background jobs where you only need the final text. A single response is simpler to handle.
Example
The -N flag tells curl not to buffer, so you see chunks as they arrive. Check Streaming on luv13 for the exact behavior luv13 supports.
curl -N https://api.luv13.ai/v1/chat/completions \
-H "Authorization: Bearer $LUV13_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "luv13/glm-5.3-flash",
"messages": [{"role": "user", "content": "Count from 1 to 5."}],
"stream": true
}'