Input tokens
These are the tokens in your request: the system message, the chat history, your new message, and any tool definitions or tool results. Because a model has no memory between calls, a long chat re-sends its history each time, so input tokens grow with every turn.
Output tokens
These are the tokens the model generates in its reply, including any tool calls it makes. You can set a ceiling with the max_tokens parameter. If the reply hits that ceiling, it stops early, and finish_reason is length instead of stop.
Reading the counts
Each response reports both counts:
"usage": {
"prompt_tokens": 850,
"completion_tokens": 120,
"total_tokens": 970
}These numbers are an example. prompt_tokens is input, completion_tokens is output, and total_tokens is the sum.
How pricing uses them
Your cost for a request is the input tokens times the input price plus the output tokens times the output price. On luv13 the input and output prices are the same, a flat $0.33 per 1 million tokens on every model, so you can just multiply total_tokens by that one rate. See Pricing.
For the example above, 970 total tokens cost 970 / 1,000,000 x $0.33, or about $0.00032.
Example
This request caps the reply at 100 output tokens. The usage field in the response shows both counts.
curl https://api.luv13.ai/v1/chat/completions \
-H "Authorization: Bearer $LUV13_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "luv13/glm-5.3-flash",
"messages": [{"role": "user", "content": "Explain tokens in two sentences."}],
"max_tokens": 100
}'