What it does
At each step, a model gives every possible next token a probability. Temperature reshapes those probabilities before one is picked:
- Below 1, the likely tokens get even more likely. Output becomes steadier and more predictable.
- At 1, the probabilities are used as the model gave them.
- Above 1, the odds flatten out, so less likely tokens get picked more often. Output becomes more surprising, and at high values it can drift into nonsense.
Picking a value
These are rough starting points, not rules:
| Task | Example temperature |
|---|---|
| Code, data extraction, factual answers | 0 to 0.3 |
| General chat and writing | 0.5 to 0.8 |
| Brainstorming and creative writing | 0.9 to 1.2 |
Some models, especially Reasoning Models, have their own recommended settings or ignore temperature. Test on your own prompts.
Example
curl https://api.luv13.ai/v1/chat/completions \
-H "Authorization: Bearer $LUV13_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "luv13/glm-5.3-flash",
"messages": [{"role": "user", "content": "Suggest a name for a coffee shop."}],
"temperature": 1.0
}'Run it a few times, then try 0.2, and compare how much the answers change.
