How it works
At each step the model ranks every possible next token by probability. Nucleus sampling:
- Sorts the tokens from most to least likely.
- Adds up their probabilities until the total reaches
top_p. - Throws away everything below that line (the long tail).
- Picks the next token at random from what's left, weighted by probability.
The group that's left is the "nucleus". When the model is confident, the nucleus may be one or two tokens. When it's unsure, the nucleus is bigger. That's the main difference from a fixed "top-k" cutoff, which always keeps the same number of tokens.
Picking a value
As an example, many apps leave top_p at 1 and tune temperature instead. If you do use it, values from 0.8 to 0.95 trim the oddest word choices while still leaving room for variety.
On luv13
Check Request Parameters to confirm whether luv13 passes top_p through for your model.
Example
curl https://api.luv13.ai/v1/chat/completions \
-H "Authorization: Bearer $LUV13_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "luv13/glm-5.3-flash",
"messages": [{"role": "user", "content": "Write one line about the desert."}],
"top_p": 0.9
}'