What counts toward it
Every token in the request counts, including:
- The system message and every user and assistant message you include
- Tool or function definitions and tool results
- The tokens the model writes in its reply
So if a model's window is 100,000 tokens and your prompt uses 95,000, only about 5,000 are left for the answer. (These numbers are only an example. Check your model's real limit.)
Why it matters
A bigger window lets you send longer documents, more code or a longer conversation in one go. But more tokens in means a higher cost per request, since you pay for every input token. See What Is a Token.
Staying inside the limit
- Trim old messages. Drop or summarize the earliest turns of a long chat.
- Send only what's needed. Paste the relevant section of a file, not the whole thing.
- Cap the reply. Set
max_tokensso the reply can't use more room than you've left for it.
On luv13
Context windows depend on the model. luv13's model list at https://api.luv13.ai/v1/models returns only id, object, created and owned_by for each model. It doesn't report a context length, so this page gives no per-model numbers. See Listing Models for the model list itself.
