Common routing rules
- By task. Short classification or formatting goes to a fast model. Hard reasoning or long code goes to a larger one.
- By length. Very long prompts go to a model with more room. luv13's list doesn't publish context lengths, so test before relying on this. See Context Window.
- By result. Try a fast model first. If the answer fails a check (bad JSON, failed test), retry on a stronger model.
- By availability. If one model errors out, switch to another. See Retrying Requests.
On luv13
All luv13 models share one base URL, one key and one flat price per token (see Pricing), so routing on luv13 is about speed and quality rather than cost per token. The ids come from the live list at https://api.luv13.ai/v1/models. See Model IDs.
Example
A tiny router in Python. The rule is only an example. Set LUV13_STRONG_MODEL to another id from the live list that you've tested, or leave it unset to use the same model for both routes.
import os
from openai import OpenAI
client = OpenAI(base_url="https://api.luv13.ai/v1",
api_key=os.environ["LUV13_API_KEY"])
def pick_model(task: str) -> str:
if task in ("classify", "extract", "rewrite"):
return "luv13/glm-5.3-flash"
return os.environ.get("LUV13_STRONG_MODEL", "luv13/glm-5.3-flash")
resp = client.chat.completions.create(
model=pick_model("classify"),
messages=[{"role": "user", "content": "Is this spam? 'You won a free cruise!' Reply yes or no."}],
)
print(resp.choices[0].message.content)Related
This page covers the general idea. For how luv13 handles a request when a model is unavailable, see Model Fallback.
