How it works
- Collect examples. Hundreds to thousands of input and output pairs that show exactly what you want.
- Train. Start from a base model and train it a bit more on your examples.
- Evaluate. Compare the tuned model with the original on test cases it didn't train on.
- Serve. Host the new weights so you can call them.
Lighter methods such as LoRA train a small add-on instead of every weight. They need far less memory and are common with Open-Weight Models.
Fine-tuning vs. other options
| Goal | Usually best |
|---|---|
| Answer from your documents or fresh facts | Retrieval-augmented generation |
| Follow a format a few times | Few-shot examples in the prompt |
| Match a narrow style every time, at scale | Fine-tuning |
| Teach new facts | Retrieval, not fine-tuning. Tuning is poor at adding reliable facts. |
On luv13
luv13 serves the models in its live list at https://api.luv13.ai/v1/models as they are. To steer them, use prompts and examples. See Endpoints for what luv13 serves.
Example
A few-shot prompt is often enough to get the style you'd otherwise fine-tune for:
curl https://api.luv13.ai/v1/chat/completions \
-H "Authorization: Bearer $LUV13_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "luv13/glm-5.3-flash",
"messages": [
{"role": "system", "content": "Rewrite product names in our house style: all lowercase, words joined by dots."},
{"role": "user", "content": "Blue Water Bottle"},
{"role": "assistant", "content": "blue.water.bottle"},
{"role": "user", "content": "Travel Coffee Mug"}
]
}'