Docs
Open-weight models behind one OpenAI-compatible API.
Point any tool that accepts an OpenAI base URL at luv13 and use one of our model IDs.
https://api.luv13.ai/v1Keys
Get a key
Sign in to the dashboard with Google or email, then create a key. The full key starts with sk-luv13- and is shown once, at creation. Keep it in an environment variable, never in client-side code.
Can’t use the dashboard? Email [email protected].
Price
Price
One flat price: $0.33 / 1M tokens on every model. Input costs the same as output.
800k input + 200k output = 1.0M tokens = $0.33.
Top up any amount from $5 in your dashboard; usage draws down the balance. See all models.
First request
Your first request
Export your key as LUV13_API_KEY, then:
curl https://api.luv13.ai/v1/chat/completions \
-H "Authorization: Bearer $LUV13_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "luv13/glm-5.3-flash",
"messages": [
{"role": "user", "content": "ping"}
]
}'Limits
Limits and status
luv13 runs on single-provider capacity. When capacity is saturated, requests queue or fail. Retry with backoff rather than hammering. If a model is unavailable, the error names the model; switch to another ID.
Balance and recent usage are in the dashboard.
Library
A–Z
Every docs page, alphabetically. Pick a letter, or scan the lists below. Greyed letters have no pages yet.
A
An AI agent is a program that lets a model work toward a goal over several steps, choosing and using tools and checking the results as it goes.
Aider is an open-source AI pair-programming tool for the terminal that can use luv13 as an OpenAI-compatible endpoint.
API key best practices are the habits that keep your luv13 key, which starts with sk-luv13- and spends your prepaid credit, from leaking or being misused.
Authentication on luv13 means sending your sk-luv13- API key as a Bearer token in the Authorization header of each chat completion request.
B
A base URL is the fixed start of an API's address that every endpoint path is added to.
Browser requests are calls to luv13 made from JavaScript running in a web page, which luv13 blocks for other sites' origins, so they should go through your own server.
C
Chain-of-thought prompting means asking a model to work through a problem step by step before it gives the final answer.
Chat completions is the luv13 endpoint that takes a list of messages and returns the model's next reply.
A chat template is the fixed text format a model uses to turn a list of chat messages into the single token sequence it was trained on.
Compatible tools are apps and coding agents that accept an OpenAI-style base URL and key, and so can use luv13 models once pointed at https://api.luv13.ai/v1.
Contact and support covers how to reach the luv13 team, by email at [email protected] or the form on luv13.ai, and what to include so a problem can be traced.
A context window is the most tokens a model can handle in one request, counting both what you send and what it writes back.
Conversation history is the list of earlier messages you send with each request so a model can follow an ongoing chat.
curl examples are ready-to-run terminal commands for luv13's two endpoints, listing models and sending a chat completion.
D
The luv13 dashboard at luv13.ai/dashboard is where you sign in, create API keys, top up prepaid credit and see your balance and recent usage.
Data and privacy covers what information luv13 says it handles when you use the API, and where its privacy notice and terms stand.
E
An embedding is a list of numbers that represents the meaning of a piece of text, so that similar texts get similar numbers.
luv13's endpoints are the two URL paths under https://api.luv13.ai/v1 that it serves, one to list models and one to create chat completions.
An environment variable is a named value set outside your code, such as an API key, that your program reads when it runs.
Errors and status codes are the HTTP codes and bodies luv13 returns when a request can't be served, and what each one means you should do.
Estimating costs means turning luv13 token counts into dollars with one multiplication, because every model has the same flat rate and input costs the same as output.
F
The luv13 FAQ answers the questions people ask most about the API, each in a sentence or two, using only facts luv13 has published or that were checked live.
Few-shot prompting means showing a model a few worked examples in the prompt so it copies the pattern for a new input.
Fine-tuning is further training of an existing model on your own examples so it learns a specific task, style or format.
G
The luv13 glossary defines the terms, ids and fields you meet when using the luv13 API, each in one line.
H
A hallucination is when a model states something false or made up as if it were true.
A luv13 health check is a quick request, usually GET /v1/models, that tells you whether the API is reachable before you debug your own code.
HTTP headers are the name-value lines sent with each luv13 request and response; you need two on requests, Authorization and Content-Type.
I
Image input means sending a picture to a luv13 model that accepts images; five of the seven models list image input, and all return text.
Input tokens are the tokens you send to a model, and output tokens are the tokens it writes back.
J
The JavaScript example is a short Node.js script that calls luv13 with the official OpenAI JavaScript SDK.
JSON mode is a request option that tells a model to reply with valid JSON instead of free text.
K
A luv13 account is where you create API keys, which start with sk-luv13-, and hold the prepaid credit that your requests spend.
L
LangChain is a framework for building LLM apps whose ChatOpenAI class can call luv13 by setting base_url.
Latency is how long you wait for a model's response, often measured as the time to the first token and the time to the full reply.
Listing models means calling luv13's GET /v1/models endpoint to see every model id you can use right now.
M
DeepSeek V4-Pro is DeepSeek's large MIT-licensed Mixture-of-Experts text model, listed on luv13 as luv13/deepseek-v4-pro and temporarily unavailable.
DeepSeek V4.1 Flash is DeepSeek's MIT-licensed multimodal Mixture-of-Experts model, available on luv13 as luv13/deepseek-v4.1-flash.
GLM 5.3 is Z.ai's flagship open-weight coding and agent model, available on luv13 as luv13/glm-5.3.
GLM-5.3 Flash is Z.ai's natively multimodal, MIT-licensed GLM-5 model, available on luv13 as luv13/glm-5.3-flash.
Kimi K3 is Moonshot AI's open-weight multimodal model for long coding and agent work, available on luv13 as luv13/kimi-k3.
Kimi K3 Fast is a model id luv13 lists as luv13/kimi-k3-fast, named after Moonshot AI's Kimi K3.
Migrating from OpenAI means moving code that calls the OpenAI API over to luv13 by changing the base URL, the API key and the model id.
Model fallback means trying a second luv13 model id when the first one is unavailable, instead of failing or waiting.
A model ID is the exact string, such as luv13/kimi-k3, that you put in the model field of a luv13 request to choose which model answers.
Model routing means sending each request to the model best suited for it, based on rules like task type, cost or speed.
luv13 is an OpenAI-compatible API for open-weight models at base URL https://api.luv13.ai/v1, and GET https://api.luv13.ai/v1/models returns the exact model ids you pass in the model field. Every model costs one flat $0.33 per 1M tokens, with input priced the same as output.
Qwen 3.8 27B is Alibaba's Apache-2.0 dense vision-language model from the Qwen3.8 series, available on luv13 as luv13/qwen-3.8-27b.
N
n8n is a workflow automation tool whose OpenAI credential has a Base URL field, so its OpenAI nodes can call luv13.
Node.js has a built-in fetch function that can call luv13's OpenAI-compatible API with no extra packages.
Nucleus sampling, set with top_p, makes a model pick each next token only from the smallest group of likely tokens whose probabilities add up to a set share.
O
An open-weight model is a language model whose trained weights are published, so anyone allowed by its license can download and run it.
An OpenAI-compatible API accepts the same requests and returns the same response shapes as OpenAI's API, so existing tools and code work with it after changing the base URL and key.
OpenCode is an open-source AI coding agent for the terminal that can use luv13 as a custom OpenAI-compatible provider.
P
Prompt engineering is the practice of writing and testing the instructions you give a model so it reliably produces the output you want.
Prompt injection is when text from an untrusted source, such as a web page, email or user message, contains instructions that trick a model into ignoring its real ones.
The Python example is a short script that calls luv13 with the official OpenAI Python SDK.
The Python requests library can call luv13's OpenAI-compatible API directly with plain HTTP, without an SDK.
Q
Quantization shrinks a model by storing its weights with fewer bits, which saves memory and speeds it up at some cost to accuracy.
Get a luv13 key, set base URL https://api.luv13.ai/v1, and make your first OpenAI-compatible chat completion with one curl.
R
Rate limiting is when an API caps how many requests or tokens you can use in a period of time, and rejects extra ones until the window resets.
A reasoning model is a language model trained to work through a problem in intermediate steps before it gives its final answer.
Retrieval-augmented generation (RAG) means looking up relevant text first and adding it to the prompt so the model answers from that source.
Retrying requests means sending a failed luv13 call again after a growing wait, and only for errors that can succeed on a second try.
Roo Code is an AI coding agent for VS Code that can use luv13 through its OpenAI Compatible provider.
S
Server-Sent Events (SSE) is a simple web standard for a server to push a stream of text messages to a client over one open HTTP connection.
Maker and product URLs cited by public luv13 docs. Gateway catalogs and supplier pages are not listed here.
A stop sequence is a string that tells the model to stop writing as soon as it would produce that text.
Streaming means the API sends a model's reply in small pieces as it is written, instead of all at once at the end.
A system prompt is an instruction at the start of a conversation that sets how the model should behave for every reply that follows.
T
Temperature is a sampling setting that controls how random a model's word choices are.
Throughput is how much work a model or API gets done over time, usually measured in output tokens per second.
A timeout is the longest your code will wait for a request to finish before it gives up.
A tokenizer is the part of a language model system that splits text into tokens and turns them into numbers the model can read.
Tool calling lets a model ask your code to run a function you described, then use the result in its reply.
Setup in coding tools: where to paste the base URL and key in Cursor, VS Code, Cline, Open WebUI, Codex, Hermes, Kilo Code and the OpenAI SDK.
Troubleshooting is a symptom-to-fix list for the most common problems when calling luv13, based on the responses the API actually returns.
U
Usage and billing is how luv13 charges the tokens your requests use against your prepaid credit, at a flat $0.33 per 1M tokens.
Using Claude Code with luv13 isn't possible today, because Claude Code needs an Anthropic-format API and luv13 serves only OpenAI-style chat completions.
Cline is an open-source AI coding agent for VS Code and other editors that can use luv13 through its OpenAI Compatible provider.
Using Codex with luv13 would mean setting luv13 as a custom model provider in OpenAI's Codex coding agent.
Continue is an open-source AI coding assistant for VS Code and JetBrains that can use luv13 through its openai provider with a custom apiBase.
Using Cursor with luv13 means pointing Cursor's OpenAI API key and base URL override at luv13 so local Chat and Agent run on a luv13 model.
Using Hermes Agent with luv13 means pointing Nous Research's Hermes Agent at luv13 through its custom OpenAI-compatible provider.
Using Kilo Code with luv13 means adding luv13 as an OpenAI Compatible custom provider in the Kilo Code agent.
Using Open WebUI with luv13 means adding luv13 as an OpenAI API connection so Open WebUI's chat can use luv13 models.
The official OpenAI SDKs for Python and JavaScript can call luv13 by setting their base URL to https://api.luv13.ai/v1 and using a luv13 API key.
Using VS Code with luv13 means adding luv13 to VS Code's chat as a Custom Endpoint model that uses the Chat Completions API.
V
The Vercel AI SDK is a TypeScript library for building AI features that can call luv13 through its OpenAI Compatible provider package.
Vibe coding is building software mostly by describing what you want to an AI coding tool and accepting its changes, with little reading of the code yourself.
W
A token is a small chunk of text, often a word or part of a word, that a language model reads and writes one at a time.
luv13 is an OpenAI-compatible API that serves seven open-weight models from one base URL and one API key, at one flat price per token.
X
An XML prompt uses simple XML-style tags to separate the parts of a prompt, such as instructions, documents and examples.
Y
A YAML config is a settings file written in YAML, a plain-text format that many AI tools use to store provider, model and key settings.
Z
Zed is a code editor with built-in AI features that can use luv13 as an OpenAI-compatible provider.
Zero-shot prompting means asking a model to do a task with instructions only, without giving it any examples.
