Using Kilo Code

Definition

Using Kilo Code with luv13 means adding luv13 as an OpenAI Compatible custom provider in the Kilo Code agent.

Key takeaways

  • Go to Settings (gear) > Providers > Custom provider. Set Provider API to OpenAI Compatible, Base URL to https://api.luv13.ai/v1, add your luv13 key, and pick luv13/glm-5.3-flash from the list Kilo fetches.
  • Don't pick OpenAI Responses or Anthropic Messages. luv13 returns 404 for /v1/responses and /v1/messages.
  • For tool use, token limits and other model options, edit kilo.jsonc. Keep the key in your global config (~/.config/kilo/kilo.jsonc), because Kilo ignores {env:...} in a project's config file.
  • Set tool_call: true if you want Kilo's agent to edit files and run commands. Whether each luv13 model handles tools well isn't confirmed.

Compatibility: works, with setup caveats. Kilo Code's Custom provider with Provider API OpenAI Compatible is built for Chat Completions endpoints like luv13's. Kilo can list luv13's models automatically. Two catches: luv13 doesn't publish context or output limits, and a custom model without limits turns off Kilo's automatic context compaction. Tool calling also isn't confirmed per luv13 model.

Before you start

You need:

  • Kilo Code (the VS Code extension or the Kilo CLI). These steps follow Kilo's "OpenAI Compatible" and "Custom Models" docs, read 2026-09-30. The newest tag in the Kilo-Org/kilocode repo that day was v7.8.1.
  • A luv13 API key. The Quickstart shows how to get one.
  • The model id luv13/glm-5.3-flash, from the live list at https://api.luv13.ai/v1/models. See GLM-5.3 Flash for details on the model.

Facts about luv13 that affect this setup:

  • luv13 serves one generation endpoint, POST /v1/chat/completions, plus GET /v1/models. /v1/responses, /v1/messages, /v1/completions and /v1/embeddings return 404 (checked 2026-09-30). See Endpoints.
  • Every model costs a flat $0.33 per 1M tokens, input the same as output. See Pricing.
  • The model list at https://api.luv13.ai/v1/models doesn't report context length, so luv13 publishes no per-model limits. See the model list.

Test your key and model first

Run these two checks in a terminal before you touch the tool. They take a few seconds and rule out key and model problems.

1. The model id exists. This call needs no key:

curl -s https://api.luv13.ai/v1/models

The list should include "id":"luv13/glm-5.3-flash".

2. Your key works for chat. Set the key in your shell first (export LUV13_API_KEY="your luv13 key"), then run:

curl -s https://api.luv13.ai/v1/chat/completions \
  -H "Authorization: Bearer $LUV13_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "luv13/glm-5.3-flash",
    "messages": [{"role": "user", "content": "Reply with the word ready."}],
    "max_tokens": 20
  }'

With a valid key you should get back a JSON chat completion whose choices[0].message.content holds the reply, plus a usage block. If you see {"error":{"code":401,"message":"unauthorized","type":"invalid_auth"}} instead, the key is missing or wrong, and no tool setting will fix that.

Option A: add luv13 in Settings

  1. Open Settings (the gear icon) and go to the Providers tab.
  2. Scroll to the bottom and click Custom provider.
  3. Fill in the dialog:
FieldWhat to enter
Provider IDluv13. Kilo's docs allow lowercase letters, numbers, hyphens and underscores.
Display nameluv13
Provider APIOpenAI Compatible
Base URLhttps://api.luv13.ai/v1
API keyyour luv13 key
Modelsonce the base URL is in, Kilo fetches luv13's models from /v1/models. Choose luv13/glm-5.3-flash, or add it by hand.
Headersleave empty
  1. Click Submit. luv13's models then show up in the model picker.
  2. Pick luv13/glm-5.3-flash and ask something small, like "Reply with the word ready."

To change the provider later, click Edit provider next to it in the connected providers section.

Option B: kilo.jsonc

Kilo's docs say token limits, tool calling and other model options go in kilo.jsonc. The global config is ~/.config/kilo/kilo.jsonc. On Windows it's C:\Users\<you>\.config\kilo\kilo.jsonc.

This follows the "OpenAI-compatible provider with a custom endpoint" example in Kilo's docs. It uses the documented id field so the model key doesn't need a slash, while the id sent to luv13 stays luv13/glm-5.3-flash:

{
  "$schema": "https://app.kilo.ai/config.json",
  "model": "openai-compatible/glm-5.3-flash",
  "provider": {
    "openai-compatible": {
      "options": {
        "apiKey": "{env:LUV13_API_KEY}",
        "baseURL": "https://api.luv13.ai/v1"
      },
      "models": {
        "glm-5.3-flash": {
          "id": "luv13/glm-5.3-flash",
          "name": "GLM-5.3 Flash (luv13)",
          "tool_call": true
        }
      }
    }
  }
}

Set LUV13_API_KEY in your environment before starting Kilo. Then run kilo models to check that the model is listed.

About {env:LUV13_API_KEY}: Kilo's docs say {env:...} only resolves in trusted config, meaning your global config, KILO_CONFIG or KILO_CONFIG_CONTENT, or managed config. In a project kilo.jsonc committed to a repo, it's ignored, so the provider won't authenticate. Kilo does this so a repo can't send your keys to a baseURL of its choosing.

Token limits

Kilo's docs say that a custom model with no limit and no match in its catalog ends up with context and output set to 0. When that happens:

  • Automatic compaction is off, so conversations grow until the provider rejects them.
  • The output cap falls back to 32,000 tokens and is sent as max_tokens.
  • Context usage isn't tracked.

luv13 doesn't publish per-model context or output limits, and /v1/models doesn't include them. So this guide doesn't give numbers. If you want compaction, ask luv13 support for the model's real limits and add them as "limit": { "context": ..., "output": ... }. Don't guess.

What works and what doesn't

FeatureWith luv13
Custom provider with OpenAI CompatibleWorks. It uses /v1/chat/completions.
Automatic model detectionWorks. luv13 serves /v1/models.
OpenAI Responses or Anthropic Messages Provider APINo. Both endpoints return 404 on luv13.
Agent tools (file edits, terminal)Needs tool_call: true. Not confirmed per luv13 model. See Tool Calling on luv13.
Automatic context compactionOnly after you set limit.context. See Token limits.
Voice transcription via a custom base URLNo. luv13 has no /audio/transcriptions.

Common errors

SymptomLikely causeFix
401 with {"error":{"code":401,"message":"unauthorized","type":"invalid_auth"}}The key is missing, has extra spaces, or isn't a luv13 key.Paste your luv13 key again. Run test 2 above to confirm it.
404 with an HTML page titled "404 Not Found"The URL is wrong. Common versions: https://api.luv13.ai with no /v1, a doubled path such as https://api.luv13.ai/v1/v1/..., a pasted full endpoint like .../v1/chat/completions in the base URL field, or a trailing slash.Set the base URL to exactly https://api.luv13.ai/v1. The tool adds /chat/completions itself.
404 even though the base URL is rightThe tool is calling an endpoint luv13 doesn't serve, such as /v1/responses, /v1/messages, /v1/completions or /v1/embeddings.Use the tool's OpenAI chat-completions mode, and turn off features that need other endpoints. See Endpoints.
Model not found, or the model doesn't appearThe model id is misspelled or missing the luv13/ prefix.Use the exact id luv13/glm-5.3-flash from https://api.luv13.ai/v1/models.
"Invalid API Key"Kilo's troubleshooting section lists it. The key is wrong, or {env:LUV13_API_KEY} didn't resolve.Re-enter the key. If you used {env:...}, move the provider to your global config and check the variable is set in the shell that starts Kilo.
"Model Not Found"Kilo's docs list it for an invalid model id.Make sure the id sent is exactly luv13/glm-5.3-flash (the id field in Option B).
The model doesn't appear in the pickerKilo's docs say to check credentials and that "model" matches provider/model-key.In Option B, the value is openai-compatible/glm-5.3-flash. Run kilo models to see what's active.
The conversation keeps growing and never compactsKilo's docs say this means limit.context is 0 (unset).Add a limit block once you have confirmed limits. See Token limits.
The model writes tool calls as text or never edits filesKilo's docs say to set tool_call: true for tool use.Set it. If it still fails, try another luv13 model.

General tip: put the base URL https://api.luv13.ai/v1 in the Base URL field, not the full /v1/chat/completions URL. Kilo's docs say full endpoint URLs are supported, but Kilo issue #2035 reported /chat/completions being appended twice. The issue was closed in November 2025.

Sources

Related: Using Cline, Using VS Code, Using Claude Code, Compatible Tools.