How they're used
- Run each document chunk through an embedding model and store the vectors.
- When a query comes in, embed it with the same model.
- Find the stored vectors closest to the query vector.
- Use those chunks, for example as context in a chat prompt.
Vectors from different embedding models aren't compatible. Embed your documents and your queries with the same model.
Cosine similarity
Cosine similarity compares the direction of two vectors. It ranges from -1 to 1, and higher means more alike. In practice you compare scores against each other, not against a fixed cutoff, because the typical range depends on the model.
Using embeddings with luv13
Because luv13 doesn't offer embeddings, you'd create them with a separate embedding model or service, then send the matching text to a luv13 chat model. See Retrieval-Augmented Generation for the full pattern and Listing Models for what luv13 serves today.
You can check the current status yourself. A 404 means the endpoint isn't there:
curl -i https://api.luv13.ai/v1/embeddings \
-H "Authorization: Bearer $LUV13_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "luv13/glm-5.3-flash", "input": "hello"}'