# AI embeddings

Turn text into vectors for search and similarity.

Base URL: `https://ai.inovacc.dev`

## Overview

AI embeddings turns text into vectors: lists of numbers whose distance from each other follows the meaning of the text. Two passages about the same thing land close together even when they share no words. The endpoint, `POST /v1/embeddings`, speaks the OpenAI embeddings wire format, so an existing OpenAI client can call it with the base URL `https://ai.inovacc.dev/v1` and two extra headers.

The problem it solves is semantic comparison inside your own systems. With vectors you can search by meaning, group similar items, find near-duplicates or route a message to the closest category, using whatever database or index you already run.

Use AI embeddings when **you** keep the vectors. When you would rather have Inovacc keep them, search them and return the matching passages, use the managed Vector Database of [Knowledge and vector search](/en/products/knowledge/), which embeds your documents for you. When you need generated text rather than vectors, use [AI chat](/en/products/ai.chat/).

## Concepts

**Routes.** As with chat, you never name a model. Your organization is given embedding routes, opaque ids configured by Inovacc. `GET /v1/models` lists them: an entry whose `endpoint` is `/v1/embeddings` is an embedding route. Send its id as `model`, or send `"auto"` (or no `model`) to use your organization's default route. A route serves one endpoint only; a chat route called here is refused.

**Input.** `input` is one non-empty string, or an array of strings; inside an array an empty string is accepted. The answer holds one vector per input, in `data[]`, each with the `index` of the input it belongs to.

**Vectors of one route belong together.** A vector is only comparable with vectors produced by the same route. Store the route id next to every vector you keep, and embed both your documents and your queries with the same route.

**Optional fields.** `encoding_format` and `dimensions` are passed to the model where the route's model supports them and ignored otherwise. `user` is accepted as a string and replaced, before it leaves Inovacc, by a value derived from your organization, so two organizations sending the same value never collide.

**Usage.** The answer's `usage` holds `prompt_tokens` and `total_tokens`. This endpoint carries no `x-usage-*` headers; the route that answered is in `x-route-id`.

## How it works

Every call carries `Authorization: Bearer <key>` and an `X-Operation-Id`. Your organization comes from the key. The service checks the operation id and the key, then whether that operation id was already used, then the body: only `model`, `input`, `encoding_format`, `dimensions` and `user` are accepted, and any other field is refused by name. It resolves the route, checks that the route serves embeddings, validates `input`, estimates the input size (a quarter of the characters of every input string) against your organization's per-request cap, and reserves the call against your organization's quotas. Then the model is called.

The answer is `{"object":"list","data":[{"object":"embedding","index":0,"embedding":[...]}],"model":"<route id>","usage":{...}}`. `model` is always the route id.

There is no streaming. A successful answer is stored for 24 hours under its operation id: the same id sent again returns the same vectors with `x-idempotent-replay: true`, without a second call or charge. An error is never stored, so a failed call can be retried with the same id. What is metered is input tokens.

## Get started

You need an API key and an embedding route enabled for your organization ([Authentication](/en/guides/authentication/)).

1. **Find your embedding route.** Call `GET /v1/models` and pick an entry whose `endpoint` is `/v1/embeddings`.
2. **Embed two passages.** Call `POST /v1/embeddings` with that route as `model`, an `input` array of two strings and a fresh `X-Operation-Id` (see [the samples](#example)). The answer has two entries in `data`, with `index` 0 and 1, and `model` is your route id.
3. **Compare them.** Compute the cosine similarity of the two vectors in your code. Embed a third passage on another subject and compare again: its score against the first is lower.
4. **Replay.** Send step 2 again with the same operation id; the vectors are identical and the answer carries `x-idempotent-replay: true`.

## Use cases

**Search in your own database.** A product catalogue stores one vector per product description in the database it already uses. A shopper's question is embedded with the same route at query time, and the nearest products are shown, including those whose description uses different words.

**Near-duplicate detection.** A ticketing system embeds each new ticket and compares it with the open ones; a score above a threshold you calibrate links the ticket to the existing case instead of opening a second one.

**Routing by similarity.** A small set of example messages per team is embedded once. Each incoming message is embedded and sent to the team whose examples are closest, with no model call per message beyond the embedding.

## Limits and pricing

| Limit | Value |
|---|---|
| Request body | 1 MiB by default; your organization may be set between 1 KiB and 20 MiB |
| Input tokens per request | your organization's cap (`400 input_too_large` above it) |
| Time to the model's first response | 60 seconds |
| Stored answer for a replay | 24 hours |
| Requests, tokens, cost, concurrency, daily budgets | as set for your organization |

**Pricing:** on request. The pricing unit is **tokens**.

## Errors

| Status | Code | What it means and what to do |
|---|---|---|
| 400 | `missing_operation_id`, `invalid_operation_id` | Send a valid `X-Operation-Id`. |
| 400 | `unsupported_field` | A field other than `model`, `input`, `encoding_format`, `dimensions`, `user`; remove it. |
| 400 | `invalid_body` | Malformed JSON, `input` of the wrong shape, or a non-string `user`. |
| 400 | `model_not_allowed` | `model` is not a valid route id. |
| 400 | `input_too_large` | The input is over your per-request cap; split it into several calls. |
| 401 | `missing_credentials`, `invalid_credentials` | Send a valid key. |
| 403 | `route_forbidden` | The route is not yours, is disabled, or does not serve embeddings. |
| 403 | `key_disabled`, `organization_disabled` | Check the key. |
| 409 | `operation_in_progress` | That operation id is still running; wait and retry. |
| 413 | `request_too_large` | The body is over your size cap. |
| 429 | `rate_limited`, `budget_exhausted` | Wait the `Retry-After` seconds. |
| 429 | `quota_exceeded`, `concurrency_limited` | An allowance is used up; no `Retry-After`. |
| 502, 503, 504 | `upstream_error`, `service_unavailable`, `timeout_error` | Retry with the same operation id and backoff. |

See [Errors](/en/guides/errors/) for the envelope.

## Best practices

- **Batch inputs**: send many passages in one `input` array, within your size and token caps, rather than one call per passage.
- **Keep the route id with the vector**, and re-embed everything if you change route: vectors of two routes are not comparable.
- **Chunk long documents** into passages of a few paragraphs before embedding; one vector for a whole document blurs its meaning.
- **Derive the operation id from the content** (for example a hash of the batch) so a restarted job replays instead of paying again.
- **Retry `502`, `503` and `504` with the same operation id**; wait `Retry-After` on `429 rate_limited`.
- **Calibrate thresholds on your own data**: similarity scores are relative, not probabilities.

## Related

- [Knowledge and vector search](/en/products/knowledge/): the managed Vector Database that embeds and searches for you
- [AI chat](/en/products/ai.chat/)
- [Authentication](/en/guides/authentication/), [Errors](/en/guides/errors/), [Limits](/en/guides/limits/)

## Endpoints

- POST /v1/embeddings

## Example

```sh
curl -X POST "https://ai.inovacc.dev/v1/embeddings" -H "Authorization: Bearer $INOVACC_API_KEY" -H "X-Operation-Id: $(uuidgen)" -H "Content-Type: application/json" -d '{}'
```
