# Knowledge and vector search

Ingest documents, build knowledge collections and query them by meaning.

Base URL: `https://ai.inovacc.dev`

## Overview

Knowledge and vector search keeps your organization's knowledge where an application can ask it questions by meaning. It has three parts, all on `https://ai.inovacc.dev`:

- the **Vector Database**: a managed store of text passages you ingest and search, with filters, scopes and optional reranking;
- **knowledge collections**: send a batch of mixed files (documents, images, video) and get them processed into knowledge packages, then download the result, query it, chat with it with citations, or have a report written over it;
- **knowledge jobs**: turn a long recording into a knowledge base you download as one archive.

The problem it solves is everything between "we have files" and "the assistant can answer from them": extracting text, cutting it into passages, embedding, indexing, keeping the index consistent, and citing the source of every answer. Inovacc does those steps and you call a handful of endpoints.

Use it when answers must come from your own material. When you already run your own vector index, [AI embeddings](/en/products/ai.embeddings/) gives you the vectors alone. When the answer is free text with no sources, use [AI chat](/en/products/ai.chat/). Each part is enabled for your organization by Inovacc; a part that is not enabled answers `403` (`vector_database_disabled`, `collections_disabled` or `knowledge_disabled`).

## Concepts

**Vector Database and its profile.** Your organization creates one database with `POST /v1/vector/databases`, choosing an **embedding profile** from `GET /v1/vector/profiles`. A profile states the modality (`text`), the languages (`english` or `multilingual`), the dimensions and the distance metric; ids look like `text-multilingual-1024`. You choose a profile, never a model. The profile cannot change: to use another, delete the database, create a new one and ingest again.

**Documents, scopes and modes.** You ingest passages with your own `id`, their `text` and optional flat `metadata`. A passage goes to a **scope**: `shared` (every service of your organization's workspace) or `local` (only the API key that wrote it). Queries take a **mode**: `shared`, `local`, or `hybrid` to search both and merge by score. Six metadata keys can be filtered on: `kind`, `source`, `category`, `tag`, `language`, `created_at`. Chunks that share a `document_id` form a document with a `generation` (your version number), which lets you delete or replace a whole document at once.

**Collections and their mode.** A collection is created in one of three modes, fixed for its life: `download` (process files, download the packages, nothing stays afterwards), `cloud` (Inovacc keeps the processed content and its vectors so you can `query`, `chat` and run an `analysis`) or `local_vectors` (Inovacc keeps vectors only, never text: you `embed` your own chunks and `query` returns chunk ids and scores). Collections are kept 30 days by default; an organization set to `zero` retention has them deleted one hour after the download starts.

**Items.** A file joins a collection in the request (up to 16 MiB), from a completed upload (video above 16 MiB), by public URL, or by a signed URL sealed to a single-use key so that it is never stored. Each item goes `queued`, `processing`, then `done` or `failed`; a failed item does not stop the others.

**Eventual consistency and `searchable`.** The vector index applies writes asynchronously: a vector is readable some seconds, sometimes minutes, after it was written. `GET /v1/vector/databases` and every collection answer report `searchable`; a query sent while it is `false` may miss the latest passages. A `cloud` collection's `indexing.complete` is `true` only when every chunk is written **and** readable.

**Jobs, power and budget.** A knowledge job takes a completed upload of a recording, an optional transcript and notes, a `power` (`light`, `medium` or `max`) and a required `max_price_usd`. A job over your organization's cap is refused before any paid work.

## How it works

Every call carries `Authorization: Bearer <key>`. Your organization comes from the key. `X-Operation-Id` is required on the collection, upload and job `POST` calls and makes them idempotent: a retry with the same id returns the first answer (`X-Idempotent-Replay: true`) and adds nothing; the same id with a different file is `409 operation_id_reused`. On the Vector Database it is optional, because ingest and delete are idempotent by construction: ingesting an existing `id` replaces it.

A **query** (`POST /v1/vector/query`) takes `mode`, `query` (1 to 8,000 characters), `top_k` (1 to 50) and an optional `filter`; it returns matches best first, each with `id`, `score`, `text` and `metadata`. With `"retrieval": "vector_rerank"` and a `rerank_profile`, a wider candidate list is re-ordered by a reranking stage; the answer's `retrieval` object says whether the rerank applied or fell back to vector order.

A **collection** works in steps you can observe: create it, add items (`202`, `queued`), read it until `status` is `complete`, then download, query or chat. Reading a `cloud` collection also advances its indexing, a bounded amount per call, and every answer carries the `indexing` progress. `chat` answers only from the retrieved sources, numbers each citation with the item, package, chunk and locator, and says it cannot answer when the sources do not cover the question. An `analysis` is built over several requests: repeat `POST` or `GET` every few seconds until it answers `200` with the report.

A **job** is asynchronous: `POST /v1/knowledge/jobs` answers `202` with `job_id`, `status_url` and `poll_after_seconds` (also as `Retry-After`). Read the status when told; the platform forecasts the time to done and includes `eta_seconds` while it can. When `state` is `done`, download the archive; it is kept 24 hours.

What is metered: embedding tokens and vectors written for ingest and indexing, queries, reranks as their own unit, chat calls on your chat route, and jobs by their price. `POST /v1/knowledge/estimates` prices a recording before you upload it, at no cost.

## Get started

You need an API key and the part you want enabled for your organization ([Authentication](/en/guides/authentication/)).

1. **Create the database.** `GET /v1/vector/profiles`, pick a profile, then `POST /v1/vector/databases` with `{"profile":"<id>"}`. The answer is `201` with the database.
2. **Ingest passages.** `POST /v1/vector/ingest` with `"scope":"shared"` and two or three documents (see [the samples](#example)). The answer says how many were `ingested` and lists any `skipped` with a reason.
3. **Wait for `searchable`.** `GET /v1/vector/databases` until `searchable` is `true`.
4. **Query.** `POST /v1/vector/query` with `"mode":"shared"` and a question worded differently from your passages. The closest passage comes first, with its `score`.
5. **Try a collection.** `POST /v1/collections` with `{"mode":"cloud"}`, add a PDF to `/items`, read the collection until `indexing.complete` is `true`, then `POST /v1/collections/{id}/chat` and check the `citations`.

## Use cases

**An assistant grounded in your handbook.** Your intranet ingests each policy page as chunks under its `document_id`, with `category` metadata. The assistant queries with a `category` filter and passes the top passages to its own prompt. When a page changes, it is ingested again at a higher `generation`, and the chunks the new version no longer has stop appearing at once.

**Questions over a client's file drop.** A consultant creates a `cloud` collection per engagement, adds the contracts and spreadsheets the client sent, and asks questions through `chat`. Each answer cites the item and page it rests on, so every statement can be checked before it goes into a report; `analysis` produces a summary, entities and contradictions across the files.

**Meeting recordings to a knowledge base.** After each recorded meeting, the app estimates the price, uploads the video in 64 MiB parts with the call's own transcript as a reference, and starts a job with `max_price_usd`. When the job is `done`, the archive (audio, key frames and the knowledge base) is filed with the meeting.

## Limits and pricing

| Limit | Value |
|---|---|
| Vector Database request body | 1 MiB |
| Documents per ingest, ids per delete | 100 |
| One document (`id` + `text` + `metadata`) | 10 KiB |
| Query text; `top_k` | 8,000 characters; 50 |
| Filter `$in` values | 20 |
| Item in a request body, or fetched by URL | 16 MiB |
| Items per collection | 100; 1 GiB of files sent in requests; 20 GiB of video by upload |
| Collection query `top_k`; chat messages | 20; 1 to 20 messages of up to 8,000 characters |
| Chunks per `embed` call | 100, each up to 8,000 characters |
| Analysis model calls per report | 32 |
| Upload part size | 64 MiB (every part but the last) |
| Non-video upload (transcript, notes) | 16 MiB |
| Recording upload | your organization's cap |
| Fetch keys held unused | 16, each valid once for 60 seconds |
| Job result kept | 24 hours |
| Collection retention | 30 days by default |

**Pricing:** on request. A recording processed by a job is priced in **video minutes**, and `POST /v1/knowledge/estimates` returns its price line by line before you upload anything. Chat over a collection is metered as [AI chat](/en/products/ai.chat/) on your chat route.

## Errors

| Status | Code | What it means and what to do |
|---|---|---|
| 400 | `invalid_body`, `unsupported_field` | A field or value the route does not accept; check the route's fields. |
| 400 | `invalid_profile` | Pick a profile listed by `GET /v1/vector/profiles`. |
| 400 | `confirmation_required` | Deleting the database needs `{"confirm":"delete-database-and-all-vectors"}`. |
| 400 | `invalid_url`, `signed_url_must_be_sealed` | Only public `https` URLs on port 443; a signed URL must be sealed. |
| 400 | `fetch_key_invalid`, `invalid_sealed_url` | Request a new fetch key and seal again. |
| 403 | `vector_database_disabled`, `collections_disabled`, `knowledge_disabled` | That part is not enabled for your organization. |
| 403 | `budget_exceeded` | The job's `max_price_usd` is over your organization's cap. |
| 404 | `collection_not_found`, `job_not_found`, `upload_not_found`, `analysis_not_found` | Check the id; a job's download is also `404` until it is `done`. |
| 409 | `vector_database_not_created`, `vector_database_exists`, `vector_database_deleting` | Create the database first; one per organization; wait for a delete to finish. |
| 409 | `mode_not_supported` | The operation does not exist in this collection's mode. |
| 409 | `collection_incomplete`, `collection_full` | Wait for items to finish; the collection is at its limit. |
| 409 | `upload_incomplete`, `operation_id_reused`, `operation_in_progress` | Finish the upload; use a new operation id for a different file; wait. |
| 410 | `collection_expired`, `result_expired` | The retention ended; process again. |
| 413 | `request_too_large`, `url_too_large` | Use `POST /v1/uploads` for large files. |
| 415, 422 | `unsupported_media_type`, `input_rejected`, `archive_rejected` | The file type is not read, or the file was refused (`reason` says why). |
| 429 | `too_many_fetch_keys` and the quota codes | Use or let expire the keys you hold; see [Errors](/en/guides/errors/). |
| 502, 504 | `url_fetch_failed`, `url_fetch_timeout`, `upstream_error`, `timeout_error` | The URL or the processing failed; retry. |
| 503 | `service_unavailable` | Retry later; nothing was guessed. |

## Best practices

- **Group chunks under a `document_id` and bump `generation`** when the source changes; deletion and replacement then work per document.
- **Send `content_sha256` with each chunk** and compare it with `GET /v1/vector/documents` to resync your copy without re-ingesting everything.
- **Check `searchable` (or `indexing.complete`) before relying on the newest passages.**
- **Keep the answer of a passage within its first 2,000 characters** when you use reranking; only that much is read by the reranker.
- **Compare `vector_rerank` with plain `vector` on your own data**, especially outside English, before relying on it.
- **Always estimate before a job and set `max_price_usd`**; poll at `poll_after_seconds`, not faster.
- **Seal signed URLs**; never send a URL that carries a login. Upload the bytes instead.
- **Check citations**: a model can cite a real source for something it does not say.

## Related

- [AI embeddings](/en/products/ai.embeddings/), [AI chat](/en/products/ai.chat/)
- [Files](/en/products/files/): store the source files themselves
- [Authentication](/en/guides/authentication/), [Errors](/en/guides/errors/), [Limits](/en/guides/limits/)

## Endpoints

- POST /v1/vector/ingest
- POST /v1/vector/query
- POST /v1/collections
- POST /v1/collections/{id}/query
- POST /v1/knowledge/jobs

## Example

```sh
curl -X POST "https://ai.inovacc.dev/v1/vector/ingest" -H "Authorization: Bearer $INOVACC_API_KEY" -H "X-Operation-Id: $(uuidgen)" -H "Content-Type: application/json" -d '{}'
```
