# Inovacc developer documentation: full text Updated 2026-10-10. --- Product: AI chat URL: https://developer.inovacc.dev/en/products/ai.chat/ Base URL: https://ai.inovacc.dev Chat completions from the platform's language models. Endpoints: - POST /v1/chat/completions - GET /v1/models - POST /v1/run Example: ```sh curl -X POST "https://ai.inovacc.dev/v1/chat/completions" -H "Authorization: Bearer $INOVACC_API_KEY" -H "X-Operation-Id: $(uuidgen)" -H "Content-Type: application/json" -d '{}' ``` ## Overview AI chat gives your application conversational answers from the language models Inovacc runs for your organization. The endpoint, `POST /v1/chat/completions`, speaks the OpenAI chat-completions wire format, so an existing OpenAI client library can call it by changing the base URL to `https://ai.inovacc.dev/v1` and adding two headers. You send a list of messages and get back the assistant's reply, either as one JSON document or as a stream of events while it is being written. The problem it solves is access without plumbing: your organization does not hold model-provider accounts, keys or contracts. Inovacc decides which model serves each of your routes, applies your organization's limits and budgets before a call is made, and records the usage of every call so you can see what each key and route spends. Use AI chat when the answer is free text: an assistant, a summary, a draft, an extraction into JSON, a reply that calls your tools. Another product fits better when: - you need a structured verdict against criteria (a score, a choice, a yes or no, with confidence): use the [Decision engine](/en/products/decision/); - the answer must come from your own documents, with citations: ask a cloud collection of [Knowledge and vector search](/en/products/knowledge/); - you need vectors rather than text: use [AI embeddings](/en/products/ai.embeddings/). ## Concepts **Routes, not models.** You never name a model. Your organization is given **routes**: opaque ids such as `r00`, configured by Inovacc for your organization alone. A route decides which model answers, its output ceiling and the fields it accepts. `GET /v1/models` lists the routes your organization may use, each with the `endpoint` it serves; the `id` is always the route id, never the name of what backs it. Send the id as the `model` field; send `"auto"`, or leave `model` out, and your organization's default route answers. A route serves exactly one endpoint: a chat route called from `/v1/embeddings` or `/v1/run` is refused like a route you are not allowed to use. **Fields.** The common OpenAI fields are always accepted: `messages`, `model`, `stream`, `stream_options`, `max_tokens`, `max_completion_tokens`, `temperature`, `top_p`, `stop`, the penalties, `logit_bias`, `logprobs`, `tools`, `tool_choice`, `response_format`, `seed`, `reasoning_effort` and a few more. Seven fields (`n`, `service_tier`, `store`, `metadata`, `modalities`, `audio`, `web_search_options`) are accepted only when your organization has been granted them. Any other field is refused by name. Some routes take a narrower set (`messages`, `model`, `stream`, `stream_options`, the two output limits, `temperature`, `top_p`); on those, `tools` and the other fields are refused for that route. **Messages and images.** `content` is a string, or an array of text parts and `image_url` parts. An image is a `data:` URL (PNG, JPEG, WebP or GIF, base64) or an `https://` URL; a route whose model cannot fetch a URL refuses the `https://` form. One request carries at most 20 images, and each counts as an estimated 1,600 input tokens. **Output ceiling.** `max_tokens` and `max_completion_tokens` can only lower the ceiling the route and your organization set; a higher value is clamped, never refused. Reasoning tokens a model spends are inside `completion_tokens`. **Streaming.** With `"stream": true` the answer is a `text/event-stream` of OpenAI `chat.completion.chunk` events, each a `data:` line, ending with `data: [DONE]`. Set `stream_options.include_usage` to receive a final chunk with the token counts. **Usage headers.** A non-streaming answer carries `x-route-id`, `x-usage-input-tokens` and `x-usage-output-tokens`, and every answer carries `x-operation-id`. ## How it works Every call carries `Authorization: Bearer ` and an `X-Operation-Id`. Your organization comes from the key. The operation id is your idempotency key: 8 to 128 letters, digits, `_` or `-`. The service checks, in order: the operation id, the key, whether that operation id was already used, the body and its fields, the route, the messages, the estimated input size (a quarter of the characters of the whole request body, plus the image estimate) and then reserves the call against your organization's quotas. Only when all of that passes is a model called. The answer's `id` is always `chatcmpl-` followed by your operation id, and its `model` is always the route id. A reply whose whole output budget went to reasoning is still `200`, with empty `content` and `finish_reason: "length"`; a reply stopped by the model's safety filter is `200` with `finish_reason: "content_filter"`. A successful non-streaming answer is stored for 24 hours under its operation id. Send the same id again and you receive the stored answer, with `x-idempotent-replay: true`, without a second model call, a second charge or a quota change. An error is never stored, so retrying a failed call with the same id runs it again. While a call is still running, a second call with its id is `409 operation_in_progress`. A stream is never stored. If the model fails partway, the stream ends with one error event (`upstream_error`) and `data: [DONE]`; a stream that ends without `[DONE]` was cut and is incomplete. Read `delta.content` on every chunk, including the first and the one that carries `finish_reason`. What is metered is tokens: input and output, as reported by the model, or the estimate when no count is available. `GET /v1/usage/summary` sums your organization's usage per key, route, day or hour, and `GET /v1/capabilities` tells an agent what your organization may call; neither is charged. ## Get started You need an API key and at least one chat route enabled for your organization. Keys are created in the Inovacc console; see [Authentication](/en/guides/authentication/). 1. **List your routes.** Call `GET /v1/models`. Each entry with `"endpoint": "/v1/chat/completions"` is a route you can chat with; note its `id`. 2. **Send a first message.** Call `POST /v1/chat/completions` with that id as `model`, one `user` message and a fresh `X-Operation-Id` (see [the samples](#example)). The answer's `id` ends with your operation id and `x-usage-input-tokens` shows what was counted. 3. **Retry it.** Send the same request with the same operation id. The answer is identical and carries `x-idempotent-replay: true`: nothing ran twice. 4. **Stream.** Add `"stream": true` and `"stream_options": {"include_usage": true}` with a new operation id. Chunks arrive as the reply is written, the last before `[DONE]` carrying `usage`. 5. **Check the spend.** Call `GET /v1/usage/summary?group_by=route` to see the calls and tokens you just made. ## Use cases **A support assistant.** Your help widget streams the reply so the person sees it appear at once. Each user turn is one call with a new operation id; your server keeps the conversation and sends the recent messages each time. If the connection drops before `[DONE]`, the widget shows the reply as cut and offers to ask again. **Structured extraction in bulk.** A nightly job reads incoming documents and asks for a JSON object with `response_format`. Each document gets an operation id derived from its own id, so when the job is restarted after a crash, finished documents come back from the stored answers in seconds and are not charged again. **An agent that calls your tools.** On a route that accepts `tools`, the model answers with `tool_calls`; your code runs the tool and sends the result back as a `tool` message. Each step is its own operation id, so a retried step never runs a tool twice on the model's side. ## Limits and pricing | Limit | Value | |---|---| | Request body | 1 MiB by default; your organization may be set between 1 KiB and 20 MiB | | Images per request | 20 | | Estimated input cost of one image | 1,600 tokens | | Input tokens per request | your organization's cap (`400 input_too_large` above it) | | Output tokens per request | the lower of the route's and your organization's ceiling | | Time to the model's first response | 60 seconds | | Stored answer for a replay | 24 hours | | A call held as in progress | 5 minutes at most | | Requests per minute, per day; tokens and cost per day and month; concurrent calls; daily budgets | as set for your organization | **Pricing:** on request. The pricing unit is **tokens**. ## Errors | Status | Code | What it means and what to do | |---|---|---| | 400 | `missing_operation_id`, `invalid_operation_id` | Send `X-Operation-Id` with 8 to 128 letters, digits, `_` or `-`. | | 400 | `invalid_body` | The body is not a JSON object, or `user` is not a string. Fix the body. | | 400 | `invalid_messages` | `messages` is empty or a message is malformed. | | 400 | `unsupported_field` | A top-level field this endpoint does not accept; remove it. | | 400 | `unsupported_field_for_route` | The route takes a narrower set of fields, or cannot fetch an image URL; remove the field or send the image as `data:`. | | 400 | `model_not_allowed` | The `model` value is not a valid route id. Use an id from `GET /v1/models`. | | 400 | `input_too_large` | The input is over your organization's per-request cap, or there are more than 20 images. Shorten it. | | 401 | `missing_credentials`, `invalid_credentials` | Send a valid key. | | 403 | `key_disabled`, `organization_disabled` | The key or organization was disabled; contact your administrator. | | 403 | `field_not_allowed` | A field your organization has not been granted. | | 403 | `route_forbidden` | The route is not one your organization may use on this endpoint. | | 409 | `operation_in_progress` | That operation id is still running; wait and retry. | | 413 | `request_too_large` | The body is over your size cap. | | 429 | `rate_limited` | Too many requests; wait `Retry-After` seconds. | | 429 | `quota_exceeded`, `concurrency_limited` | A token, cost or concurrency allowance is used up; no `Retry-After`. | | 429 | `budget_exhausted` | A daily budget is spent; it resets at 00:00 UTC (`Retry-After`). | | 502 | `upstream_error` | The model did not answer usably; retry with the same operation id. | | 503 | `service_unavailable`, `route_unavailable` | Retry later; `route_unavailable` needs Inovacc to fix your route. | | 504 | `timeout_error` | The model did not start answering in 60 seconds; retry. | The envelope and every type are in [Errors](/en/guides/errors/). ## Best practices - **One operation id per logical request**, reused only to retry that same request. A reused id returns the first answer whatever the new body says. - **Retry on `502`, `503`, `504` and on network errors with the same operation id**, with backoff. On `429 rate_limited` and `budget_exhausted`, wait the `Retry-After` seconds. - **Set `max_tokens`** to what you need: it lowers the reservation made against your quotas and the cost of a runaway reply. - **Watch `X-Budget-Warning`**: it appears once your organization is past a budget's warning threshold, before calls are refused. - **Use `stream` for people, not for jobs**: a streamed reply cannot be replayed, a non-streamed one can. - **Send images as `data:` URLs** when you do not know the route's model can fetch a URL, and keep them under the 20-image limit. - **Never ship a secret key to a browser**: call AI chat from your server. ## Related - [Authentication](/en/guides/authentication/), [Errors](/en/guides/errors/), [Limits](/en/guides/limits/) - [Decision engine](/en/products/decision/): structured verdicts on the same key - [Knowledge and vector search](/en/products/knowledge/): answers from your documents, with citations - [AI embeddings](/en/products/ai.embeddings/) --- Product: AI embeddings URL: https://developer.inovacc.dev/en/products/ai.embeddings/ Base URL: https://ai.inovacc.dev Turn text into vectors for search and similarity. Endpoints: - POST /v1/embeddings Example: ```sh curl -X POST "https://ai.inovacc.dev/v1/embeddings" -H "Authorization: Bearer $INOVACC_API_KEY" -H "X-Operation-Id: $(uuidgen)" -H "Content-Type: application/json" -d '{}' ``` ## Overview AI embeddings turns text into vectors: lists of numbers whose distance from each other follows the meaning of the text. Two passages about the same thing land close together even when they share no words. The endpoint, `POST /v1/embeddings`, speaks the OpenAI embeddings wire format, so an existing OpenAI client can call it with the base URL `https://ai.inovacc.dev/v1` and two extra headers. The problem it solves is semantic comparison inside your own systems. With vectors you can search by meaning, group similar items, find near-duplicates or route a message to the closest category, using whatever database or index you already run. Use AI embeddings when **you** keep the vectors. When you would rather have Inovacc keep them, search them and return the matching passages, use the managed Vector Database of [Knowledge and vector search](/en/products/knowledge/), which embeds your documents for you. When you need generated text rather than vectors, use [AI chat](/en/products/ai.chat/). ## Concepts **Routes.** As with chat, you never name a model. Your organization is given embedding routes, opaque ids configured by Inovacc. `GET /v1/models` lists them: an entry whose `endpoint` is `/v1/embeddings` is an embedding route. Send its id as `model`, or send `"auto"` (or no `model`) to use your organization's default route. A route serves one endpoint only; a chat route called here is refused. **Input.** `input` is one non-empty string, or an array of strings; inside an array an empty string is accepted. The answer holds one vector per input, in `data[]`, each with the `index` of the input it belongs to. **Vectors of one route belong together.** A vector is only comparable with vectors produced by the same route. Store the route id next to every vector you keep, and embed both your documents and your queries with the same route. **Optional fields.** `encoding_format` and `dimensions` are passed to the model where the route's model supports them and ignored otherwise. `user` is accepted as a string and replaced, before it leaves Inovacc, by a value derived from your organization, so two organizations sending the same value never collide. **Usage.** The answer's `usage` holds `prompt_tokens` and `total_tokens`. This endpoint carries no `x-usage-*` headers; the route that answered is in `x-route-id`. ## How it works Every call carries `Authorization: Bearer ` and an `X-Operation-Id`. Your organization comes from the key. The service checks the operation id and the key, then whether that operation id was already used, then the body: only `model`, `input`, `encoding_format`, `dimensions` and `user` are accepted, and any other field is refused by name. It resolves the route, checks that the route serves embeddings, validates `input`, estimates the input size (a quarter of the characters of every input string) against your organization's per-request cap, and reserves the call against your organization's quotas. Then the model is called. The answer is `{"object":"list","data":[{"object":"embedding","index":0,"embedding":[...]}],"model":"","usage":{...}}`. `model` is always the route id. There is no streaming. A successful answer is stored for 24 hours under its operation id: the same id sent again returns the same vectors with `x-idempotent-replay: true`, without a second call or charge. An error is never stored, so a failed call can be retried with the same id. What is metered is input tokens. ## Get started You need an API key and an embedding route enabled for your organization ([Authentication](/en/guides/authentication/)). 1. **Find your embedding route.** Call `GET /v1/models` and pick an entry whose `endpoint` is `/v1/embeddings`. 2. **Embed two passages.** Call `POST /v1/embeddings` with that route as `model`, an `input` array of two strings and a fresh `X-Operation-Id` (see [the samples](#example)). The answer has two entries in `data`, with `index` 0 and 1, and `model` is your route id. 3. **Compare them.** Compute the cosine similarity of the two vectors in your code. Embed a third passage on another subject and compare again: its score against the first is lower. 4. **Replay.** Send step 2 again with the same operation id; the vectors are identical and the answer carries `x-idempotent-replay: true`. ## Use cases **Search in your own database.** A product catalogue stores one vector per product description in the database it already uses. A shopper's question is embedded with the same route at query time, and the nearest products are shown, including those whose description uses different words. **Near-duplicate detection.** A ticketing system embeds each new ticket and compares it with the open ones; a score above a threshold you calibrate links the ticket to the existing case instead of opening a second one. **Routing by similarity.** A small set of example messages per team is embedded once. Each incoming message is embedded and sent to the team whose examples are closest, with no model call per message beyond the embedding. ## Limits and pricing | Limit | Value | |---|---| | Request body | 1 MiB by default; your organization may be set between 1 KiB and 20 MiB | | Input tokens per request | your organization's cap (`400 input_too_large` above it) | | Time to the model's first response | 60 seconds | | Stored answer for a replay | 24 hours | | Requests, tokens, cost, concurrency, daily budgets | as set for your organization | **Pricing:** on request. The pricing unit is **tokens**. ## Errors | Status | Code | What it means and what to do | |---|---|---| | 400 | `missing_operation_id`, `invalid_operation_id` | Send a valid `X-Operation-Id`. | | 400 | `unsupported_field` | A field other than `model`, `input`, `encoding_format`, `dimensions`, `user`; remove it. | | 400 | `invalid_body` | Malformed JSON, `input` of the wrong shape, or a non-string `user`. | | 400 | `model_not_allowed` | `model` is not a valid route id. | | 400 | `input_too_large` | The input is over your per-request cap; split it into several calls. | | 401 | `missing_credentials`, `invalid_credentials` | Send a valid key. | | 403 | `route_forbidden` | The route is not yours, is disabled, or does not serve embeddings. | | 403 | `key_disabled`, `organization_disabled` | Check the key. | | 409 | `operation_in_progress` | That operation id is still running; wait and retry. | | 413 | `request_too_large` | The body is over your size cap. | | 429 | `rate_limited`, `budget_exhausted` | Wait the `Retry-After` seconds. | | 429 | `quota_exceeded`, `concurrency_limited` | An allowance is used up; no `Retry-After`. | | 502, 503, 504 | `upstream_error`, `service_unavailable`, `timeout_error` | Retry with the same operation id and backoff. | See [Errors](/en/guides/errors/) for the envelope. ## Best practices - **Batch inputs**: send many passages in one `input` array, within your size and token caps, rather than one call per passage. - **Keep the route id with the vector**, and re-embed everything if you change route: vectors of two routes are not comparable. - **Chunk long documents** into passages of a few paragraphs before embedding; one vector for a whole document blurs its meaning. - **Derive the operation id from the content** (for example a hash of the batch) so a restarted job replays instead of paying again. - **Retry `502`, `503` and `504` with the same operation id**; wait `Retry-After` on `429 rate_limited`. - **Calibrate thresholds on your own data**: similarity scores are relative, not probabilities. ## Related - [Knowledge and vector search](/en/products/knowledge/): the managed Vector Database that embeds and searches for you - [AI chat](/en/products/ai.chat/) - [Authentication](/en/guides/authentication/), [Errors](/en/guides/errors/), [Limits](/en/guides/limits/) --- Product: Knowledge and vector search URL: https://developer.inovacc.dev/en/products/knowledge/ Base URL: https://ai.inovacc.dev Ingest documents, build knowledge collections and query them by meaning. Endpoints: - POST /v1/vector/ingest - POST /v1/vector/query - POST /v1/collections - POST /v1/collections/{id}/query - POST /v1/knowledge/jobs Example: ```sh curl -X POST "https://ai.inovacc.dev/v1/vector/ingest" -H "Authorization: Bearer $INOVACC_API_KEY" -H "X-Operation-Id: $(uuidgen)" -H "Content-Type: application/json" -d '{}' ``` ## Overview Knowledge and vector search keeps your organization's knowledge where an application can ask it questions by meaning. It has three parts, all on `https://ai.inovacc.dev`: - the **Vector Database**: a managed store of text passages you ingest and search, with filters, scopes and optional reranking; - **knowledge collections**: send a batch of mixed files (documents, images, video) and get them processed into knowledge packages, then download the result, query it, chat with it with citations, or have a report written over it; - **knowledge jobs**: turn a long recording into a knowledge base you download as one archive. The problem it solves is everything between "we have files" and "the assistant can answer from them": extracting text, cutting it into passages, embedding, indexing, keeping the index consistent, and citing the source of every answer. Inovacc does those steps and you call a handful of endpoints. Use it when answers must come from your own material. When you already run your own vector index, [AI embeddings](/en/products/ai.embeddings/) gives you the vectors alone. When the answer is free text with no sources, use [AI chat](/en/products/ai.chat/). Each part is enabled for your organization by Inovacc; a part that is not enabled answers `403` (`vector_database_disabled`, `collections_disabled` or `knowledge_disabled`). ## Concepts **Vector Database and its profile.** Your organization creates one database with `POST /v1/vector/databases`, choosing an **embedding profile** from `GET /v1/vector/profiles`. A profile states the modality (`text`), the languages (`english` or `multilingual`), the dimensions and the distance metric; ids look like `text-multilingual-1024`. You choose a profile, never a model. The profile cannot change: to use another, delete the database, create a new one and ingest again. **Documents, scopes and modes.** You ingest passages with your own `id`, their `text` and optional flat `metadata`. A passage goes to a **scope**: `shared` (every service of your organization's workspace) or `local` (only the API key that wrote it). Queries take a **mode**: `shared`, `local`, or `hybrid` to search both and merge by score. Six metadata keys can be filtered on: `kind`, `source`, `category`, `tag`, `language`, `created_at`. Chunks that share a `document_id` form a document with a `generation` (your version number), which lets you delete or replace a whole document at once. **Collections and their mode.** A collection is created in one of three modes, fixed for its life: `download` (process files, download the packages, nothing stays afterwards), `cloud` (Inovacc keeps the processed content and its vectors so you can `query`, `chat` and run an `analysis`) or `local_vectors` (Inovacc keeps vectors only, never text: you `embed` your own chunks and `query` returns chunk ids and scores). Collections are kept 30 days by default; an organization set to `zero` retention has them deleted one hour after the download starts. **Items.** A file joins a collection in the request (up to 16 MiB), from a completed upload (video above 16 MiB), by public URL, or by a signed URL sealed to a single-use key so that it is never stored. Each item goes `queued`, `processing`, then `done` or `failed`; a failed item does not stop the others. **Eventual consistency and `searchable`.** The vector index applies writes asynchronously: a vector is readable some seconds, sometimes minutes, after it was written. `GET /v1/vector/databases` and every collection answer report `searchable`; a query sent while it is `false` may miss the latest passages. A `cloud` collection's `indexing.complete` is `true` only when every chunk is written **and** readable. **Jobs, power and budget.** A knowledge job takes a completed upload of a recording, an optional transcript and notes, a `power` (`light`, `medium` or `max`) and a required `max_price_usd`. A job over your organization's cap is refused before any paid work. ## How it works Every call carries `Authorization: Bearer `. Your organization comes from the key. `X-Operation-Id` is required on the collection, upload and job `POST` calls and makes them idempotent: a retry with the same id returns the first answer (`X-Idempotent-Replay: true`) and adds nothing; the same id with a different file is `409 operation_id_reused`. On the Vector Database it is optional, because ingest and delete are idempotent by construction: ingesting an existing `id` replaces it. A **query** (`POST /v1/vector/query`) takes `mode`, `query` (1 to 8,000 characters), `top_k` (1 to 50) and an optional `filter`; it returns matches best first, each with `id`, `score`, `text` and `metadata`. With `"retrieval": "vector_rerank"` and a `rerank_profile`, a wider candidate list is re-ordered by a reranking stage; the answer's `retrieval` object says whether the rerank applied or fell back to vector order. A **collection** works in steps you can observe: create it, add items (`202`, `queued`), read it until `status` is `complete`, then download, query or chat. Reading a `cloud` collection also advances its indexing, a bounded amount per call, and every answer carries the `indexing` progress. `chat` answers only from the retrieved sources, numbers each citation with the item, package, chunk and locator, and says it cannot answer when the sources do not cover the question. An `analysis` is built over several requests: repeat `POST` or `GET` every few seconds until it answers `200` with the report. A **job** is asynchronous: `POST /v1/knowledge/jobs` answers `202` with `job_id`, `status_url` and `poll_after_seconds` (also as `Retry-After`). Read the status when told; the platform forecasts the time to done and includes `eta_seconds` while it can. When `state` is `done`, download the archive; it is kept 24 hours. What is metered: embedding tokens and vectors written for ingest and indexing, queries, reranks as their own unit, chat calls on your chat route, and jobs by their price. `POST /v1/knowledge/estimates` prices a recording before you upload it, at no cost. ## Get started You need an API key and the part you want enabled for your organization ([Authentication](/en/guides/authentication/)). 1. **Create the database.** `GET /v1/vector/profiles`, pick a profile, then `POST /v1/vector/databases` with `{"profile":""}`. The answer is `201` with the database. 2. **Ingest passages.** `POST /v1/vector/ingest` with `"scope":"shared"` and two or three documents (see [the samples](#example)). The answer says how many were `ingested` and lists any `skipped` with a reason. 3. **Wait for `searchable`.** `GET /v1/vector/databases` until `searchable` is `true`. 4. **Query.** `POST /v1/vector/query` with `"mode":"shared"` and a question worded differently from your passages. The closest passage comes first, with its `score`. 5. **Try a collection.** `POST /v1/collections` with `{"mode":"cloud"}`, add a PDF to `/items`, read the collection until `indexing.complete` is `true`, then `POST /v1/collections/{id}/chat` and check the `citations`. ## Use cases **An assistant grounded in your handbook.** Your intranet ingests each policy page as chunks under its `document_id`, with `category` metadata. The assistant queries with a `category` filter and passes the top passages to its own prompt. When a page changes, it is ingested again at a higher `generation`, and the chunks the new version no longer has stop appearing at once. **Questions over a client's file drop.** A consultant creates a `cloud` collection per engagement, adds the contracts and spreadsheets the client sent, and asks questions through `chat`. Each answer cites the item and page it rests on, so every statement can be checked before it goes into a report; `analysis` produces a summary, entities and contradictions across the files. **Meeting recordings to a knowledge base.** After each recorded meeting, the app estimates the price, uploads the video in 64 MiB parts with the call's own transcript as a reference, and starts a job with `max_price_usd`. When the job is `done`, the archive (audio, key frames and the knowledge base) is filed with the meeting. ## Limits and pricing | Limit | Value | |---|---| | Vector Database request body | 1 MiB | | Documents per ingest, ids per delete | 100 | | One document (`id` + `text` + `metadata`) | 10 KiB | | Query text; `top_k` | 8,000 characters; 50 | | Filter `$in` values | 20 | | Item in a request body, or fetched by URL | 16 MiB | | Items per collection | 100; 1 GiB of files sent in requests; 20 GiB of video by upload | | Collection query `top_k`; chat messages | 20; 1 to 20 messages of up to 8,000 characters | | Chunks per `embed` call | 100, each up to 8,000 characters | | Analysis model calls per report | 32 | | Upload part size | 64 MiB (every part but the last) | | Non-video upload (transcript, notes) | 16 MiB | | Recording upload | your organization's cap | | Fetch keys held unused | 16, each valid once for 60 seconds | | Job result kept | 24 hours | | Collection retention | 30 days by default | **Pricing:** on request. A recording processed by a job is priced in **video minutes**, and `POST /v1/knowledge/estimates` returns its price line by line before you upload anything. Chat over a collection is metered as [AI chat](/en/products/ai.chat/) on your chat route. ## Errors | Status | Code | What it means and what to do | |---|---|---| | 400 | `invalid_body`, `unsupported_field` | A field or value the route does not accept; check the route's fields. | | 400 | `invalid_profile` | Pick a profile listed by `GET /v1/vector/profiles`. | | 400 | `confirmation_required` | Deleting the database needs `{"confirm":"delete-database-and-all-vectors"}`. | | 400 | `invalid_url`, `signed_url_must_be_sealed` | Only public `https` URLs on port 443; a signed URL must be sealed. | | 400 | `fetch_key_invalid`, `invalid_sealed_url` | Request a new fetch key and seal again. | | 403 | `vector_database_disabled`, `collections_disabled`, `knowledge_disabled` | That part is not enabled for your organization. | | 403 | `budget_exceeded` | The job's `max_price_usd` is over your organization's cap. | | 404 | `collection_not_found`, `job_not_found`, `upload_not_found`, `analysis_not_found` | Check the id; a job's download is also `404` until it is `done`. | | 409 | `vector_database_not_created`, `vector_database_exists`, `vector_database_deleting` | Create the database first; one per organization; wait for a delete to finish. | | 409 | `mode_not_supported` | The operation does not exist in this collection's mode. | | 409 | `collection_incomplete`, `collection_full` | Wait for items to finish; the collection is at its limit. | | 409 | `upload_incomplete`, `operation_id_reused`, `operation_in_progress` | Finish the upload; use a new operation id for a different file; wait. | | 410 | `collection_expired`, `result_expired` | The retention ended; process again. | | 413 | `request_too_large`, `url_too_large` | Use `POST /v1/uploads` for large files. | | 415, 422 | `unsupported_media_type`, `input_rejected`, `archive_rejected` | The file type is not read, or the file was refused (`reason` says why). | | 429 | `too_many_fetch_keys` and the quota codes | Use or let expire the keys you hold; see [Errors](/en/guides/errors/). | | 502, 504 | `url_fetch_failed`, `url_fetch_timeout`, `upstream_error`, `timeout_error` | The URL or the processing failed; retry. | | 503 | `service_unavailable` | Retry later; nothing was guessed. | ## Best practices - **Group chunks under a `document_id` and bump `generation`** when the source changes; deletion and replacement then work per document. - **Send `content_sha256` with each chunk** and compare it with `GET /v1/vector/documents` to resync your copy without re-ingesting everything. - **Check `searchable` (or `indexing.complete`) before relying on the newest passages.** - **Keep the answer of a passage within its first 2,000 characters** when you use reranking; only that much is read by the reranker. - **Compare `vector_rerank` with plain `vector` on your own data**, especially outside English, before relying on it. - **Always estimate before a job and set `max_price_usd`**; poll at `poll_after_seconds`, not faster. - **Seal signed URLs**; never send a URL that carries a login. Upload the bytes instead. - **Check citations**: a model can cite a real source for something it does not say. ## Related - [AI embeddings](/en/products/ai.embeddings/), [AI chat](/en/products/ai.chat/) - [Files](/en/products/files/): store the source files themselves - [Authentication](/en/guides/authentication/), [Errors](/en/guides/errors/), [Limits](/en/guides/limits/) --- Product: Decision engine URL: https://developer.inovacc.dev/en/products/decision/ Base URL: https://ai.inovacc.dev Structured, criteria-based decisions at a speed and precision tier you choose. Endpoints: - POST /v1/run Example: ```sh curl -X POST "https://ai.inovacc.dev/v1/run" -H "Authorization: Bearer $INOVACC_API_KEY" -H "X-Operation-Id: $(uuidgen)" -H "Content-Type: application/json" -d '{}' ``` ## Overview The Decision engine answers structured questions about a piece of content with structured verdicts. You send a **state** (the thing to judge: a text, a record, a JSON object) and a set of **questions**, each of a known type; you get back one answer per question, with the label or score chosen, the probabilities behind it and a confidence. The endpoint is `POST /v1/run` on `https://ai.inovacc.dev`, called on a decision route of your organization. The problem it solves is turning a judgement into data a program can act on. Asking a chat model "is this good?" returns prose you then have to parse and cannot compare from one call to the next. The Decision engine returns the same shape every time, so you can store it, threshold it and audit it. Use it for classification, scoring, triage and checks against criteria: does this ticket need a person, which of five categories fits, how well does this answer meet the bar. When you need generated prose, use [AI chat](/en/products/ai.chat/); when the verdict must rest on your own documents, retrieve the passages first with [Knowledge and vector search](/en/products/knowledge/) and put them in the state. ## Concepts **Decision routes.** As everywhere in Inovacc AI, you call a route, never a model. `GET /v1/models` lists your routes; an entry whose `endpoint` is `/v1/run` is a run route, and the ones Inovacc configured as decision routes answer in the shape below. **Questions and their types.** `questions` is an object of 1 to 64 entries, keyed by your own ids (letters, digits, `_`, `.` or `-`, up to 100 characters). Each question has a `type`, usually `instructions`, and `criteria` whose shape the type fixes: | Type | Asks for | `criteria` | |---|---|---| | `noul` | yes or no | an object `{"true": "...", "false": "..."}` | | `choice` | one label among several | an object `{"label": "description", ...}` | | `score` | a level on a scale | an array of levels, lowest first | The shapes are strict: a criteria string such as `"1-5"` or `"yes|no"` is refused. **Answers.** Each answer carries what its type produced: the `choice` or the `score` with `probabilities`, a `legend` and a `confidence`. A `noul` answer carries its probability `p` instead of a confidence. **Tiers.** A decision route serves one of three tiers: `fast`, `standard` or `precise`. All answer in one shape, and the answer's `engine.tier` says which one answered. **Confidence is not comparable between tiers**: each tier reports on its own scale, so set a threshold per tier. **Escalation.** With `escalate`, questions answered below a confidence threshold are asked again, alone, on the next stronger tier (`fast`, then `standard`, then `precise`), and the stronger answer replaces the weaker one. The defaults are `fast` 0.40, `standard` 0.80 and `precise` 0.55; you can pass `min_confidence` (one number, or one per tier) and `max_tier`. A `noul` answer's confidence is the distance of `p` from a coin flip, `|2p - 1|`. ## How it works Every call carries `Authorization: Bearer ` and an `X-Operation-Id`. Your organization comes from the key. The body takes only `model`, `input` and `escalate`. On a decision route `input` is `{"state", "questions", "images"?}`: `state` is any non-null JSON value, `images` at most 16 entries. Any other field inside `input` is refused with `invalid_decision_input`, and the reply never echoes it. The answer is: `{"model": "", "result": {"answers": {...}, "usage": {"input_tokens", "output_tokens"}}, "state": "Completed", "engine": {"tier", "latency_ms"}}` and its headers carry `x-route-id`, `x-usage-input-tokens` and `x-usage-output-tokens`. With escalation, each answer also says which `tier` answered it and, when it was escalated, `escalated_from` lists the lower tiers and the confidence each gave. An `escalation` object reports every attempt, how many questions were asked and how many came back below threshold, and how many answers each tier gave in the end. Each attempt is its own call, reserved against your quotas and metered at its own tier. If an attempt fails, the escalation stops, the earlier answers are kept, and the request is still `200`. There is no streaming and no 60-second timeout on this endpoint. A successful answer is stored for 24 hours under its operation id: the same id returns the same merged answer with `x-idempotent-replay: true` and calls nothing. What is metered is tokens, at the price of each tier that answered. ## Get started You need an API key and a decision route enabled for your organization ([Authentication](/en/guides/authentication/)). 1. **Find your decision route** with `GET /v1/models`: an entry whose `endpoint` is `/v1/run`. 2. **Ask one `noul` question.** `POST /v1/run` with your route as `model`, a short text as `state` and one question `{"type":"noul","instructions":"...","criteria":{"true":"...","false":"..."}}` (see [the samples](#example)). The answer has your question id under `result.answers`, with its probability `p`. 3. **Add a `score` question** with three levels to the same call. Both answers come back in one response; `engine.tier` names the tier. 4. **Escalate.** Send the call again with a new operation id and `"escalate": true`. The answer now carries `tier` per answer and an `escalation` object listing the attempts. ## Use cases **Ticket triage.** Each incoming support ticket is the `state`; a `choice` question picks the team, a `noul` question asks whether a person must answer, a `score` question rates urgency. The answers are stored with the ticket, and a ticket whose confidence is below your threshold goes to a person. **Quality checks on generated text.** Before an assistant's draft is sent, a decision call scores it against your bar ("Below the bar", "At the bar", "Above the bar") and checks a few `noul` rules. With `escalate`, only the drafts the fast tier was unsure about pay for the precise tier. **Content moderation with images.** A listing's text and up to 16 images are judged against your policy questions; the per-question confidence decides between publishing, rejecting and sending to review. ## Limits and pricing | Limit | Value | |---|---| | Questions per call | 1 to 64 | | Question id | 1 to 100 characters: letters, digits, `_`, `.`, `-` | | Images per call | 16 | | Request body | 1 MiB by default; your organization may be set between 1 KiB and 20 MiB | | Input tokens per request | your organization's cap | | Stored answer for a replay | 24 hours | | Requests, tokens, cost, concurrency, daily budgets | as set for your organization | **Pricing:** on request. The pricing unit is **tokens**, at the price of the tier that answered. ## Errors | Status | Code | What it means and what to do | |---|---|---| | 400 | `invalid_input` | `input` is missing or not a JSON object. | | 400 | `invalid_decision_input` | `input` is not `{state, questions, images?}` or a question is malformed; check types and ids. | | 400 | `invalid_escalate` | `escalate` has a key or value it does not take, or the route is not a decision route. | | 400 | `unsupported_field` | A top-level field other than `model`, `input`, `escalate`. | | 400 | `model_not_allowed`, `input_too_large` | Use a valid route id; shorten the state. | | 400 | `missing_operation_id`, `invalid_operation_id` | Send a valid `X-Operation-Id`. | | 401 | `missing_credentials`, `invalid_credentials` | Send a valid key. | | 403 | `route_forbidden` | The route is not yours, or is not a run route. | | 409 | `operation_in_progress` | That operation id is still running. | | 413 | `request_too_large` | The body is over your size cap. | | 429 | `rate_limited`, `quota_exceeded`, `concurrency_limited`, `budget_exhausted` | See [Errors](/en/guides/errors/); wait `Retry-After` where given. | | 502 | `upstream_error` | The tier failed, or a criteria shape was refused; check the criteria, then retry. | | 503 | `service_unavailable` | Decisions cannot be reached; retry later. | ## Best practices - **Write criteria in the strict shapes**: an object for `noul` and `choice`, an array for `score`. - **Keep thresholds per tier**, never one number for all tiers. - **Use `escalate` with `max_tier`** to cap the cost: most questions stop at the cheaper tier. - **Put everything the judgement needs in `state`**, including retrieved passages; the engine sees nothing else. - **Use stable question ids**, so answers can be compared across calls and stored by id. - **Derive the operation id from the item judged** so a retried batch replays instead of paying twice. ## Related - [AI chat](/en/products/ai.chat/), [Knowledge and vector search](/en/products/knowledge/) - [Authentication](/en/guides/authentication/), [Errors](/en/guides/errors/), [Limits](/en/guides/limits/) --- Product: Data URL: https://developer.inovacc.dev/en/products/data/ Base URL: https://data.inovacc.dev Databases, collections and records with schemas and rules, over one REST/JSON door. Endpoints: - GET /v1/databases - PUT /v1/databases/{database} - PUT /v1/databases/{database}/collections/{collection} - POST /v1/databases/{database}/collections/{collection}/records - POST /v1/databases/{database}/batch Example: ```sh curl -X GET "https://data.inovacc.dev/v1/databases" -H "Authorization: Bearer $INOVACC_API_KEY" ``` ## Overview Data is your application's database, served over one REST/JSON door at `https://data.inovacc.dev`. You create **databases**, give them **collections**, and read and write **records** with plain HTTP calls. A collection is either typed, with fields declared by a schema, or free, holding schemaless JSON documents. Who may do what is decided on every request, first by the credential's permissions and then, record by record, by the **rules** the collection declares. The problem it solves is running persistent data without running a database: no server, no connection pool, no migrations tool, no backups to schedule. Your application keeps its own code and stores its data here; where the data lives is decided by Inovacc and never shown. The same API serves your servers, your web pages (with a browser key) and the people signed in to your app, and a realtime listener pushes changes as they happen. Use Data for application state: users' notes, orders, settings, catalogues, anything shaped as records. Store large binary content with [Files](/en/products/files/), read what happened with the [Activity log](/en/products/activity/), and notify other systems with [Events](/en/products/events/). ## Concepts **Databases.** A database is a key your application chooses, such as `shop` or `crm/v2` (a `/` is written `%2F` in a path). A database belongs to the application whose credential created it; another application's database answers `404`, as if it did not exist. An organization-level credential sees every database of the organization. **Collections: schema or free.** A **schema** collection declares up to 64 fields, each with a type: `text`, `integer`, `number`, `bool`, `datetime`, `json`, `relation` (a record id in a `target` collection) or `blob` (the SHA-256 of a file in the same database). A field can be `required`, `unique`, `indexed`, or `fulltext` (on `text`). A **free** collection declares nothing: each record is a JSON document, queried by dotted path (`profile.city`), and it can be given a schema later once every document fits. Writing to a collection that does not exist creates a free one, when your credential may change schemas. **Records and versions.** A record is `{id, version, created_at, updated_at, ...fields}`. Its `id` is made by the service (`rec_` plus 26 sortable characters) or chosen by you. Every change raises `version` by one; updates and deletes must send `If-Match: `, so two writers never overwrite each other silently. **Rules.** A collection declares five rules, `list`, `view`, `create`, `update` and `delete`. Each is `null` (nobody), `""` (any server credential of your organization), `public` (anyone, browser keys included), `users` (the signed-in people of the owning application, and servers), or an expression such as `owner = @auth.id`. Rules apply to every credential, an administrator's included. **Credentials.** A **secret key** is for servers. A **publishable key** (`apb_...`) is for web pages: it works only from the origins its application lists and only where a rule admits it. A signed-in person's **access token** carries only record and file permissions. ## How it works Every call carries `Authorization: Bearer `. There is no organization header: your organization, and the application, come from the credential, and nothing in a path, query or header can name another. A request without a valid credential, or to a path that is not a route, gets one identical `401`. For each call the service first checks the credential's permission for the action (read, write, create, update, delete, or schema changes) on the database. Anything but "allow" is `403 forbidden` with no detail. Then the collection's rule decides per record: a list returns only the records the `list` rule admits; reading a record the `view` rule hides is `404`, the same as a missing record; a create checks the new data; an update checks both the stored record and the new data. Writes are versioned. `POST /records` creates and answers `201` with `ETag: "1"`. `PATCH` changes the named fields (on a free collection it is a JSON merge patch), `PUT` replaces the record, and `PUT` with `If-None-Match: *` creates it at your chosen id. A wrong `If-Match` is `409 version_conflict`; a missing one is `428`. Lists take `filter` (`field op value` joined by `&&` and `||`), `sort`, `limit` (up to 200) and an opaque `cursor`, and return `next_cursor` until the last page; `q` searches the `fulltext` fields. A **batch** applies up to 100 creates, updates and deletes in one transaction: all succeed, or none does. A **realtime listener** is a WebSocket on `GET .../collections/{c}/listen`: it sends `ready`, then `added`, `modified` and `removed` events for the records your rule lets you list. Every call is recorded in the [Activity log](/en/products/activity/). ## Get started You need a secret key with permission to create databases and collections ([Authentication](/en/guides/authentication/)). 1. **Create a database.** `PUT /v1/databases/notes` with no body. The answer is `201` with `{key, created_at}`; repeating it answers `200`. 2. **Declare a collection.** `PUT /v1/databases/notes/collections/items` with fields `owner` and `body` and rules (see [the samples](#example)). The answer is the collection with `schema_version` 1. 3. **Write a record.** `POST .../collections/items/records` with the fields. The answer is `201` with the record, its `id` and `version` 1. 4. **Update it.** `PATCH .../records/{id}` with `If-Match: 1`. The answer has `version` 2; sending `If-Match: 1` again is `409 version_conflict`. 5. **List.** `GET .../records?filter=owner%20%3D%20%22me%22&limit=10` returns your record and `next_cursor: null`. ## Use cases **Per-user data in a web app.** A notes app declares a collection whose five rules say `owner = @auth.id`. The page calls Data directly with the signed-in person's token; each person lists and edits only their own notes, and another person's note is a `404`. No server code enforces it: the rules do. **A public catalogue.** An online shop keeps `products` as a free collection with `list` and `view` set to `public` and the writes closed. The storefront reads it from the browser with a publishable key; the back office writes it from a server with a secret key. **Live dashboards.** An operations screen opens a listener on `orders` filtered to today's open orders. It runs the query after `ready`, then applies `added`, `modified` and `removed` events as they arrive, without polling. ## Limits and pricing | Limit | Value | |---|---| | JSON request body | 1 MiB | | One record | 900 KiB | | A `json` field | 64 KiB | | Fields per collection; collections per database | 64; 64 | | Records per page | 200 (default 50) | | Operations per batch | 100, in one transaction | | Filter | 2,048 characters, 40 values, nesting depth 8 | | Query string | 16 KiB | | Full-text search `q` | 16 terms, 256 characters | | Listeners per collection; one listener event | 1,000; 512 KiB | | Rate | 600 requests per minute per credential holder, per location (`429 rate_limited`) | **Pricing:** on request. ## Errors | Status | Code | What it means and what to do | |---|---|---| | 400 | `invalid_body`, `invalid_name`, `unknown_field`, `missing_field`, `invalid_value` | The body or a name breaks the rules; fix it. | | 400 | `record_too_large`, `field_too_large`, `reserved_field` | Shrink the record; do not send `id`, `version`, `created_at`, `updated_at` in a body. | | 400 | `invalid_schema`, `invalid_rule`, `schema_change_unsupported` | The collection definition is not valid, or a field was changed in place. | | 400 | `invalid_filter`, `invalid_sort`, `invalid_cursor`, `invalid_limit`, `invalid_search` | Fix the list query; start again without the cursor. | | 400 | `invalid_if_match`, `batch_too_large`, `cross_shard_batch` | Check the header or split the batch. | | 401 | `invalid_credentials` | Send a valid credential. | | 403 | `forbidden` | The credential's permissions or the collection's rule refuse it. | | 403 | `origin_not_allowed`, `secret_key_in_browser` | A browser key from an unlisted origin, or a secret key in a browser. | | 404 | `database_not_found`, `collection_not_found`, `record_not_found` | It does not exist, or you may not see it. | | 409 | `version_conflict` | Someone changed the record; read it again and retry. | | 409 | `already_exists`, `unique_violation`, `not_empty`, `schema_violation` | The id or value is taken; empty it before deleting; documents do not fit the schema. | | 412, 428 | `precondition_failed`, `precondition_required` | The record already exists; send `If-Match`. | | 413 | `payload_too_large` | The body is over 1 MiB. | | 429 | `rate_limited`, `too_many_listeners` | Slow down; close listeners you do not need. | | 503 | `service_unavailable` | Retry later; nothing was written. | ## Best practices - **Always send `If-Match`** with the version you read, and on `409 version_conflict` re-read before retrying. - **Close every collection by default** and open it rule by rule; a collection born from a write admits any server credential until you tighten it. - **Never put a secret key in a web page**: use a publishable key and `public` or `@auth` rules. - **Choose your own record ids** for imports, so a retried import creates nothing twice (`409 already_exists`). - **Use a batch** when several writes must succeed together. - **Index the fields you filter and sort on**, and follow `next_cursor` instead of large pages. - **With a listener, wait for `ready` before querying**, then drop events not newer than what the query returned. ## Related - [Files](/en/products/files/): binary content beside your records - [Activity log](/en/products/activity/): what was written, by whom - [Events](/en/products/events/): tell other systems about a change - [Authentication](/en/guides/authentication/), [Errors](/en/guides/errors/), [Limits](/en/guides/limits/) --- Product: Files URL: https://developer.inovacc.dev/en/products/files/ Base URL: https://data.inovacc.dev Content-addressed file storage, with multipart upload for large files. Endpoints: - PUT /v1/databases/{database}/blobs/{sha256} - GET /v1/databases/{database}/blobs/{sha256} - DELETE /v1/databases/{database}/blobs/{sha256} - POST /v1/databases/{database}/blobs/{sha256}/uploads Example: ```sh curl -X PUT "https://data.inovacc.dev/v1/databases/{database}/blobs/{sha256}" -H "Authorization: Bearer $INOVACC_API_KEY" -H "Content-Type: application/json" -d '{}' ``` ## Overview Files stores your application's binary content (images, documents, exports, recordings) on `https://data.inovacc.dev`, with uploads in parts for anything large. It offers two ways to address a file: - **by path, in your application's folder**: `PUT /v1/files/users/123/avatar.png`, the way a cloud storage bucket names files, with listings by folder, user metadata and folders shared with other applications of your workspace; - **by content, inside a database**: `PUT /v1/databases/{db}/blobs/{sha256}`, where the file's own SHA-256 is its name, so a record of [Data](/en/products/data/) can point at it with a `blob` field. The problem it solves is keeping files next to your data without running storage: Inovacc chooses where bytes live, verifies them against their hash, removes duplicates inside your folder, and serves them back with headers that stop a downloaded file from running in a browser. Use Files for anything that is bytes rather than fields. Keep the structured description of a file (owner, title, status) as a record in [Data](/en/products/data/). To make a file searchable by meaning, add it to a collection of [Knowledge and vector search](/en/products/knowledge/). ## Concepts **Your application's folder.** Every application has its own folder. The folder is chosen from the credential, never from the request: a credential bound to an application uses that application's folder, and a credential bound to no application gets `403 application_required` on every path route. Paths are UTF-8, 1 to 1,024 bytes, separated by `/`, case-sensitive, with no empty, `.` or `..` piece; a path that is not already in that form is refused, never silently rewritten. **Metadata.** A write may carry a `Content-Type` and your own `X-File-Meta-` headers (2 KiB in total). Both come back on every read. A write replaces the file's content, type and metadata whole. **Hashes.** Every file is identified by its SHA-256. Send `X-File-Sha256` with a write and the bytes are checked against it (`400 hash_mismatch`, nothing stored). Inside one application's folder, two paths with the same bytes share one stored copy; the copy is kept until the last path goes. Deduplication never crosses applications or organizations. **Database blobs.** A blob is addressed by the SHA-256 in its path inside one database. Uploading it again answers `200` instead of `201`. Anyone with read permission on the database who knows the hash can read it: record rules do not apply to blobs. **Shared folders.** An application can create a folder `team` that appears to every application with access as `shared/team/...`. The creator owns it and grants other applications of the same workspace `read` or `read_write`. An application without access cannot tell the folder exists: every operation answers `404`. **Large files.** Up to 100 MiB goes in one request. Above that, up to 5 GiB, you upload in parts and then link the result to a path (or it becomes the blob). ## How it works Every call carries `Authorization: Bearer `; there is no organization header, because your organization and application come from the credential. The permission checked is blob read, write or delete. A web page can use these routes with a publishable key from an origin its application lists. **Writing by path.** `PUT /v1/files/{path}` with the bytes and a `Content-Length` (required). The answer is `201` for a new path or `200` for a replacement: `{"path","size","sha256","content_type","updated_at","metadata"}` and `ETag: ""`. **Reading.** `GET /v1/files/{path}` returns the whole file (ranges are not supported) with `Content-Type`, `ETag`, `X-File-Sha256`, `X-File-Updated-At` and your `X-File-Meta-*` headers. Every download also carries `X-Content-Type-Options: nosniff`, a `Content-Security-Policy` that sandboxes it and `Cache-Control: no-store`, so an uploaded HTML page can never run as your site. `HEAD` returns the headers alone. **Listing.** `GET /v1/files?prefix=photos/&delimiter=/` lists the files directly under a folder as `items` and each sub-folder once in `prefixes`. Follow `next_cursor` until it is `null`: a page may be shorter than `limit` and still have a next one. **In parts.** `POST /v1/files-uploads/{sha256}` starts an upload (or answers `200` with `exists: true` when your folder already holds those bytes). Send parts 1 to 10,000 with `PUT .../{upload_id}/parts/{n}`, each between 5 MiB and 100 MiB except the last, then `POST .../complete` with the list of parts and their `etag`s. Inovacc reads the whole object back and checks its SHA-256 and size. Finally `PUT /v1/files/{path}` with `X-File-Sha256` and an empty body gives it a path. Database blobs use the same four steps under `/v1/databases/{db}/blobs/{sha256}/uploads`. Uploads, downloads and deletes are recorded in the [Activity log](/en/products/activity/) with the file's hash and size. ## Get started You need a secret key bound to an application, with blob permissions ([Authentication](/en/guides/authentication/)). 1. **Upload a file.** `PUT /v1/files/reports/2026-10.csv` with the file as the body and `Content-Type: text/csv` (see [the samples](#example)). The answer is `201` with its `sha256`. 2. **Read it back.** `GET /v1/files/reports/2026-10.csv`. The bytes arrive with `ETag` equal to the hash from step 1. 3. **List the folder.** `GET /v1/files?prefix=reports/&delimiter=/`. Your file is in `items`. 4. **Replace it.** `PUT` the same path with new bytes. The answer is `200`, with a new `sha256`. 5. **Delete it.** `DELETE /v1/files/reports/2026-10.csv` answers `204`; a second `GET` is `404 file_not_found`. ## Use cases **Profile pictures from the browser.** A web app uploads each avatar to `users//avatar.png` with a publishable key and keeps the returned `sha256` in the user's record, so a page can tell when the picture changed. **Exports shared with a partner app.** A reporting application creates the shared folder `exports`, grants the billing application `read`, and writes monthly files under `shared/exports/`. The billing application lists and downloads them; revoking the grant takes effect on its very next request. **Large media with integrity.** A video tool hashes a 2 GiB recording on the client, uploads it in 100 MiB parts and completes it. Inovacc stores it only if the SHA-256 matches, so a corrupted transfer never becomes a file. ## Limits and pricing | Limit | Value | |---|---| | File or blob in one request | 100 MiB, `Content-Length` required | | File or blob uploaded in parts | 5 GiB; parts 5 MiB to 100 MiB (except the last), numbered 1 to 10,000 | | Path | 1,024 bytes | | Listing prefix | 256 bytes | | Metadata per file | 2 KiB | | Listing page | 1 to 200 (default 50) | | Shared folder name | 1 to 64 letters, digits, `_` or `-` | | Rate | 600 requests per minute per credential holder, per location | **Pricing:** on request. ## Errors | Status | Code | What it means and what to do | |---|---|---| | 400 | `invalid_path`, `invalid_prefix` | The path or prefix is not in canonical form; fix it on your side. | | 400 | `invalid_metadata` | `Content-Type` or `X-File-Meta-*` breaks the rules or is over 2 KiB. | | 400 | `hash_mismatch` | The bytes do not match the SHA-256 you sent; nothing was stored. | | 400 | `invalid_part`, `invalid_limit`, `invalid_cursor`, `invalid_grantee` | Fix the upload parts, the listing or the grant. | | 401 | `invalid_credentials` | Send a valid credential. | | 403 | `application_required` | Path routes need a credential bound to an application. | | 403 | `forbidden` | Missing permission, or a `read` grantee tried to write. | | 404 | `file_not_found`, `blob_not_found`, `share_not_found`, `upload_not_found` | It does not exist, or you have no access to it. | | 409 | `already_exists`, `not_empty` | The shared folder name is taken; empty the folder before deleting it. | | 411 | `length_required` | Send `Content-Length`. | | 413 | `payload_too_large` | Over 100 MiB in one request (use parts) or over 5 GiB. | | 429 | `rate_limited` | Slow down. | | 503 | `service_unavailable` | Retry later. | ## Best practices - **Send `X-File-Sha256` on every write** so a damaged upload is refused instead of stored. - **Use parts above 100 MiB**, and start with `POST .../uploads`: it answers `exists: true` when the bytes are already there. - **Store the `sha256`, not just the path**, in your records; it tells you exactly which content a record refers to. - **Follow `next_cursor` to the end** when listing; short pages are normal. - **Grant `read` unless a partner must write**, and revoke grants you no longer need. - **Do not rely on blobs being private by rule**: anyone who may read the database and knows the hash can read a blob. ## Related - [Data](/en/products/data/): records that point at files with `blob` fields - [Activity log](/en/products/activity/): uploads and downloads with their hashes - [Knowledge and vector search](/en/products/knowledge/): make documents searchable - [Authentication](/en/guides/authentication/), [Errors](/en/guides/errors/), [Limits](/en/guides/limits/) --- Product: Activity log URL: https://developer.inovacc.dev/en/products/activity/ Base URL: https://data.inovacc.dev Read your organization's own log of API calls and file and record events. Endpoints: - GET /v1/activity Example: ```sh curl -X GET "https://data.inovacc.dev/v1/activity" -H "Authorization: Bearer $INOVACC_API_KEY" ``` ## Overview The Activity log is your organization's own record of what happened on Inovacc: every API call, and every file and record event those calls caused. You read it with one endpoint, `GET /v1/activity` on `https://data.inovacc.dev`, filtered by time, kind of event or file hash. The problem it solves is answering "who did what, and when" without building your own audit trail. The log is kept by Inovacc for every call to the Data, AI and Events APIs, whether or not your application remembers to write anything. Recording is a fact of the service, not an option, and an organization only ever sees its own records. Use it for audits, for investigating a failed integration ("which calls were refused this morning, and with what code?"), for proving a file was delivered (the download event carries the file's SHA-256), and for a per-key view of traffic. To see what your AI usage **cost**, use the usage summary of [AI chat](/en/products/ai.chat/) (`GET /v1/usage/summary`): the Activity log records events, never prices. ## Concepts **Events and kinds.** Each record is one event with a `kind`: | Kind | Written when | |---|---| | `api.call` | any authenticated call, refused ones (`403`, `429`) included | | `record.write`, `record.delete` | a record is created, changed or deleted (one per operation of a batch) | | `blob.write`, `blob.read`, `blob.delete` | a database blob is uploaded, downloaded or deleted | | `file.upload.completed`, `file.download`, `file.delete` | a file is written, downloaded or deleted | | `file.upload.started`, `file.processed` | written by other Inovacc services, such as AI uploads and conversions | **What a record holds.** The time (`at_ms`, Unix milliseconds), the organization, who acted (`actor_kind`: `user`, `key` or `service`, and `actor_id`), the `source` service, the `method`, the `route` (a template such as `/v1/databases/{database}/collections/{collection}/records/{record_id}`, never the ids themselves), the `status`, the `outcome`, the `error_code` of a failure, the latency, sizes in and out, and for a file its size, type and SHA-256. `related` points at the database, collection and record, or the upload or job, concerned. Every key is always present, `null` when it does not apply. **What it never holds.** A prompt, a response body or a file's content. AI calls never record a file name. **Country.** An `api.call` record carries `country`: the caller's country as Inovacc's network edge sees it (two letters, `XX` when unknown). It cannot be set by the caller, and the log holds no city, IP address or coordinates. **Retention.** Records are kept one year. The last 30 days answer at once; older records come from the archive and take longer. ## How it works Call `GET /v1/activity` with `Authorization: Bearer `. The credential needs the activity read permission at the organization level; a credential limited to records does not have it. The organization is always the credential's: there is no parameter that names another. Every parameter is optional: | Parameter | Meaning | |---|---| | `from`, `to` | Unix milliseconds, UTC; `from` inclusive, `to` exclusive | | `kind` | one or more kinds, comma separated | | `file_sha256` | 64 lowercase hex characters: every event of that file | | `limit` | 1 to 500, default 100 | | `cursor` | the `next_cursor` of the previous page, unchanged | Any other parameter, a repeated or empty one, or an invalid value is `400 invalid_query`, and nothing is read. The answer is newest first: `{"items":[...],"next_cursor":,"archive_searched":}`. Follow `next_cursor` until it is `null`. `archive_searched` tells you the read reached past the last 30 days. A read can reach back at most 366 days per call. Records appear a few seconds after the call, not instantly. Recording happens after the response and never changes it: if the log cannot be written, the call is answered as usual. A download from Data or Files is recorded when it starts; a download from the AI API is recorded when the transfer completes. The AI API also offers `GET /v1/activity` on `https://ai.inovacc.dev` for its own calls, with `from` and `to` as RFC 3339 instants; reading it there is not billed. ## Get started You need a credential with the activity read permission ([Authentication](/en/guides/authentication/)). 1. **Make a call to log.** Write a record in [Data](/en/products/data/) or upload a file in [Files](/en/products/files/). 2. **Read the latest events.** `GET /v1/activity?limit=10` (see [the samples](#example)). Your call is there as `api.call`, with its route template and status, followed by its `record.write` or `file.upload.completed`. 3. **Filter by kind.** `GET /v1/activity?kind=api.call&limit=50` shows only calls; look at `status` and `error_code` to find refusals. 4. **Follow one file.** Take the `file_sha256` of your upload and call `GET /v1/activity?file_sha256=`: every upload, download and delete of that content is listed. 5. **Page.** Pass the `next_cursor` back as `cursor` until it is `null`. ## Use cases **An audit answer.** A customer asks who deleted a record last Tuesday. Filter `kind=record.delete` with `from` and `to` around that day; the event names the actor and `related` names the database, collection and record. **Proof of delivery.** A compliance team must show a report reached a partner. The `file.download` event carries the SHA-256 and size of the bytes sent, and the time, which can be matched against the original file. **Integration health.** An integration starts failing at night. Filtering `kind=api.call` for that window shows the `status` and `error_code` of every call by the integration's key, for example a run of `429 rate_limited` that points at a missing backoff. ## Limits and pricing | Limit | Value | |---|---| | Retention | one year; the last 30 days answer at once | | Range of one read | 366 days | | Page size | 1 to 500 (default 100) | | Delay before an event is readable | a few seconds | | Rate | 600 requests per minute per credential holder, per location | **Pricing:** on request. ## Errors | Status | Code | What it means and what to do | |---|---|---| | 400 | `invalid_query` | A parameter is unknown, repeated, empty or invalid; fix the query. | | 401 | `invalid_credentials` | Send a valid credential. | | 403 | `forbidden` | The credential lacks the activity read permission. | | 429 | `rate_limited` | Slow down. | | 503 | `service_unavailable` | The log cannot be read now; retry later. | ## Best practices - **Filter on the server**: pass `kind`, `from` and `to` instead of reading everything and filtering in your code. - **Keep the newest `at_ms` you processed** and read from there on the next run, following cursors to the end. - **Use `file_sha256`** to trace a file across uploads, downloads and AI processing. - **Expect a few seconds of delay**; do not treat a missing just-made event as a failure. - **Ignore keys you do not know**: records may gain fields. - **Give each integration its own key**, so `actor_id` tells you which one acted. ## Related - [Data](/en/products/data/), [Files](/en/products/files/), [Events](/en/products/events/) - [AI chat](/en/products/ai.chat/): usage and cost per key and route - [Authentication](/en/guides/authentication/), [Errors](/en/guides/errors/), [Limits](/en/guides/limits/) --- Product: Events URL: https://developer.inovacc.dev/en/products/events/ Base URL: https://events.inovacc.dev Publish events and deliver them to subscribers by webhook, pull or live stream, with dead letters and replay. Endpoints: - POST /v1/workspaces/{workspace_id}/events - GET /v1/workspaces/{workspace_id}/subscriptions - POST /v1/workspaces/{workspace_id}/subscriptions - DELETE /v1/workspaces/{workspace_id}/subscriptions/{subscription_id} - POST /v1/workspaces/{workspace_id}/subscriptions/{subscription_id}/rotate-secret - POST /v1/workspaces/{workspace_id}/subscriptions/{subscription_id}/pull - POST /v1/workspaces/{workspace_id}/subscriptions/{subscription_id}/ack - GET /v1/workspaces/{workspace_id}/subscriptions/{subscription_id}/stream - GET /v1/workspaces/{workspace_id}/dead - POST /v1/workspaces/{workspace_id}/dead/{event_id}/replay - GET /v1/workspaces/{workspace_id}/stats Example: ```sh curl -X POST "https://events.inovacc.dev/v1/workspaces/{workspace_id}/events" -H "Authorization: Bearer $INOVACC_API_KEY" -H "Content-Type: application/json" -d '{}' ``` ## Overview Events lets one part of your system announce that something happened, and lets every interested part hear about it, without the two knowing each other. A publisher sends an event such as `order.paid` to `https://events.inovacc.dev`; every **subscription** whose pattern matches the event's type receives it, by **webhook** to your HTTPS endpoint, by **pull** from your worker, or by **live stream** over a WebSocket. Events that cannot be delivered are kept as **dead letters**, which you can inspect and replay. The problem it solves is reliable fan-out. Calling each interested service directly couples them, loses messages when one is down and retries badly. Events accepts the publish once, delivers it at least once to each subscription, retries failed deliveries with backoff, deduplicates repeated publishes, and keeps what it could not deliver. Use Events to react to what happens in your system. Send an email when an order is paid, refresh a cache when a record changes, push a notification to a page. Keep the state itself in [Data](/en/products/data/) and the bytes in [Files](/en/products/files/): an event should say *what happened*, not carry the object. For the full webhook receiver story, see the [Webhooks guide](/en/guides/webhooks/). ## Concepts **Workspace.** Every route is under `/v1/workspaces/{workspace}`. The workspace in the path is a claim checked against your credential: an application key may name only its own workspace, and any other answers `403 forbidden`, the same as one that does not exist. **Events and the envelope.** You publish `type` (dotted lowercase words, such as `invoice.created`, up to 128 bytes), `data` (any JSON value) and an optional `idempotency_key`. Subscribers receive an envelope with `event_id` (`evt_` and 26 characters, assigned by Inovacc), `type`, `source` (your application, set from the credential), `time`, `tenant`, `idempotency_key` and `data`. **Tiers.** A publish may name a `tier`: `basic` (the default), `reliable` or `advanced`. The tier sets the largest `data` an event may carry: 16 KiB for `basic` and `reliable`, 32 KiB for `advanced`. **Subscriptions and patterns.** A subscription lists 1 to 20 `types` patterns, each `*` (everything), an exact type, or a prefix ending in `.*` (`order.*` matches `order.paid` and `order.item.added`), and one delivery mode. **Delivery modes.** `webhook` posts each event to your HTTPS URL, signed with the subscription's secret. `pull` keeps an inbox your code reads and acknowledges. `websocket` keeps the same inbox and also streams it live; pull keeps working beside the stream. **At least once.** Each event reaches each matching subscription once under normal operation, but a retry or a lost acknowledgement can deliver it again. Receivers deduplicate by `event_id`. **Dead letters.** A webhook delivery that keeps failing is moved to your workspace's dead letters, with the last error and the number of attempts, and kept 30 days. ## How it works Every call carries `Authorization: Bearer `. The credential is the tenant: organization, account and application come from it, and nothing in a body can name another. A **secret key** is for servers and may use every route. A **publishable key** (`apb_...`) may only open a stream, from an origin its application lists. An end user's token may use nothing. Each route needs its own permission, such as `events:event.publish`, `events:subscription.write` or `events:dead.replay`. **Publish.** `POST /events` with `{"event": {...}}` or `{"events": [...]}` (1 to 100). The answer is `202` with `accepted` (each with its `event_id` and `duplicate`) and `rejected` (each with an index and a code); a bad event never blocks the others. Sending an `idempotency_key` already used in the workspace in the last 168 hours returns the original `event_id` with `duplicate: true` and delivers nothing new. **Webhook delivery.** Inovacc posts the envelope to your URL with `x-event-id`, `x-event-type`, `x-event-timestamp` and `x-event-signature: v1=`, an HMAC-SHA256 of the timestamp and the raw body. Any `2xx` within 10 seconds is a delivery; anything else is retried after 10 seconds, the delay doubling up to 10 minutes, until the retry cap of 5 is reached and the event becomes a dead letter. Redirects are never followed, and your URL must be public HTTPS on port 443. **Pull.** `POST /subscriptions/{id}/pull` with `max` (1 to 100, default 10) and `lease_seconds` (5 to 300, default 30) returns the oldest messages, each with its `delivery_count`. A pulled message is hidden for its lease; acknowledge it with `POST .../ack` and the ids, or it returns when the lease ends. **Stream.** `GET /subscriptions/{id}/stream` upgrades to a WebSocket with the subprotocol `inovacc.v1`; a browser passes its key as a second subprotocol, `bearer.`. Each frame is one envelope; reply `{"ack": [...]}` to acknowledge. An unacknowledged frame is sent again after its lease. **Monitor.** `GET /stats` returns `published`, `completed`, `depth`, `oldest_pending_age_seconds`, `retries`, `retry_rate` and `dlq_inflow` for the workspace. Every call is recorded in the [Activity log](/en/products/activity/). What is metered is the delivered event. ## Get started You need a secret key bound to your workspace with the events permissions ([Authentication](/en/guides/authentication/)). 1. **Create a pull subscription.** `POST /v1/workspaces/{workspace}/subscriptions` with `{"types":["demo.*"],"delivery":{"mode":"pull"}}` (see [the samples](#example)). The answer is `201` with a `subscription_id`. 2. **Publish.** `POST .../events` with `{"event":{"type":"demo.hello","data":{"n":1}}}`. The answer is `202` with one `accepted` entry and its `event_id`. 3. **Pull.** `POST .../subscriptions/{id}/pull` with `{}`. Your event is in `messages`, with `delivery_count` 1. 4. **Acknowledge.** `POST .../subscriptions/{id}/ack` with `{"ids":[""]}` answers `{"acked":1}`; a second pull is empty. 5. **Add a webhook.** Create a second subscription with `{"mode":"webhook","url":"https://"}`, keep the `signing_secret` the answer shows once, and follow the [Webhooks guide](/en/guides/webhooks/) to verify deliveries. ## Use cases **Order fulfilment.** The checkout publishes `order.paid` with the order id in `data` and the payment id as `idempotency_key`. A webhook subscription for `order.*` reaches the warehouse system; a pull subscription feeds the invoicing job. A checkout retried by the customer publishes once. **Live updates in a page.** A dashboard opens a stream on a `websocket` subscription for `ticket.*` with a publishable key from its own origin, shows each new ticket as the frame arrives, and acknowledges it. **Recovering from an outage.** A partner's endpoint was down for an hour. Its deliveries ended in dead letters with `last_error` `webhook_timeout`. Once it is back, your team lists `GET /dead` and replays each event; it goes only to the subscriptions that never received it. ## Limits and pricing | Limit | Value | |---|---| | Events per publish | 100 | | Request body | 5 MiB | | `data` per event | 16 KiB (`basic`, `reliable`), 32 KiB (`advanced`) | | Deduplication window for `idempotency_key` | 168 hours | | Subscriptions per workspace; patterns per subscription | 50; 20 | | Messages per pull; lease | 100; 5 to 300 seconds | | Webhook response time | 10 seconds | | Retries | from 10 seconds, doubling to 10 minutes; cap 5, then dead letter | | Dead letters kept | 30 days | | Stream frame from the client | 49,152 bytes | | Rate | 600 calls per minute per credential, per location | **Pricing:** on request. The pricing unit is the **delivered event**. ## Errors | Status | Code | What it means and what to do | |---|---|---| | 400 | `invalid_body` | The body is malformed, or names `source` (it comes from your credential). | | 400 | `invalid_query` | Only `GET /dead` takes a query (`limit`). | | 400 | `invalid_subscription`, `invalid_delivery`, `invalid_limit`, `invalid_id` | Fix the subscription, the delivery or the parameter. | | 400 | `invalid_url`, `url_not_https`, `url_has_credentials`, `url_port_not_allowed`, `url_destination_blocked` | The webhook URL must be public HTTPS on port 443, without credentials. | | 401 | `invalid_credentials` | Send a valid credential. | | 403 | `forbidden`, `account_required` | Missing permission or wrong workspace; use an application key. | | 403 | `publishable_key_not_allowed`, `origin_not_allowed`, `secret_key_in_browser` | A publishable key may only stream, from a listed origin; keep secret keys on servers. | | 404 | `subscription_not_found`, `dead_event_not_found` | It does not exist in your workspace. | | 409 | `wrong_delivery_mode`, `too_many_subscriptions`, `secret_not_managed` | The operation does not fit the delivery mode, or you are at 50 subscriptions. | | 413 | `payload_too_large` | The body is over 5 MiB. | | 426 | `upgrade_required` | The stream route needs a WebSocket upgrade. | | 429 | `rate_limited`, `too_many_streams` | Slow down; close sockets you do not need. | | 503 | `service_unavailable`, `events_disabled`, `signing_unavailable` | Retry later; a publish retried with the same keys reuses its ids. | A rejected event inside a `202` carries its own code: `invalid_event`, `unknown_field`, `invalid_type`, `invalid_idempotency_key`, `payload_too_large` or `data_required`. ## Best practices - **Send an `idempotency_key`** derived from what happened (the payment id, the record version), so a retried publish is not a second event. - **Deduplicate by `event_id`** in every receiver: delivery is at least once. - **Put ids in `data`, not objects**: fetch the current state from [Data](/en/products/data/) when you handle the event. - **Answer webhooks fast** with a `2xx` and do the work afterwards; 10 seconds is the limit. - **Verify every webhook signature** before parsing the body ([Webhooks guide](/en/guides/webhooks/)). - **Acknowledge pulled messages** after processing, and choose a lease longer than your processing time. - **Watch `depth` and `dlq_inflow`** in `/stats`, and replay dead letters once the cause is fixed. ## Related - [Webhooks guide](/en/guides/webhooks/): receive, verify and replay deliveries - [Data](/en/products/data/), [Activity log](/en/products/activity/) - [Authentication](/en/guides/authentication/), [Errors](/en/guides/errors/), [Limits](/en/guides/limits/) --- Guide: Authentication URL: https://developer.inovacc.dev/en/guides/authentication/ # Authentication Every call to an Inovacc API is authenticated by one header, `Authorization: Bearer `. The credential identifies your organization, and for an application's credential the application too, so you never name them in the request. Calls that change something on the AI API also carry an `X-Operation-Id`, which makes them safe to retry. This guide explains the credentials, the headers each API reads, how idempotency works, and how to keep keys safe over their whole life. ## The headers | Header | Where | Value | |---|---|---| | `Authorization` | every call, every API | `Bearer ` | | `X-Operation-Id` | AI API (`ai.inovacc.dev`): every `POST`, `PUT` and `DELETE`; optional on `GET` | a unique id you choose: 8 to 128 letters, digits, `_` or `-` | Data (`data.inovacc.dev`) and Events (`events.inovacc.dev`) read no organization or operation header at all: the tenant comes only from the credential, and any header that tries to name another is ignored. Data uses HTTP versions (`If-Match`, `If-None-Match`) and Events uses an `idempotency_key` in the body for safe retries; see their product pages. A request without a credential, with a malformed one, or with one that is not recognised is refused before anything runs: `401` with `missing_credentials` or `invalid_credentials`. On every API this answer is the same whether the path you called exists or not, so nothing about the API can be learned without a credential. ## Get a key **Secret keys** are created in the Inovacc console by a person of your organization and shown **once**: Inovacc stores only a hash, so a lost key cannot be shown again, only replaced. Access is by invitation today: ask your Inovacc contact to invite you, then create a key under your organization. ## Credentials | Credential | Looks like | For | Accepted by | |---|---|---|---| | Secret API key | `aik_` followed by 64 hex characters | your servers | AI, Data, Events | | OAuth client access token | a signed token issued to an application's client | your servers | AI, Data, Events | | Publishable key | `apb_...` | web pages, from the origins its application lists | Data (records and files, where rules allow), Events (streams only) | | End user access token | issued when a person signs in to your application | that person's browser or app | Data (records and files, by the collection's rules) | A key may be **bound to an application**. Such a key reaches only what that application has enabled: on the AI API each route needs a capability (`ai.chat`, `ai.embeddings`, `knowledge`, `decision`), refused with `403 capability_not_enabled` otherwise; on Data it sees only the databases its application created and its application's file folder; on Events it may name only its own workspace. **Publishable keys are not secrets.** They are meant to ship inside a web page. What protects your data is the list of allowed origins and, above all, the rules of each collection: a publishable key can do nothing a rule does not admit. ## Idempotency On the AI API, `X-Operation-Id` is an idempotency key. The first successful answer to an operation id is stored for 24 hours. Send the same id again and you receive that same answer, marked `x-idempotent-replay: true`, without the work running twice and without a second charge. That is what makes a retry after a timeout or a dropped connection safe. - An **error is never stored**: retrying a failed call with the same id runs it again. - While the first call is **still running**, a second call with the same id is `409 operation_in_progress`; wait and retry. - On chat, embeddings and run, a reused id **returns the stored answer whatever the new body says**. On collection items, the same id with a different file is `409 operation_id_reused`. - Generate a new id for every new operation, and reuse an id only to retry the call it was created for. Every AI answer echoes the id it ran under in `x-operation-id`; on a `GET` without one, the service makes one up. A streamed chat reply is never stored, so it cannot be replayed. ## Example ```sh export INOVACC_API_KEY="" ``` ```sample GET https://ai.inovacc.dev/v1/models ``` The answer lists the routes configured for your organization. Routes are named per organization, so the list you see is yours. A call that changes something adds its operation id: ```sh curl -X POST "https://ai.inovacc.dev/v1/chat/completions" \ -H "Authorization: Bearer $INOVACC_API_KEY" \ -H "Content-Type: application/json" \ -H "X-Operation-Id: $(uuidgen)" \ -d '{"model":"auto","messages":[{"role":"user","content":"Hello"}]}' ``` ## Key lifecycle A key moves through four stages, and each has its own care. 1. **Creation.** Create one key per application and per environment (production, staging, a developer's laptop), never one key shared by everything. Copy the key straight from the console into your secret store. 2. **Use.** Load the key from the environment or a secret manager at start-up. Send it only in the `Authorization` header, only over HTTPS, and only from a server: a secret key that reaches the Data or Events API from a browser (with an `Origin` header) is refused with `403 secret_key_in_browser`. 3. **Rotation.** Create a second key, deploy it everywhere the first is used, confirm in the [Activity log](/en/products/activity/) that the old key's id no longer appears as `actor_id`, then have the old key disabled. A disabled key answers `403 key_disabled`; revoking an application's credential takes effect within a minute. 4. **Retirement.** Remove the key from every secret store and pipeline when an application or environment goes away. One behaviour to plan for: in the Vector Database, the `local` scope belongs to the key that wrote it. A new key sees a new, empty `local` scope; use the `shared` scope for knowledge that must survive a rotation. OAuth access tokens expire. When a call answers `401 invalid_credentials` with a token that used to work, obtain a new token and retry once; do not retry a `401` in a loop. ## Key hygiene - **Never commit a key**, not even to a private repository or a test file. Keep `.env` files out of version control. - **Never log the `Authorization` header**, and redact it from error reports and traces. - **Give each integration its own key**, so its traffic is visible in the Activity log and it can be rotated alone. - **Prefer an application-bound key** with only the capabilities it needs over an organization-wide key. - **Rotate on a schedule**, and at once when someone who had access leaves or a key may have leaked. - **In a browser, use a publishable key or a signed-in person's token**, and close every collection with rules. ## See also - [Errors](/en/guides/errors/): the error envelope and the codes you can handle. - [Limits](/en/guides/limits/): request sizes, timeouts and rates. - [Webhooks](/en/guides/webhooks/): verifying deliveries with a signing secret. - [API reference](/en/reference/): every product and its endpoints. --- Guide: Errors URL: https://developer.inovacc.dev/en/guides/errors/ # Errors When an Inovacc API refuses or cannot complete a call, it answers with an HTTP status and a JSON error envelope. This guide describes the envelope, the error types, the codes that every API shares, and what to do for each status: fix the request, wait, or retry. ## The envelope ```json { "error": { "message": "Human-readable explanation", "type": "validation_error", "code": "invalid_body" } } ``` - **`code`** is what your program should branch on. Codes are part of each API's contract: within a version path such as `/v1`, new codes may be added, and an existing code keeps its meaning. - **`type`** groups codes into families, useful for logging and broad handling. - **`message`** is a fixed English sentence for people. It never repeats a value from your request (at most the name of a refused field) and never describes what failed inside Inovacc. Do not parse it. Some answers add a field to the envelope: a refused file carries `reason` (for example `executable` or `password_protected`), and a schema conversion that does not fit carries `failing_documents`, a count. Ignore fields you do not know. Two answers do not use the envelope. On the AI API, a path that does not exist answers `404 {"error":"not_found"}`. A WebSocket (an Events stream or a Data listener) refused before the upgrade answers an ordinary HTTP error; once open, it ends with a close code. ## Types | `type` | Means | Typical status | |---|---|---| | `validation_error` | The request is malformed: a header, a field or a value breaks the rules. | 400 | | `invalid_request_error` | The request is well formed but cannot be served as sent: too large, an unsupported file type, a URL that may not be fetched. | 400, 413, 415, 422 | | `authentication_error` | No credential, or one that is not valid. | 401 | | `authorization_error`, `permission_error`, `forbidden` | The credential is valid but may not do this. | 403 | | `not_found_error`, `not_found` | The thing does not exist, or you may not know it exists. | 404, 410 | | `conflict_error`, `conflict` | The request conflicts with the current state: a version, a mode, a limit on a collection. | 409 | | `idempotency_error` | An operation id is in use or was used for something else. | 409 | | `rate_limit_error`, `rate_limited` | Too many calls, or an allowance or budget is spent. | 429 | | `api_error`, `service_error`, `service_unavailable`, `internal_error` | Something on Inovacc's side, or a source it called for you, failed. | 500, 502, 503, 504 | The Data, AI and Events APIs use the families above; the exact `type` string can differ between APIs for the same family, so handle the `code`, and fall back on the HTTP status. ## Codes every API shares | Status | Code | Meaning | What to do | |---|---|---|---| | 401 | `missing_credentials`, `invalid_credentials` | No key, or a key that is not valid, expired or revoked. | Fix the credential; do not retry unchanged. | | 403 | `forbidden`, `route_forbidden`, `key_disabled`, `capability_not_enabled` | The credential may not do this. | Ask for the permission, route or capability. | | 403 | `origin_not_allowed`, `secret_key_in_browser` | A browser call from an origin not listed, or a secret key in a browser. | List the origin; use a publishable key. | | 400 | `missing_operation_id`, `invalid_operation_id` | AI: a `POST`, `PUT` or `DELETE` without a valid `X-Operation-Id`. | Send 8 to 128 letters, digits, `_` or `-`. | | 400 | `invalid_body`, `unsupported_field`, `invalid_field` | The body or a parameter is not valid for this endpoint. | Fix the request. | | 404 | `*_not_found` | It does not exist, or it is not yours. | Check the id. | | 409 | `operation_in_progress` | The same operation id is still running. | Wait, then retry with the same id. | | 409 | `operation_id_reused` | AI: an operation id already used for a different file in a collection. | Use a new operation id for a new operation. | | 409 | `version_conflict` | Data: the record changed since you read it. | Re-read, then retry. | | 413 | `request_too_large`, `payload_too_large` | The body is over the size limit. | Make it smaller, or upload in parts. | | 429 | `rate_limited` | Too many requests in the window. | Wait `Retry-After` seconds, when given. | | 429 | `quota_exceeded`, `concurrency_limited` | AI: a token, cost or concurrency allowance is used up. | Slow down; raise the allowance with Inovacc. | | 429 | `budget_exhausted` | AI: a daily budget is spent. | Wait until 00:00 UTC (`Retry-After`). | | 502 | `upstream_error` | A source called for you failed. | Retry with backoff. | | 503 | `service_unavailable` | A dependency could not be reached; nothing was done or guessed. | Retry with backoff. | | 504 | `timeout_error` | AI: the model did not start answering in time. | Retry with backoff. | Each product page lists every code it can return, with what to do. ## Retry policy | Status | Retry? | How | |---|---|---| | 400, 401, 403, 404, 410, 413, 415, 422 | No | The same request will fail the same way. Fix it first. | | 409 `operation_in_progress` | Yes | After a short wait, with the **same** operation id. | | 409 `version_conflict` | Yes, after re-reading | Read the record, apply your change to the new version, send its `If-Match`. | | 409, other codes | No | The state must change first (create the database, wait for items, empty the folder). | | 429 with `Retry-After` | Yes | Not before the number of seconds it gives. | | 429 without `Retry-After` | Later | An allowance is used up: back off for minutes, not seconds, or raise it. | | 500, 502, 503, 504 | Yes | With exponential backoff and jitter (for example 1, 2, 4, 8 seconds), a few times at most. | | Network error, no answer | Yes | As for 503. | **Retry safely.** On the AI API, retry a mutating call with the **same** `X-Operation-Id`: if the first attempt did succeed, you receive its stored answer instead of a second run. On Data, retry a create with your own record id or `If-None-Match: *`, and an update with the same `If-Match`. On Events, retry a publish with the same `idempotency_key`: the original `event_id` comes back with `duplicate: true`. ## See also - [Authentication](/en/guides/authentication/): credentials and idempotency. - [Limits](/en/guides/limits/): the sizes and rates behind `413` and `429`. - [Webhooks](/en/guides/webhooks/): how failed deliveries are retried. --- Guide: Limits URL: https://developer.inovacc.dev/en/guides/limits/ # Limits Every Inovacc API bounds the size of what you send, how long a call may take and how often you may call. This guide lists those limits per API and says what happens when you reach each one, so your code can stay under them or handle the refusal. Values marked "per organization" are set for your organization by Inovacc; the others are fixed. ## How limits answer | You reach | Answer | What to do | |---|---|---| | A size limit (body, file, record, item) | `413` (`request_too_large`, `payload_too_large`, `url_too_large`), or `400` for a field or document that is too large | Send less, or use the upload in parts | | A count limit (documents per call, operations per batch, items per collection) | `400 invalid_body`, `400 batch_too_large` or `409 collection_full` | Split the call; start a new collection | | A rate limit | `429 rate_limited`, with `Retry-After` on the AI API | Wait, then retry | | A token, cost or concurrency allowance (AI) | `429 quota_exceeded` or `429 concurrency_limited`, no `Retry-After` | Slow down for minutes, or raise the allowance | | A daily budget (AI) | `429 budget_exhausted`, `Retry-After` until 00:00 UTC | Wait for the next UTC day | | A per-request input cap (AI) | `400 input_too_large` | Shorten the input | A call refused at a limit is refused before its work starts: nothing is called on your behalf and nothing is written. ## AI | Limit | Value | |---|---| | Request body (chat, embeddings, run) | 1 MiB by default; per organization, 1 KiB to 20 MiB | | Input tokens per request | per organization (`400 input_too_large`) | | Output tokens per request | the lower of the route's and your organization's ceiling; a higher `max_tokens` is lowered, not refused | | Time until the model starts answering | 60 seconds (`504 timeout_error`); not on `/v1/run` | | Images in one chat request | 20 | | Estimated input cost of one image | 1,600 tokens | | Stored answer for an idempotent replay | 24 hours | | A call held as in progress | 5 minutes | | Vector Database: documents per ingest, ids per delete | 100 | | Vector Database: one document (`id` + `text` + `metadata`) | 10 KiB | | Vector Database: request body | 1 MiB | | Vector query: text; `top_k` | 8,000 characters; 50 | | Collection item in a request body, or by URL | 16 MiB | | Collection | 100 items; 1 GiB of files sent in requests; 20 GiB of video by upload | | Upload part | 64 MiB, every part but the last | | Upload that is not a video (transcript, notes) | 16 MiB | | Recording upload | per organization | | Knowledge job result | kept 24 hours | | Activity and usage reads | up to 366 days and 92 days per call | **Quotas per organization.** Each of these is set for your organization, and unlimited when not set: | Allowance | Window | When reached | |---|---|---| | Requests | per UTC clock minute, per UTC day | `429 rate_limited`, `Retry-After` to the next minute or day | | Tokens | per UTC day, per UTC month | `429 quota_exceeded` | | Cost | per UTC day, per UTC month | `429 quota_exceeded` | | Concurrent calls | at any moment | `429 concurrency_limited` | | Daily budget, per provider and optionally per application | per UTC day | `429 budget_exhausted`, `Retry-After` to 00:00 UTC | A request counts against the per-minute and per-day request limits as soon as it is admitted, even if it fails later. Token and cost limits allow a call that reaches the limit exactly. From a budget's warning level (80% by default), answers carry `X-Budget-Warning` before calls start being refused. ## Data and Files | Limit | Value | |---|---| | JSON request body | 1 MiB (`413 payload_too_large`) | | One record | 900 KiB (`400 record_too_large`) | | A `json` field | 64 KiB | | Fields per collection; collections per database | 64; 64 | | Records per page | 200 (default 50) | | Operations per batch | 100, one transaction | | Filter | 2,048 characters, 40 values, nesting depth 8 | | Query string | 16 KiB (`400 invalid_query`) | | File or blob in one request | 100 MiB, with `Content-Length` (`411` without it) | | File or blob in parts | 5 GiB; parts of 5 MiB to 100 MiB, numbered 1 to 10,000 | | File path; listing prefix; file metadata | 1,024 bytes; 256 bytes; 2 KiB | | Realtime listeners per collection; one event | 1,000; 512 KiB | | Activity log page | 1 to 500 (default 100) | | Rate | 600 requests per minute per credential holder, per location (`429 rate_limited`) | The rate limit is counted per location and is a brake, not an exact count. ## Events | Limit | Value | |---|---| | Request body | 5 MiB (`413 payload_too_large`) | | Events per publish | 100 | | `data` per event | 16 KiB (`basic`, `reliable`), 32 KiB (`advanced`) | | Deduplication window of an `idempotency_key` | 168 hours | | Subscriptions per workspace; patterns per subscription | 50; 20 | | Messages per pull; pull lease | 100; 5 to 300 seconds | | Webhook answer time | 10 seconds | | Webhook retries | from 10 seconds, doubling to 10 minutes; cap 5 | | Dead letters kept | 30 days | | Rate | 600 calls per minute per credential, per location (`429 rate_limited`) | ## Staying under the limits - **Batch**: send up to 100 documents, operations or events per call instead of one each. - **Upload large files in parts** instead of raising body sizes. - **Back off on `429`**: honour `Retry-After`, and add jitter so many clients do not retry at the same instant. - **Set `max_tokens`** on chat calls: it lowers what each call reserves against your quotas. - **Watch `X-Budget-Warning`** and the AI usage summary to see an allowance coming before it is spent. ## See also - [Errors](/en/guides/errors/): every code and the retry policy. - [Authentication](/en/guides/authentication/): credentials and idempotency. - Each product page lists its own limits in full. --- Guide: Webhooks URL: https://developer.inovacc.dev/en/guides/webhooks/ # Webhooks A webhook is the way [Events](/en/products/events/) delivers to a server you run: every event that matches a subscription arrives at your HTTPS endpoint as a signed `POST`. This guide covers creating a webhook subscription, the request your endpoint receives, verifying its signature, how retries and dead letters work, and rotating the signing secret without dropping a delivery. ## What a webhook subscription is A subscription says which events a receiver wants and how they reach it. A **webhook subscription** has a list of type patterns and a delivery block with `"mode": "webhook"` and your `url`. From then on, each event published in the workspace whose `type` matches one of the patterns is posted to that URL, once per subscription. Delivery is **at least once**. Under normal operation each event arrives once; a retry, a slow answer or a lost connection can make it arrive again. Your receiver is therefore built to accept a duplicate harmlessly, keyed on the event id. The URL must be `https`, on port 443, with no user name or password in it, no fragment, and a public host: private, loopback, link-local and local-use names are refused when the subscription is created and checked again at every delivery. Redirects are never followed. ## Create one Call `POST https://events.inovacc.dev/v1/workspaces/{workspace}/subscriptions` with a secret key that has the subscription write permission: ```json {"types": ["order.*"], "delivery": {"mode": "webhook", "url": "https://hooks.example.com/inovacc"}} ``` The answer is `201`: ```json {"subscription_id": "sub_...", "types": ["order.*"], "delivery": {"mode": "webhook", "url": "https://hooks.example.com/inovacc", "secret_ref": "managed"}, "created_at": "2026-10-06T12:00:00.000Z", "signing_secret": "whsec_..."} ``` `signing_secret` is the key your endpoint uses to verify deliveries. **It is shown once, in this answer, and never again**: listings show `"secret_ref": "managed"` and nothing more. Store it in your secret manager at once. If it is lost, rotate it (below). A workspace holds at most 50 subscriptions, each with 1 to 20 patterns. ## The delivery request Each delivery is: ```http POST https://hooks.example.com/inovacc content-type: application/json x-event-id: evt_01K... x-event-type: order.paid x-event-timestamp: 1791201600 x-event-signature: v1=<64 lowercase hex characters> {"event_id":"evt_01K...","type":"order.paid", ... } ``` `x-event-timestamp` is the Unix time in seconds at which the delivery was signed. During a secret rotation, `x-event-signature` carries one entry per live secret, newest first, separated by commas: `v1=,v1=`. Answer with any `2xx` status to confirm the delivery. Anything else (a `4xx`, a `5xx`, no answer within 10 seconds, a network error) is a failed attempt and will be retried. ## The event envelope The body is the event envelope, exactly as subscribers of every mode receive it: | Field | Meaning | |---|---| | `event_id` | `evt_` followed by 26 characters, assigned by Inovacc when the event was published. Unique per event. | | `type` | The event's type, such as `order.paid`: dotted lowercase words, up to 128 bytes. | | `source` | The application that published the event, taken from its credential. | | `time` | When it was published, RFC 3339 UTC with milliseconds. | | `tenant` | The organization, account, workspace and, when known, application the event belongs to. | | `idempotency_key` | The publisher's key; when it sent none, the `event_id`. | | `data` | The publisher's JSON payload: what happened, usually ids rather than whole objects. | Read only the fields you need, and ignore fields you do not know. ## Verify the signature Verify every delivery **before** you parse or act on it: 1. Read `x-event-timestamp` and the **raw body bytes**, exactly as received, before any JSON parsing. 2. Compute `HMAC-SHA256(secret, timestamp + "." + body)`, where `secret` is the whole `whsec_...` string as UTF-8, and write it as lowercase hex. 3. Split `x-event-signature` on commas and accept the delivery when **any** `v1=` entry equals your value, comparing in constant time. 4. Reject a timestamp more than five minutes away from your clock, so an old delivery cannot be replayed at you. **Worked example.** With the secret `example-key-0001`, the timestamp `1791201600` and the body `{"hello":"world"}`, the signed text is `1791201600.{"hello":"world"}` and the signature is the 64 hex characters below, printed here in two halves that you join without a space: ```text 05080ab29a6abfb44033966c7837c2a1 889fbf1a909f9cbd12b9dd5d12bb2bf5 ``` so the header reads `v1=` followed by both halves. Run your verifier on this example before you deploy it. **TypeScript** (Node.js 18 or later): ```ts import { createHmac, timingSafeEqual } from "node:crypto"; export function verifyInovaccWebhook( secret: string, timestamp: string, signatureHeader: string, rawBody: Buffer, nowSeconds: number = Math.floor(Date.now() / 1000), ): boolean { const ts = Number(timestamp); if (!Number.isInteger(ts) || Math.abs(nowSeconds - ts) > 300) return false; const expected = createHmac("sha256", secret) .update(`${timestamp}.`) .update(rawBody) .digest(); return signatureHeader.split(",").some((entry) => { const part = entry.trim(); if (!part.startsWith("v1=")) return false; const given = Buffer.from(part.slice(3), "hex"); return given.length === expected.length && timingSafeEqual(given, expected); }); } ``` **Python** (3.10 or later): ```python import hashlib import hmac import time def verify_inovacc_webhook(secret: str, timestamp: str, signature_header: str, raw_body: bytes, now: int | None = None) -> bool: try: ts = int(timestamp) except ValueError: return False now = int(time.time()) if now is None else now if abs(now - ts) > 300: return False expected = hmac.new(secret.encode(), timestamp.encode() + b"." + raw_body, hashlib.sha256).hexdigest() return any( part.strip().startswith("v1=") and hmac.compare_digest(part.strip()[3:], expected) for part in signature_header.split(",") ) ``` Both functions take the raw body: configure your web framework to hand you the bytes, not a parsed object, on this route. ## Retries and timeouts Your endpoint has **10 seconds** to answer. A failed attempt is retried after 10 seconds, and the delay doubles on each further failure up to 10 minutes. When the retry cap of 5 is reached, the event moves to your workspace's dead letters. A failure affects only the delivery that failed: other events, and other subscriptions of the same event, keep flowing. Because a delivery can be repeated, a `2xx` that arrives after your endpoint already processed the event is harmless as long as you deduplicate by `x-event-id`. ## Rotate the secret `POST /v1/workspaces/{workspace}/subscriptions/{id}/rotate-secret` returns a new secret, once: ```json {"subscription_id": "sub_...", "signing_secret": "whsec_...", "previous_valid_until": "2026-10-07T12:00:00.000Z"} ``` Until `previous_valid_until` (24 hours by default), **both secrets sign**: each delivery carries `v1=,v1=`. Deploy the new secret to your receiver within that window; because the verifier accepts any matching entry, nothing is dropped while you switch. Rotating again inside the window keeps every unexpired secret. Rotation applies to webhook subscriptions only (`409 wrong_delivery_mode` otherwise). ## Dead letters and replay `GET /v1/workspaces/{workspace}/dead?limit=50` lists dead letters, newest first, with the event, `attempts`, `dead_at`, `replay_count` and `last_error`: | `last_error` | Meaning | |---|---| | `webhook_status_` | Your endpoint answered that status, for example `webhook_status_500`. | | `webhook_timeout` | No answer within 10 seconds. | | `webhook_network` | The connection failed. | | `url_destination_blocked` | The URL now points somewhere deliveries may not go. | | `secret_missing` | The subscription's signing secret could not be used. | Fix the cause, then `POST /v1/workspaces/{workspace}/dead/{event_id}/replay`. The answer is `202`, and the same envelope, with the same `event_id`, is delivered again to the subscriptions that never received it. Dead letters are kept 30 days. ## Best practices - **Verify before parsing**: compute the signature over the raw bytes, then parse. - **Make the receiver idempotent**: store each processed `x-event-id` and answer `2xx` at once for one you have seen. - **Answer fast**: confirm with a `2xx` and do slow work after, from your own job. - **Reject stale timestamps** (more than five minutes off) and keep your server clock synchronised. - **Keep the secret in a secret manager** and rotate it when someone who knew it leaves, or on a schedule. - **Monitor dead letters** and `dlq_inflow` in `GET /stats`; replay once the endpoint is healthy. - **Return `2xx` only when you have the event safely**: a `2xx` ends the retries. ## See also - [Events](/en/products/events/): publishing, subscriptions, pull and stream - [Errors](/en/guides/errors/), [Authentication](/en/guides/authentication/), [Limits](/en/guides/limits/) --- Guide: Inovacc MCP URL: https://developer.inovacc.dev/en/guides/connect-your-ai-agent/ # Inovacc MCP The [Model Context Protocol](https://modelcontextprotocol.io) (MCP) is an open standard that lets AI assistants call tools exposed by a server. The Inovacc MCP server is a program, `inovacc mcp serve`, that runs on your machine with this documentation built in, so your assistant can look up Inovacc products, endpoints and guides and answer with their sources. It works with any MCP-compatible client, including Cursor, Claude Code, VS Code (GitHub Copilot), Windsurf, Cline, Claude Desktop, JetBrains IDEs, Codex CLI and Gemini CLI. To let your agent do the whole setup itself, give it this line: ```text Fetch and execute the appropriate instructions to set me up for Inovacc from https://developer.inovacc.dev/agent-setup/prompt.md ``` ## Capabilities With the server connected, your assistant can: - search the documentation: contract sections, products, guides and quickstarts - explore endpoints: the contract section for any method and path, such as `POST /v1/vector/query` - get a quickstart for a product: base URL, headers and a curl skeleton - tell you which documentation version it is answering from It does not execute requests and does not touch your account. The server is read-only, needs no API key, and answers from the documentation embedded in the program. ## Prerequisites - Windows on x86_64. macOS and Linux are coming; there is no date yet. - An MCP-compatible client from the list below. - The release files for the `inovacc` program. The repository is private until launch: download the release your Inovacc contact shares with you. ## Install the inovacc CLI 1. Download `inovacc-windows-x86_64.exe` and `inovacc-windows-x86_64.exe.sha256` from the release. 2. Check the download in PowerShell. The two values must be identical: ```powershell (Get-FileHash .\inovacc-windows-x86_64.exe -Algorithm SHA256).Hash.ToLower() Get-Content .\inovacc-windows-x86_64.exe.sha256 ``` 3. Rename the file to `inovacc.exe` and put it in a folder on your `PATH`. 4. Open a new terminal and verify: ```powershell inovacc docs version ``` The output has this shape; the version numbers differ by release: ```text embedded 2026.10.09-0 installed none active embedded 2026.10.09-0 previous none published unknown; run `inovacc docs update --check` data dir C:\Users\\AppData\Local\Inovacc\knowledge ``` ## Connect your editor or chat app Every client needs the same three facts: the transport is stdio, the command is `inovacc`, and the arguments are `mcp` and `serve`. Restart the client after changing its configuration. ### Cursor Create `.cursor/mcp.json` in a project, or `~/.cursor/mcp.json` for every project: ```json { "mcpServers": { "inovacc": { "command": "inovacc", "args": ["mcp", "serve"] } } } ``` ### Claude Code ```sh claude mcp add inovacc -- inovacc mcp serve ``` The `--` separates Claude's own options from the command that runs the server. By default the server is added in the `local` scope: available only to you, in the current project. Choose another scope with `-s` before the name: ```sh claude mcp add -s project inovacc -- inovacc mcp serve claude mcp add -s user inovacc -- inovacc mcp serve ``` `project` writes `.mcp.json` in the project root, shared with everyone who uses the repository. `user` makes the server available in all your projects. ### VS Code (GitHub Copilot) Create `.vscode/mcp.json` in a workspace, or run **MCP: Open User Configuration** for your user profile: ```json { "servers": { "inovacc": { "type": "stdio", "command": "inovacc", "args": ["mcp", "serve"] } } } ``` **Important:** `.vscode/mcp.json` uses `servers`, not `mcpServers`. VS Code's own documentation sets `"type": "stdio"` in its stdio examples, so this page does too. A project-root `.mcp.json` uses `mcpServers`, like the other clients. ### Windsurf Windsurf's documentation now lives with Devin Desktop. Add the server to `mcp_config.json`, at `%APPDATA%\devin\mcp_config.json` on Windows: ```json { "mcpServers": { "inovacc": { "command": "inovacc", "args": ["mcp", "serve"] } } } ``` ### Cline In the Cline panel, click the MCP Servers icon in the top toolbar, open the Configure tab and click **Configure MCP Servers**. Add the server to the settings file that opens (`cline_mcp_settings.json`): ```json { "mcpServers": { "inovacc": { "command": "inovacc", "args": ["mcp", "serve"] } } } ``` ### Claude Desktop Open the Claude menu, then **Settings**, the **Developer** tab and **Edit Config**. The file is `%APPDATA%\Claude\claude_desktop_config.json` on Windows: ```json { "mcpServers": { "inovacc": { "command": "inovacc", "args": ["mcp", "serve"] } } } ``` Quit Claude Desktop completely and start it again. If the app cannot find `inovacc`, give the full path to `inovacc.exe` as the command, with each backslash doubled in JSON. ### JetBrains IDEs Open **Settings | Tools | AI Assistant | Model Context Protocol (MCP)**, click **Add**, choose the **STDIO** connection type and enter: ```json { "mcpServers": { "inovacc": { "command": "inovacc", "args": ["mcp", "serve"] } } } ``` Click **OK**, then **Apply**. ### Codex CLI ```sh codex mcp add inovacc -- inovacc mcp serve ``` or write it in `~/.codex/config.toml` (a trusted project can use `.codex/config.toml`): ```toml [mcp_servers.inovacc] command = "inovacc" args = ["mcp", "serve"] ``` **Important:** Codex uses TOML and `mcp_servers`, not JSON and `mcpServers`. ### Gemini CLI ```sh gemini mcp add inovacc inovacc mcp serve ``` or add the server to `~/.gemini/settings.json` (a project can use `.gemini/settings.json`): ```json { "mcpServers": { "inovacc": { "command": "inovacc", "args": ["mcp", "serve"] } } } ``` ### Any other client Use the stdio transport with the command `inovacc` and the arguments `mcp` and `serve`. Check your client's MCP documentation for where that goes. ## Verify your setup Ask your assistant: "What does POST /v1/vector/query take?" A working setup answers from the contract and names the document it came from and the knowledge version. Without an assistant, search from the terminal: ```sh inovacc docs search "vector query" ``` Each hit shows the document id, its title and its citation (repository, path and source commit). ## Troubleshooting - **`inovacc` is not found.** The folder holding `inovacc.exe` is not on your `PATH`, or the terminal was opened before you changed it. Open a new terminal. In the client's configuration you can give the full path to `inovacc.exe` as the command. - **The client does not list the server.** Restart the client completely, then check that the command in its configuration runs in a terminal: `inovacc mcp serve` should wait silently for input (press Ctrl+C to leave). - **Answers look old.** Run `inovacc docs version` to see what is in use and what is published, then `inovacc docs update`. - **Windows SmartScreen warns on first run.** The program is not code-signed today, so Windows may show "Windows protected your PC". Check the checksum as described above before you choose to run it. ## Available MCP tools | Tool | What it does | |---|---| | `search_documentation` | Full-text search over contract sections, capabilities, quickstarts and guides | | `get_document` | One document in full, by id, with its provenance | | `get_api_reference` | The contract section for one endpoint, such as `POST /v1/vector/query` | | `list_products` | The products and capabilities with documentation, and how many documents each has | | `get_quickstart` | Base URL, endpoints, scopes, the headers every call needs and a curl skeleton | | `get_knowledge_version` | Which documentation bundle is answering: version, build time, pinned commits and document count | ## Keep the documentation current The program carries the documentation it was built with and can install newer, signed documentation without a new program. ```sh inovacc docs version # embedded, installed, active, previous and published versions inovacc docs update # download, verify and install the newest published documentation inovacc docs update --check inovacc docs rollback # go back to the previous installed documentation ``` - Documentation is published as signed bundles. A bundle is installed only if its size, SHA-256 and Ed25519 signature check out, and a bundle that is not newer than the one in use is refused, so a downgrade cannot be forced. - The automatic check runs at most once every 24 hours, never delays an answer, installs nothing and at most prints one line telling you a newer version exists. - Until launch, updates reach you through the release your Inovacc contact shares with you. ## Configuration reference | Setting | What it does | |---|---| | `INOVACC_KNOWLEDGE_DIR` | The folder where installed documentation is kept. Default on Windows: `%LOCALAPPDATA%\Inovacc\knowledge` | | `INOVACC_NO_UPDATE_CHECK` | Set to `1` to turn off the automatic check for newer documentation. The flag `--no-update-check` does the same for one run | | `INOVACC_UPDATE_URL` | The address of `latest.json` that updates are read from; it must be HTTPS | ## FAQ **Do I need an API key?** No. The server reads documentation only and never uses your account. **Does it send my questions anywhere?** No. Questions are answered on your machine. The only network use is the automatic check for newer documentation, at most once a day, which you can turn off. **Can my assistant call the Inovacc API through it?** No. It tells the assistant how an endpoint works and how to call it; it does not make the call. **Is there a hosted MCP server?** A remote server at `https://mcp.inovacc.dev` is planned. It does not exist yet. **Which platforms are supported?** Windows x86_64 today. macOS and Linux are coming, with no date yet. ## See also - [Authentication](/en/guides/authentication/): the headers every API call carries. - [API reference](/en/reference/): every product and its endpoints. - [llms.txt](/llms.txt): the same documentation as an index for language models.