# AI chat

Chat completions from the platform's language models.

Base URL: `https://ai.inovacc.dev`

## Overview

AI chat gives your application conversational answers from the language models Inovacc runs for your organization. The endpoint, `POST /v1/chat/completions`, speaks the OpenAI chat-completions wire format, so an existing OpenAI client library can call it by changing the base URL to `https://ai.inovacc.dev/v1` and adding two headers. You send a list of messages and get back the assistant's reply, either as one JSON document or as a stream of events while it is being written.

The problem it solves is access without plumbing: your organization does not hold model-provider accounts, keys or contracts. Inovacc decides which model serves each of your routes, applies your organization's limits and budgets before a call is made, and records the usage of every call so you can see what each key and route spends.

Use AI chat when the answer is free text: an assistant, a summary, a draft, an extraction into JSON, a reply that calls your tools. Another product fits better when:

- you need a structured verdict against criteria (a score, a choice, a yes or no, with confidence): use the [Decision engine](/en/products/decision/);
- the answer must come from your own documents, with citations: ask a cloud collection of [Knowledge and vector search](/en/products/knowledge/);
- you need vectors rather than text: use [AI embeddings](/en/products/ai.embeddings/).

## Concepts

**Routes, not models.** You never name a model. Your organization is given **routes**: opaque ids such as `r00`, configured by Inovacc for your organization alone. A route decides which model answers, its output ceiling and the fields it accepts. `GET /v1/models` lists the routes your organization may use, each with the `endpoint` it serves; the `id` is always the route id, never the name of what backs it. Send the id as the `model` field; send `"auto"`, or leave `model` out, and your organization's default route answers. A route serves exactly one endpoint: a chat route called from `/v1/embeddings` or `/v1/run` is refused like a route you are not allowed to use.

**Fields.** The common OpenAI fields are always accepted: `messages`, `model`, `stream`, `stream_options`, `max_tokens`, `max_completion_tokens`, `temperature`, `top_p`, `stop`, the penalties, `logit_bias`, `logprobs`, `tools`, `tool_choice`, `response_format`, `seed`, `reasoning_effort` and a few more. Seven fields (`n`, `service_tier`, `store`, `metadata`, `modalities`, `audio`, `web_search_options`) are accepted only when your organization has been granted them. Any other field is refused by name. Some routes take a narrower set (`messages`, `model`, `stream`, `stream_options`, the two output limits, `temperature`, `top_p`); on those, `tools` and the other fields are refused for that route.

**Messages and images.** `content` is a string, or an array of text parts and `image_url` parts. An image is a `data:` URL (PNG, JPEG, WebP or GIF, base64) or an `https://` URL; a route whose model cannot fetch a URL refuses the `https://` form. One request carries at most 20 images, and each counts as an estimated 1,600 input tokens.

**Output ceiling.** `max_tokens` and `max_completion_tokens` can only lower the ceiling the route and your organization set; a higher value is clamped, never refused. Reasoning tokens a model spends are inside `completion_tokens`.

**Streaming.** With `"stream": true` the answer is a `text/event-stream` of OpenAI `chat.completion.chunk` events, each a `data:` line, ending with `data: [DONE]`. Set `stream_options.include_usage` to receive a final chunk with the token counts.

**Usage headers.** A non-streaming answer carries `x-route-id`, `x-usage-input-tokens` and `x-usage-output-tokens`, and every answer carries `x-operation-id`.

## How it works

Every call carries `Authorization: Bearer <key>` and an `X-Operation-Id`. Your organization comes from the key. The operation id is your idempotency key: 8 to 128 letters, digits, `_` or `-`. The service checks, in order: the operation id, the key, whether that operation id was already used, the body and its fields, the route, the messages, the estimated input size (a quarter of the characters of the whole request body, plus the image estimate) and then reserves the call against your organization's quotas. Only when all of that passes is a model called.

The answer's `id` is always `chatcmpl-` followed by your operation id, and its `model` is always the route id. A reply whose whole output budget went to reasoning is still `200`, with empty `content` and `finish_reason: "length"`; a reply stopped by the model's safety filter is `200` with `finish_reason: "content_filter"`.

A successful non-streaming answer is stored for 24 hours under its operation id. Send the same id again and you receive the stored answer, with `x-idempotent-replay: true`, without a second model call, a second charge or a quota change. An error is never stored, so retrying a failed call with the same id runs it again. While a call is still running, a second call with its id is `409 operation_in_progress`.

A stream is never stored. If the model fails partway, the stream ends with one error event (`upstream_error`) and `data: [DONE]`; a stream that ends without `[DONE]` was cut and is incomplete. Read `delta.content` on every chunk, including the first and the one that carries `finish_reason`.

What is metered is tokens: input and output, as reported by the model, or the estimate when no count is available. `GET /v1/usage/summary` sums your organization's usage per key, route, day or hour, and `GET /v1/capabilities` tells an agent what your organization may call; neither is charged.

## Get started

You need an API key and at least one chat route enabled for your organization. Keys are created in the Inovacc console; see [Authentication](/en/guides/authentication/).

1. **List your routes.** Call `GET /v1/models`. Each entry with `"endpoint": "/v1/chat/completions"` is a route you can chat with; note its `id`.
2. **Send a first message.** Call `POST /v1/chat/completions` with that id as `model`, one `user` message and a fresh `X-Operation-Id` (see [the samples](#example)). The answer's `id` ends with your operation id and `x-usage-input-tokens` shows what was counted.
3. **Retry it.** Send the same request with the same operation id. The answer is identical and carries `x-idempotent-replay: true`: nothing ran twice.
4. **Stream.** Add `"stream": true` and `"stream_options": {"include_usage": true}` with a new operation id. Chunks arrive as the reply is written, the last before `[DONE]` carrying `usage`.
5. **Check the spend.** Call `GET /v1/usage/summary?group_by=route` to see the calls and tokens you just made.

## Use cases

**A support assistant.** Your help widget streams the reply so the person sees it appear at once. Each user turn is one call with a new operation id; your server keeps the conversation and sends the recent messages each time. If the connection drops before `[DONE]`, the widget shows the reply as cut and offers to ask again.

**Structured extraction in bulk.** A nightly job reads incoming documents and asks for a JSON object with `response_format`. Each document gets an operation id derived from its own id, so when the job is restarted after a crash, finished documents come back from the stored answers in seconds and are not charged again.

**An agent that calls your tools.** On a route that accepts `tools`, the model answers with `tool_calls`; your code runs the tool and sends the result back as a `tool` message. Each step is its own operation id, so a retried step never runs a tool twice on the model's side.

## Limits and pricing

| Limit | Value |
|---|---|
| Request body | 1 MiB by default; your organization may be set between 1 KiB and 20 MiB |
| Images per request | 20 |
| Estimated input cost of one image | 1,600 tokens |
| Input tokens per request | your organization's cap (`400 input_too_large` above it) |
| Output tokens per request | the lower of the route's and your organization's ceiling |
| Time to the model's first response | 60 seconds |
| Stored answer for a replay | 24 hours |
| A call held as in progress | 5 minutes at most |
| Requests per minute, per day; tokens and cost per day and month; concurrent calls; daily budgets | as set for your organization |

**Pricing:** on request. The pricing unit is **tokens**.

## Errors

| Status | Code | What it means and what to do |
|---|---|---|
| 400 | `missing_operation_id`, `invalid_operation_id` | Send `X-Operation-Id` with 8 to 128 letters, digits, `_` or `-`. |
| 400 | `invalid_body` | The body is not a JSON object, or `user` is not a string. Fix the body. |
| 400 | `invalid_messages` | `messages` is empty or a message is malformed. |
| 400 | `unsupported_field` | A top-level field this endpoint does not accept; remove it. |
| 400 | `unsupported_field_for_route` | The route takes a narrower set of fields, or cannot fetch an image URL; remove the field or send the image as `data:`. |
| 400 | `model_not_allowed` | The `model` value is not a valid route id. Use an id from `GET /v1/models`. |
| 400 | `input_too_large` | The input is over your organization's per-request cap, or there are more than 20 images. Shorten it. |
| 401 | `missing_credentials`, `invalid_credentials` | Send a valid key. |
| 403 | `key_disabled`, `organization_disabled` | The key or organization was disabled; contact your administrator. |
| 403 | `field_not_allowed` | A field your organization has not been granted. |
| 403 | `route_forbidden` | The route is not one your organization may use on this endpoint. |
| 409 | `operation_in_progress` | That operation id is still running; wait and retry. |
| 413 | `request_too_large` | The body is over your size cap. |
| 429 | `rate_limited` | Too many requests; wait `Retry-After` seconds. |
| 429 | `quota_exceeded`, `concurrency_limited` | A token, cost or concurrency allowance is used up; no `Retry-After`. |
| 429 | `budget_exhausted` | A daily budget is spent; it resets at 00:00 UTC (`Retry-After`). |
| 502 | `upstream_error` | The model did not answer usably; retry with the same operation id. |
| 503 | `service_unavailable`, `route_unavailable` | Retry later; `route_unavailable` needs Inovacc to fix your route. |
| 504 | `timeout_error` | The model did not start answering in 60 seconds; retry. |

The envelope and every type are in [Errors](/en/guides/errors/).

## Best practices

- **One operation id per logical request**, reused only to retry that same request. A reused id returns the first answer whatever the new body says.
- **Retry on `502`, `503`, `504` and on network errors with the same operation id**, with backoff. On `429 rate_limited` and `budget_exhausted`, wait the `Retry-After` seconds.
- **Set `max_tokens`** to what you need: it lowers the reservation made against your quotas and the cost of a runaway reply.
- **Watch `X-Budget-Warning`**: it appears once your organization is past a budget's warning threshold, before calls are refused.
- **Use `stream` for people, not for jobs**: a streamed reply cannot be replayed, a non-streamed one can.
- **Send images as `data:` URLs** when you do not know the route's model can fetch a URL, and keep them under the 20-image limit.
- **Never ship a secret key to a browser**: call AI chat from your server.

## Related

- [Authentication](/en/guides/authentication/), [Errors](/en/guides/errors/), [Limits](/en/guides/limits/)
- [Decision engine](/en/products/decision/): structured verdicts on the same key
- [Knowledge and vector search](/en/products/knowledge/): answers from your documents, with citations
- [AI embeddings](/en/products/ai.embeddings/)

## Endpoints

- POST /v1/chat/completions
- GET /v1/models
- POST /v1/run

## Example

```sh
curl -X POST "https://ai.inovacc.dev/v1/chat/completions" -H "Authorization: Bearer $INOVACC_API_KEY" -H "X-Operation-Id: $(uuidgen)" -H "Content-Type: application/json" -d '{}'
```
