# Decision engine

Structured, criteria-based decisions at a speed and precision tier you choose.

Base URL: `https://ai.inovacc.dev`

## Overview

The Decision engine answers structured questions about a piece of content with structured verdicts. You send a **state** (the thing to judge: a text, a record, a JSON object) and a set of **questions**, each of a known type; you get back one answer per question, with the label or score chosen, the probabilities behind it and a confidence. The endpoint is `POST /v1/run` on `https://ai.inovacc.dev`, called on a decision route of your organization.

The problem it solves is turning a judgement into data a program can act on. Asking a chat model "is this good?" returns prose you then have to parse and cannot compare from one call to the next. The Decision engine returns the same shape every time, so you can store it, threshold it and audit it.

Use it for classification, scoring, triage and checks against criteria: does this ticket need a person, which of five categories fits, how well does this answer meet the bar. When you need generated prose, use [AI chat](/en/products/ai.chat/); when the verdict must rest on your own documents, retrieve the passages first with [Knowledge and vector search](/en/products/knowledge/) and put them in the state.

## Concepts

**Decision routes.** As everywhere in Inovacc AI, you call a route, never a model. `GET /v1/models` lists your routes; an entry whose `endpoint` is `/v1/run` is a run route, and the ones Inovacc configured as decision routes answer in the shape below.

**Questions and their types.** `questions` is an object of 1 to 64 entries, keyed by your own ids (letters, digits, `_`, `.` or `-`, up to 100 characters). Each question has a `type`, usually `instructions`, and `criteria` whose shape the type fixes:

| Type | Asks for | `criteria` |
|---|---|---|
| `noul` | yes or no | an object `{"true": "...", "false": "..."}` |
| `choice` | one label among several | an object `{"label": "description", ...}` |
| `score` | a level on a scale | an array of levels, lowest first |

The shapes are strict: a criteria string such as `"1-5"` or `"yes|no"` is refused.

**Answers.** Each answer carries what its type produced: the `choice` or the `score` with `probabilities`, a `legend` and a `confidence`. A `noul` answer carries its probability `p` instead of a confidence.

**Tiers.** A decision route serves one of three tiers: `fast`, `standard` or `precise`. All answer in one shape, and the answer's `engine.tier` says which one answered. **Confidence is not comparable between tiers**: each tier reports on its own scale, so set a threshold per tier.

**Escalation.** With `escalate`, questions answered below a confidence threshold are asked again, alone, on the next stronger tier (`fast`, then `standard`, then `precise`), and the stronger answer replaces the weaker one. The defaults are `fast` 0.40, `standard` 0.80 and `precise` 0.55; you can pass `min_confidence` (one number, or one per tier) and `max_tier`. A `noul` answer's confidence is the distance of `p` from a coin flip, `|2p - 1|`.

## How it works

Every call carries `Authorization: Bearer <key>` and an `X-Operation-Id`. Your organization comes from the key. The body takes only `model`, `input` and `escalate`. On a decision route `input` is `{"state", "questions", "images"?}`: `state` is any non-null JSON value, `images` at most 16 entries. Any other field inside `input` is refused with `invalid_decision_input`, and the reply never echoes it.

The answer is:

`{"model": "<route id>", "result": {"answers": {...}, "usage": {"input_tokens", "output_tokens"}}, "state": "Completed", "engine": {"tier", "latency_ms"}}`

and its headers carry `x-route-id`, `x-usage-input-tokens` and `x-usage-output-tokens`.

With escalation, each answer also says which `tier` answered it and, when it was escalated, `escalated_from` lists the lower tiers and the confidence each gave. An `escalation` object reports every attempt, how many questions were asked and how many came back below threshold, and how many answers each tier gave in the end. Each attempt is its own call, reserved against your quotas and metered at its own tier. If an attempt fails, the escalation stops, the earlier answers are kept, and the request is still `200`.

There is no streaming and no 60-second timeout on this endpoint. A successful answer is stored for 24 hours under its operation id: the same id returns the same merged answer with `x-idempotent-replay: true` and calls nothing. What is metered is tokens, at the price of each tier that answered.

## Get started

You need an API key and a decision route enabled for your organization ([Authentication](/en/guides/authentication/)).

1. **Find your decision route** with `GET /v1/models`: an entry whose `endpoint` is `/v1/run`.
2. **Ask one `noul` question.** `POST /v1/run` with your route as `model`, a short text as `state` and one question `{"type":"noul","instructions":"...","criteria":{"true":"...","false":"..."}}` (see [the samples](#example)). The answer has your question id under `result.answers`, with its probability `p`.
3. **Add a `score` question** with three levels to the same call. Both answers come back in one response; `engine.tier` names the tier.
4. **Escalate.** Send the call again with a new operation id and `"escalate": true`. The answer now carries `tier` per answer and an `escalation` object listing the attempts.

## Use cases

**Ticket triage.** Each incoming support ticket is the `state`; a `choice` question picks the team, a `noul` question asks whether a person must answer, a `score` question rates urgency. The answers are stored with the ticket, and a ticket whose confidence is below your threshold goes to a person.

**Quality checks on generated text.** Before an assistant's draft is sent, a decision call scores it against your bar ("Below the bar", "At the bar", "Above the bar") and checks a few `noul` rules. With `escalate`, only the drafts the fast tier was unsure about pay for the precise tier.

**Content moderation with images.** A listing's text and up to 16 images are judged against your policy questions; the per-question confidence decides between publishing, rejecting and sending to review.

## Limits and pricing

| Limit | Value |
|---|---|
| Questions per call | 1 to 64 |
| Question id | 1 to 100 characters: letters, digits, `_`, `.`, `-` |
| Images per call | 16 |
| Request body | 1 MiB by default; your organization may be set between 1 KiB and 20 MiB |
| Input tokens per request | your organization's cap |
| Stored answer for a replay | 24 hours |
| Requests, tokens, cost, concurrency, daily budgets | as set for your organization |

**Pricing:** on request. The pricing unit is **tokens**, at the price of the tier that answered.

## Errors

| Status | Code | What it means and what to do |
|---|---|---|
| 400 | `invalid_input` | `input` is missing or not a JSON object. |
| 400 | `invalid_decision_input` | `input` is not `{state, questions, images?}` or a question is malformed; check types and ids. |
| 400 | `invalid_escalate` | `escalate` has a key or value it does not take, or the route is not a decision route. |
| 400 | `unsupported_field` | A top-level field other than `model`, `input`, `escalate`. |
| 400 | `model_not_allowed`, `input_too_large` | Use a valid route id; shorten the state. |
| 400 | `missing_operation_id`, `invalid_operation_id` | Send a valid `X-Operation-Id`. |
| 401 | `missing_credentials`, `invalid_credentials` | Send a valid key. |
| 403 | `route_forbidden` | The route is not yours, or is not a run route. |
| 409 | `operation_in_progress` | That operation id is still running. |
| 413 | `request_too_large` | The body is over your size cap. |
| 429 | `rate_limited`, `quota_exceeded`, `concurrency_limited`, `budget_exhausted` | See [Errors](/en/guides/errors/); wait `Retry-After` where given. |
| 502 | `upstream_error` | The tier failed, or a criteria shape was refused; check the criteria, then retry. |
| 503 | `service_unavailable` | Decisions cannot be reached; retry later. |

## Best practices

- **Write criteria in the strict shapes**: an object for `noul` and `choice`, an array for `score`.
- **Keep thresholds per tier**, never one number for all tiers.
- **Use `escalate` with `max_tier`** to cap the cost: most questions stop at the cheaper tier.
- **Put everything the judgement needs in `state`**, including retrieved passages; the engine sees nothing else.
- **Use stable question ids**, so answers can be compared across calls and stored by id.
- **Derive the operation id from the item judged** so a retried batch replays instead of paying twice.

## Related

- [AI chat](/en/products/ai.chat/), [Knowledge and vector search](/en/products/knowledge/)
- [Authentication](/en/guides/authentication/), [Errors](/en/guides/errors/), [Limits](/en/guides/limits/)

## Endpoints

- POST /v1/run

## Example

```sh
curl -X POST "https://ai.inovacc.dev/v1/run" -H "Authorization: Bearer $INOVACC_API_KEY" -H "X-Operation-Id: $(uuidgen)" -H "Content-Type: application/json" -d '{}'
```
