Inovacc Developer

API reference

AI embeddings

Turn text into vectors for search and similarity.

Base URL https://ai.inovacc.dev

Overview

AI embeddings turns text into vectors: lists of numbers whose distance from each other follows the meaning of the text. Two passages about the same thing land close together even when they share no words. The endpoint, POST /v1/embeddings, speaks the OpenAI embeddings wire format, so an existing OpenAI client can call it with the base URL https://ai.inovacc.dev/v1 and two extra headers.

The problem it solves is semantic comparison inside your own systems. With vectors you can search by meaning, group similar items, find near-duplicates or route a message to the closest category, using whatever database or index you already run.

Use AI embeddings when you keep the vectors. When you would rather have Inovacc keep them, search them and return the matching passages, use the managed Vector Database of Knowledge and vector search, which embeds your documents for you. When you need generated text rather than vectors, use AI chat.

Concepts

Routes. As with chat, you never name a model. Your organization is given embedding routes, opaque ids configured by Inovacc. GET /v1/models lists them: an entry whose endpoint is /v1/embeddings is an embedding route. Send its id as model, or send "auto" (or no model) to use your organization's default route. A route serves one endpoint only; a chat route called here is refused.

Input. input is one non-empty string, or an array of strings; inside an array an empty string is accepted. The answer holds one vector per input, in data[], each with the index of the input it belongs to.

Vectors of one route belong together. A vector is only comparable with vectors produced by the same route. Store the route id next to every vector you keep, and embed both your documents and your queries with the same route.

Optional fields. encoding_format and dimensions are passed to the model where the route's model supports them and ignored otherwise. user is accepted as a string and replaced, before it leaves Inovacc, by a value derived from your organization, so two organizations sending the same value never collide.

Usage. The answer's usage holds prompt_tokens and total_tokens. This endpoint carries no x-usage-* headers; the route that answered is in x-route-id.

How it works

Every call carries Authorization: Bearer <key> and an X-Operation-Id. Your organization comes from the key. The service checks the operation id and the key, then whether that operation id was already used, then the body: only model, input, encoding_format, dimensions and user are accepted, and any other field is refused by name. It resolves the route, checks that the route serves embeddings, validates input, estimates the input size (a quarter of the characters of every input string) against your organization's per-request cap, and reserves the call against your organization's quotas. Then the model is called.

The answer is {"object":"list","data":[{"object":"embedding","index":0,"embedding":[...]}],"model":"<route id>","usage":{...}}. model is always the route id.

There is no streaming. A successful answer is stored for 24 hours under its operation id: the same id sent again returns the same vectors with x-idempotent-replay: true, without a second call or charge. An error is never stored, so a failed call can be retried with the same id. What is metered is input tokens.

Get started

You need an API key and an embedding route enabled for your organization (Authentication).

  1. Find your embedding route. Call GET /v1/models and pick an entry whose endpoint is /v1/embeddings.
  2. Embed two passages. Call POST /v1/embeddings with that route as model, an input array of two strings and a fresh X-Operation-Id (see the samples). The answer has two entries in data, with index 0 and 1, and model is your route id.
  3. Compare them. Compute the cosine similarity of the two vectors in your code. Embed a third passage on another subject and compare again: its score against the first is lower.
  4. Replay. Send step 2 again with the same operation id; the vectors are identical and the answer carries x-idempotent-replay: true.

Use cases

Search in your own database. A product catalogue stores one vector per product description in the database it already uses. A shopper's question is embedded with the same route at query time, and the nearest products are shown, including those whose description uses different words.

Near-duplicate detection. A ticketing system embeds each new ticket and compares it with the open ones; a score above a threshold you calibrate links the ticket to the existing case instead of opening a second one.

Routing by similarity. A small set of example messages per team is embedded once. Each incoming message is embedded and sent to the team whose examples are closest, with no model call per message beyond the embedding.

Limits and pricing

LimitValue
Request body1 MiB by default; your organization may be set between 1 KiB and 20 MiB
Input tokens per requestyour organization's cap (400 input_too_large above it)
Time to the model's first response60 seconds
Stored answer for a replay24 hours
Requests, tokens, cost, concurrency, daily budgetsas set for your organization

Pricing: on request. The pricing unit is tokens.

Errors

StatusCodeWhat it means and what to do
400missing_operation_id, invalid_operation_idSend a valid X-Operation-Id.
400unsupported_fieldA field other than model, input, encoding_format, dimensions, user; remove it.
400invalid_bodyMalformed JSON, input of the wrong shape, or a non-string user.
400model_not_allowedmodel is not a valid route id.
400input_too_largeThe input is over your per-request cap; split it into several calls.
401missing_credentials, invalid_credentialsSend a valid key.
403route_forbiddenThe route is not yours, is disabled, or does not serve embeddings.
403key_disabled, organization_disabledCheck the key.
409operation_in_progressThat operation id is still running; wait and retry.
413request_too_largeThe body is over your size cap.
429rate_limited, budget_exhaustedWait the Retry-After seconds.
429quota_exceeded, concurrency_limitedAn allowance is used up; no Retry-After.
502, 503, 504upstream_error, service_unavailable, timeout_errorRetry with the same operation id and backoff.

See Errors for the envelope.

Best practices

  • Batch inputs: send many passages in one input array, within your size and token caps, rather than one call per passage.
  • Keep the route id with the vector, and re-embed everything if you change route: vectors of two routes are not comparable.
  • Chunk long documents into passages of a few paragraphs before embedding; one vector for a whole document blurs its meaning.
  • Derive the operation id from the content (for example a hash of the batch) so a restarted job replays instead of paying again.
  • Retry 502, 503 and 504 with the same operation id; wait Retry-After on 429 rate_limited.
  • Calibrate thresholds on your own data: similarity scores are relative, not probabilities.

Authentication

Every call carries your key; your organization comes from it. See the authentication guide.

HeaderAuthorizationBearer <API key>
HeaderX-Operation-Ida unique id you choose, on every call except GET

Endpoints

POST /v1/embeddings

Example

Language

⋮
POST /v1/embeddingsExample

cURL

curl -X POST "https://ai.inovacc.dev/v1/embeddings" \
  -H "Authorization: Bearer $INOVACC_API_KEY" \
  -H "X-Operation-Id: $(uuidgen)" \
  -H "Content-Type: application/json" \
  -d '{
  "input": [
    "first passage",
    "second passage"
  ],
  "model": "r-embed"
}'

TypeScript

import crypto from "node:crypto";

const body: Record<string, unknown> = {
  "input": [
    "first passage",
    "second passage"
  ],
  "model": "r-embed"
};

const url = "https://ai.inovacc.dev/v1/embeddings";

const response = await fetch(url, {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.INOVACC_API_KEY}`,
    "X-Operation-Id": crypto.randomUUID(),
    "Content-Type": "application/json",
  },
  body: JSON.stringify(body),
});

console.log(response.status, await response.text());

Python

import json
import os
import urllib.request
import uuid

body = {
    "input": [
        "first passage",
        "second passage",
    ],
    "model": "r-embed",
}

request = urllib.request.Request(
    "https://ai.inovacc.dev/v1/embeddings",
    data=json.dumps(body).encode(),
    method="POST",
    headers={
        "User-Agent": "inovacc-python-sample",
        "Authorization": "Bearer " + os.environ["INOVACC_API_KEY"],
        "X-Operation-Id": str(uuid.uuid4()),
        "Content-Type": "application/json",
    },
)

with urllib.request.urlopen(request) as response:
    print(response.status, response.read().decode())

Go

package main

import (
	"crypto/rand"
	"fmt"
	"io"
	"net/http"
	"os"
	"strings"
)

const url = "https://ai.inovacc.dev/v1/embeddings"

const body = `{
  "input": [
    "first passage",
    "second passage"
  ],
  "model": "r-embed"
}`

// newUUID returns a random (version 4) UUID, from the standard library alone.
func newUUID() string {
	b := make([]byte, 16)
	if _, err := rand.Read(b); err != nil {
		panic(err)
	}
	b[6] = b[6]&0x0f | 0x40
	b[8] = b[8]&0x3f | 0x80
	return fmt.Sprintf("%x-%x-%x-%x-%x", b[0:4], b[4:6], b[6:8], b[8:10], b[10:])
}

func main() {
	req, err := http.NewRequest("POST", url, strings.NewReader(body))
	if err != nil {
		panic(err)
	}
	req.Header.Set("Authorization", "Bearer "+os.Getenv("INOVACC_API_KEY"))
	req.Header.Set("X-Operation-Id", newUUID())
	req.Header.Set("Content-Type", "application/json")

	res, err := http.DefaultClient.Do(req)
	if err != nil {
		panic(err)
	}
	defer res.Body.Close()

	out, err := io.ReadAll(res.Body)
	if err != nil {
		panic(err)
	}
	fmt.Println(res.Status, string(out))
}

Rust

// Cargo.toml: reqwest = { version = "0.12", features = ["blocking", "json"] }
// Cargo.toml: serde_json = "1"
// Cargo.toml: uuid = { version = "1", features = ["v4"] }

fn main() -> Result<(), Box<dyn std::error::Error>> {
    let api_key = std::env::var("INOVACC_API_KEY")?;
    let body = serde_json::json!({
        "input": [
            "first passage",
            "second passage"
        ],
        "model": "r-embed"
    });
    let response = reqwest::blocking::Client::new()
        .post("https://ai.inovacc.dev/v1/embeddings")
        .bearer_auth(api_key)
        .header("X-Operation-Id", uuid::Uuid::new_v4().to_string())
        .json(&body)
        .send()?;
    println!("{} {}", response.status(), response.text()?);
    Ok(())
}

JavaScript

import crypto from "node:crypto";

const body = {
  "input": [
    "first passage",
    "second passage"
  ],
  "model": "r-embed"
};

const url = "https://ai.inovacc.dev/v1/embeddings";

const response = await fetch(url, {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.INOVACC_API_KEY}`,
    "X-Operation-Id": crypto.randomUUID(),
    "Content-Type": "application/json",
  },
  body: JSON.stringify(body),
});

console.log(response.status, await response.text());

PHP

<?php

$body = <<<'JSON'
{
  "input": [
    "first passage",
    "second passage"
  ],
  "model": "r-embed"
}
JSON;

function uuid4(): string
{
    $bytes = random_bytes(16);
    $bytes[6] = chr(ord($bytes[6]) & 0x0f | 0x40);
    $bytes[8] = chr(ord($bytes[8]) & 0x3f | 0x80);
    return vsprintf('%s%s-%s-%s-%s-%s%s%s', str_split(bin2hex($bytes), 4));
}

$curl = curl_init('https://ai.inovacc.dev/v1/embeddings');
curl_setopt_array($curl, [
    CURLOPT_CUSTOMREQUEST => 'POST',
    CURLOPT_RETURNTRANSFER => true,
    CURLOPT_HTTPHEADER => [
        'Authorization: Bearer ' . getenv('INOVACC_API_KEY'),
        'X-Operation-Id: ' . uuid4(),
        'Content-Type: application/json',
    ],
    CURLOPT_POSTFIELDS => $body,
]);

$response = curl_exec($curl);
echo curl_getinfo($curl, CURLINFO_HTTP_CODE), ' ', $response, "\n";
curl_close($curl);

Ruby

require "net/http"
require "securerandom"
require "uri"

uri = URI('https://ai.inovacc.dev/v1/embeddings')
body = <<~'JSON'
{
  "input": [
    "first passage",
    "second passage"
  ],
  "model": "r-embed"
}
JSON

request = Net::HTTP::Post.new(uri)
request["Authorization"] = "Bearer #{ENV.fetch('INOVACC_API_KEY')}"
request["X-Operation-Id"] = SecureRandom.uuid
request["Content-Type"] = "application/json"
request.body = body

response = Net::HTTP.start(uri.host, uri.port, use_ssl: uri.scheme == "https") do |http|
  http.request(request)
end

puts "#{response.code} #{response.body}"

Java

import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
import java.util.UUID;

public class Main {
    public static void main(String[] args) throws Exception {
        String apiKey = System.getenv("INOVACC_API_KEY");
        String body = "{" +
                "\"input\": [" +
                "\"first passage\"," +
                "\"second passage\"" +
                "]," +
                "\"model\": \"r-embed\"" +
                "}";

        HttpRequest request = HttpRequest.newBuilder()
                .uri(URI.create("https://ai.inovacc.dev/v1/embeddings"))
                .header("Authorization", "Bearer " + apiKey)
                .header("X-Operation-Id", UUID.randomUUID().toString())
                .header("Content-Type", "application/json")
                .method("POST", HttpRequest.BodyPublishers.ofString(body))
                .build();

        HttpResponse<String> response = HttpClient.newHttpClient()
                .send(request, HttpResponse.BodyHandlers.ofString());
        System.out.println(response.statusCode() + " " + response.body());
    }
}

C#

using System.Text;

var body = "{" +
    "\"input\": [" +
    "\"first passage\"," +
    "\"second passage\"" +
    "]," +
    "\"model\": \"r-embed\"" +
    "}";

var url = "https://ai.inovacc.dev/v1/embeddings";
var apiKey = Environment.GetEnvironmentVariable("INOVACC_API_KEY");

using var client = new HttpClient();
using var request = new HttpRequestMessage(HttpMethod.Post, url);
request.Headers.Add("Authorization", $"Bearer {apiKey}");
request.Headers.Add("X-Operation-Id", Guid.NewGuid().ToString());
request.Content = new StringContent(body, Encoding.UTF8, "application/json");

using var response = await client.SendAsync(request);
Console.WriteLine($"{(int)response.StatusCode} {await response.Content.ReadAsStringAsync()}");

Kotlin

import java.net.URI
import java.net.http.HttpClient
import java.net.http.HttpRequest
import java.net.http.HttpResponse
import java.util.UUID

fun main() {
    val apiKey = System.getenv("INOVACC_API_KEY")
    val body = "{" +
        "\"input\": [" +
        "\"first passage\"," +
        "\"second passage\"" +
        "]," +
        "\"model\": \"r-embed\"" +
        "}"

    val request = HttpRequest.newBuilder()
        .uri(URI.create("https://ai.inovacc.dev/v1/embeddings"))
        .header("Authorization", "Bearer " + apiKey)
        .header("X-Operation-Id", UUID.randomUUID().toString())
        .header("Content-Type", "application/json")
        .method("POST", HttpRequest.BodyPublishers.ofString(body))
        .build()

    val response = HttpClient.newHttpClient()
        .send(request, HttpResponse.BodyHandlers.ofString())
    println("${response.statusCode()} ${response.body()}")
}

Swift

import Foundation
#if canImport(FoundationNetworking)
import FoundationNetworking
#endif

let body = #"""
{
  "input": [
    "first passage",
    "second passage"
  ],
  "model": "r-embed"
}
"""#

let apiKey = ProcessInfo.processInfo.environment["INOVACC_API_KEY"] ?? ""

let url = URL(string: "https://ai.inovacc.dev/v1/embeddings")!
var request = URLRequest(url: url)
request.httpMethod = "POST"
request.setValue("Bearer \(apiKey)", forHTTPHeaderField: "Authorization")
request.setValue(UUID().uuidString, forHTTPHeaderField: "X-Operation-Id")
request.setValue("application/json", forHTTPHeaderField: "Content-Type")
request.httpBody = Data(body.utf8)

let (data, response) = try await URLSession.shared.data(for: request)
let status = (response as? HTTPURLResponse)?.statusCode ?? 0
print(status, String(decoding: data, as: UTF8.self))
Source details

Catalogue entry

Id
identity/component-taxonomy/capabilities#ai.embeddings
Repository
identity
Path
contracts/component-taxonomy/catalog/capabilities.json
Commit
08ef3cd75d82

Example

Id
identity/component-taxonomy/quickstart#ai.embeddings
Repository
identity
Path
contracts/component-taxonomy/catalog/capabilities.json
Commit
08ef3cd75d82
Updated 2026-10-10.