Limits
Every Inovacc API bounds the size of what you send, how long a call may take and how often you may call. This guide lists those limits per API and says what happens when you reach each one, so your code can stay under them or handle the refusal. Values marked "per organization" are set for your organization by Inovacc; the others are fixed.
How limits answer
| You reach | Answer | What to do |
|---|---|---|
| A size limit (body, file, record, item) | 413 (request_too_large, payload_too_large, url_too_large), or 400 for a field or document that is too large | Send less, or use the upload in parts |
| A count limit (documents per call, operations per batch, items per collection) | 400 invalid_body, 400 batch_too_large or 409 collection_full | Split the call; start a new collection |
| A rate limit | 429 rate_limited, with Retry-After on the AI API | Wait, then retry |
| A token, cost or concurrency allowance (AI) | 429 quota_exceeded or 429 concurrency_limited, no Retry-After | Slow down for minutes, or raise the allowance |
| A daily budget (AI) | 429 budget_exhausted, Retry-After until 00:00 UTC | Wait for the next UTC day |
| A per-request input cap (AI) | 400 input_too_large | Shorten the input |
A call refused at a limit is refused before its work starts: nothing is called on your behalf and nothing is written.
AI
| Limit | Value |
|---|---|
| Request body (chat, embeddings, run) | 1 MiB by default; per organization, 1 KiB to 20 MiB |
| Input tokens per request | per organization (400 input_too_large) |
| Output tokens per request | the lower of the route's and your organization's ceiling; a higher max_tokens is lowered, not refused |
| Time until the model starts answering | 60 seconds (504 timeout_error); not on /v1/run |
| Images in one chat request | 20 |
| Estimated input cost of one image | 1,600 tokens |
| Stored answer for an idempotent replay | 24 hours |
| A call held as in progress | 5 minutes |
| Vector Database: documents per ingest, ids per delete | 100 |
Vector Database: one document (id + text + metadata) | 10 KiB |
| Vector Database: request body | 1 MiB |
Vector query: text; top_k | 8,000 characters; 50 |
| Collection item in a request body, or by URL | 16 MiB |
| Collection | 100 items; 1 GiB of files sent in requests; 20 GiB of video by upload |
| Upload part | 64 MiB, every part but the last |
| Upload that is not a video (transcript, notes) | 16 MiB |
| Recording upload | per organization |
| Knowledge job result | kept 24 hours |
| Activity and usage reads | up to 366 days and 92 days per call |
Quotas per organization. Each of these is set for your organization, and unlimited when not set:
| Allowance | Window | When reached |
|---|---|---|
| Requests | per UTC clock minute, per UTC day | 429 rate_limited, Retry-After to the next minute or day |
| Tokens | per UTC day, per UTC month | 429 quota_exceeded |
| Cost | per UTC day, per UTC month | 429 quota_exceeded |
| Concurrent calls | at any moment | 429 concurrency_limited |
| Daily budget, per provider and optionally per application | per UTC day | 429 budget_exhausted, Retry-After to 00:00 UTC |
A request counts against the per-minute and per-day request limits as soon as it is admitted, even if it fails later. Token and cost limits allow a call that reaches the limit exactly. From a budget's warning level (80% by default), answers carry X-Budget-Warning before calls start being refused.
Data and Files
| Limit | Value |
|---|---|
| JSON request body | 1 MiB (413 payload_too_large) |
| One record | 900 KiB (400 record_too_large) |
A json field | 64 KiB |
| Fields per collection; collections per database | 64; 64 |
| Records per page | 200 (default 50) |
| Operations per batch | 100, one transaction |
| Filter | 2,048 characters, 40 values, nesting depth 8 |
| Query string | 16 KiB (400 invalid_query) |
| File or blob in one request | 100 MiB, with Content-Length (411 without it) |
| File or blob in parts | 5 GiB; parts of 5 MiB to 100 MiB, numbered 1 to 10,000 |
| File path; listing prefix; file metadata | 1,024 bytes; 256 bytes; 2 KiB |
| Realtime listeners per collection; one event | 1,000; 512 KiB |
| Activity log page | 1 to 500 (default 100) |
| Rate | 600 requests per minute per credential holder, per location (429 rate_limited) |
The rate limit is counted per location and is a brake, not an exact count.
Events
| Limit | Value |
|---|---|
| Request body | 5 MiB (413 payload_too_large) |
| Events per publish | 100 |
data per event | 16 KiB (basic, reliable), 32 KiB (advanced) |
Deduplication window of an idempotency_key | 168 hours |
| Subscriptions per workspace; patterns per subscription | 50; 20 |
| Messages per pull; pull lease | 100; 5 to 300 seconds |
| Webhook answer time | 10 seconds |
| Webhook retries | from 10 seconds, doubling to 10 minutes; cap 5 |
| Dead letters kept | 30 days |
| Rate | 600 calls per minute per credential, per location (429 rate_limited) |
Staying under the limits
- Batch: send up to 100 documents, operations or events per call instead of one each.
- Upload large files in parts instead of raising body sizes.
- Back off on 429: honour
Retry-After, and add jitter so many clients do not retry at the same instant. - Set
max_tokenson chat calls: it lowers what each call reserves against your quotas. - Watch
X-Budget-Warningand the AI usage summary to see an allowance coming before it is spent.
See also
- Errors: every code and the retry policy.
- Authentication: credentials and idempotency.
- Each product page lists its own limits in full.