Rate limits + credits
Three layers — monthly plan credits, concurrent cap, and bought credits.
Monthly quota
[02]Each plan has a fixed number of credits per UTC month. Quota is per account, not per key. It resets at 00:00 UTC on the 1st of each month. Check current usage anytime via GET /v1/usage.
What each call costs:
1 per call
1 per URL; every ok: false item is refunded
0, and it takes no concurrency slot
Credits are taken before the cache lookup, so a cache hit still costs credits. A request that fails is refunded.
When you run out, behaviour depends on the tier. Free is hard-capped: it never spends credits and returns RATE_LIMITED 429, never 402. The body's retry_after is an ISO date for the end of the UTC month. Paid tiers spend bought credits; when those are empty they get PAYMENT_REQUIRED 402.
The check for a batch is all-or-nothing. If a batch of N URLs needs more than you have left, all N go to bought credits on a paid plan, and Free gets RATE_LIMITED.
X-RateLimit-Remaining is the monthly plan credits left, and X-RateLimit-Reset is the end of the UTC month. They describe the monthly quota, not a short rate-limit window. X-RateLimit-Remaining stays 0 while you are billed from bought credits.
Concurrent requests
[03]Independent of the monthly counter — limits how many requests can be in flight at the same instant per account.
Exceed it and you get CONCURRENT_LIMIT 429. The body includes in_flight, limit and retry_after: 1 (always 1). A slot is held for the whole request, and one request can hold it for a while: robots.txt check up to 5 s, page fetch up to 15 s, probe up to 8 s. Implementation: atomic Redis INCR with 30s safety TTL — a crashed function never permanently consumes a slot.
A batch holds one slot, however many URLs it has. /v1/score takes a slot even though it is free. /v1/usage does not.
Credits (overflow)
[04]Paid tiers run as soft caps. Once monthly plan credits are used up, each call debits its credit cost from your bought credits. Out of both → 402 PAYMENT_REQUIRED. Buy credit packs from /read/billing#credits:
Per-request response headers tell you which bucket got billed: X-Onto-Billed: plan (in quota) or X-Onto-Billed: credit (bought credits). X-Credits-Remaining is sent only when the request was billed to credits, and shows the new balance.
Cache behavior
[05]/v1/read, /v1/read-and-score, /v1/extract, /v1/map and /v1/batch are cached for 1 hour. The X-Onto-Cache header says HIT or MISS. Cache hits still cost credits. /v1/score is never cached.
On read, read-and-score and extract, set "fresh": true to skip the cache lookup; the new result is still written to the cache. Map and batch have no fresh option.
The cache key includes an engine version that changes on every deploy, so every deploy empties the cache.