# Rate limits + credits

> Three layers — monthly quota, concurrent cap, and credit overflow.

Source: https://docs.buildonto.dev/api/limits
Section: Read API

---

Rate limits + credits\[01\]

## Rate limits + credits

Three layers — monthly plan credits, concurrent cap, and bought credits.

### Monthly quota

\[02\]

Each plan has a fixed number of credits per UTC month. Quota is per account, not per key. It resets at 00:00 UTC on the 1st of each month. Check current usage anytime via [GET /v1/usage](/api/usage).

Plan · Credits / month

Free\[1,000\]

Starter\[10,000\]

Growth\[100,000\]

Scale\[500,000\]

Enterprise\[Unlimited\]

What each call costs:

Endpoint · Credits

/v1/read\[1\]

/v1/read-and-score\[1\]

/v1/extract\[1\]

/v1/map

1 per call

/v1/batch

1 per URL; every ok: false item is refunded

/v1/score\[0\]

/v1/usage

0, and it takes no concurrency slot

Credits are taken before the cache lookup, so a cache hit still costs credits. A request that fails is refunded.

When you run out, behaviour depends on the tier. `Free` is **hard-capped**: it never spends credits and returns `RATE_LIMITED 429`, never 402. The body's `retry_after` is an ISO date for the end of the UTC month. Paid tiers spend bought credits; when those are empty they get `PAYMENT_REQUIRED 402`.

The check for a batch is all-or-nothing. If a batch of N URLs needs more than you have left, all N go to bought credits on a paid plan, and Free gets `RATE_LIMITED`.

`X-RateLimit-Remaining` is the monthly plan credits left, and `X-RateLimit-Reset` is the end of the UTC month. They describe the monthly quota, not a short rate-limit window. `X-RateLimit-Remaining` stays 0 while you are billed from bought credits.

### Concurrent requests

\[03\]

Independent of the monthly counter — limits how many requests can be **in flight at the same instant** per account.

Plan · Concurrent

Free\[2\]

Starter\[5\]

Growth\[20\]

Scale\[50\]

Enterprise\[100\]

Exceed it and you get `CONCURRENT_LIMIT 429`. The body includes `in_flight`, `limit` and `retry_after: 1` (always 1). A slot is held for the whole request, and one request can hold it for a while: robots.txt check up to 5 s, page fetch up to 15 s, probe up to 8 s. Implementation: atomic Redis INCR with 30s safety TTL — a crashed function never permanently consumes a slot.

A batch holds one slot, however many URLs it has. `/v1/score` takes a slot even though it is free. `/v1/usage` does not.

**Concurrency is per account, not per API key or per IP.** Every process using your key, and the remote MCP connector, draw from the same pool. Upgrade the tier for more slots.

### Credits (overflow)

\[04\]

Paid tiers run as **soft caps**. Once monthly plan credits are used up, each call debits its credit cost from your bought credits. Out of both → 402 `PAYMENT_REQUIRED`. Buy credit packs from `/read/billing#credits`:

Pack · Price · Credits · Per credit · Bonus

Top-up S\[$5\]\[500\]\[$0.0100\]\[—\]

Top-up M\[$20\]\[2,200\]\[$0.0091\]\[10% bonus\]

Top-up L\[$50\]\[6,000\]\[$0.0083\]\[20% bonus\]

Top-up XL\[$200\]\[28,000\]\[$0.0071\]\[40% bonus\]

**Free never spends credits.** Only paid tiers can buy top-ups, and only paid tiers spend them. A Free account that runs out gets `RATE_LIMITED 429` until the month resets or it upgrades.

Per-request response headers tell you which bucket got billed: `X-Onto-Billed: plan` (in quota) or `X-Onto-Billed: credit` (bought credits). `X-Credits-Remaining` is sent only when the request was billed to credits, and shows the new balance.

### Cache behavior

\[05\]

`/v1/read`, `/v1/read-and-score`, `/v1/extract`, `/v1/map` and `/v1/batch` are cached for **1 hour**. The `X-Onto-Cache` header says HIT or MISS. Cache hits **still cost credits**. `/v1/score` is never cached.

On read, read-and-score and extract, set `"fresh": true` to skip the cache lookup; the new result is still written to the cache. Map and batch have no `fresh` option.

The cache key includes an engine version that changes on every deploy, so every deploy empties the cache.

Copy as Markdown[](/api/limits.md "Open the raw Markdown")

---
## Structured Data (JSON-LD)
```json
{
  "@context": "https://schema.org",
  "@type": "WebSite",
  "name": "Onto Docs",
  "url": "https://docs.buildonto.dev",
  "description": "How to use Onto: serve AI agents Markdown from your Next.js site, call the Read API, connect over MCP, and read the AIO score.",
  "inLanguage": "en",
  "publisher": {
    "@type": "Organization",
    "name": "Onto",
    "url": "https://buildonto.dev"
  }
}
```