# POST /v1/read

> Fetch any public URL and return clean Markdown plus extraction stats.

Source: https://docs.buildonto.dev/api/read
Section: Read API

---

POST /v1/read\[01\]

## POST /v1/read

Fetch any public URL and return clean Markdown plus extraction stats. 1 credit per call. No score — use [/v1/read-and-score](/api/read-and-score) for that.

Do this firstTerminalRunCopy

curl -X POST https://api.buildonto.dev/v1/read \\
  -H "Authorization: Bearer onto\_sk\_live\_YOUR\_KEY" \\  -d '{"url":"https://stripe.com/pricing"}'

Returns clean Markdown plus extraction stats.

**Cached for 1 hour** per (URL, engine version). Every deploy changes the engine version, so a deploy empties the cache. Set `"fresh": true` to skip the cache lookup; the new result is still written to the cache. A cache hit still costs 1 credit.

### Endpoint

\[02\]


```
POST https://api.buildonto.dev/v1/read
Authorization: Bearer onto_sk_live_YOUR_KEY
Content-Type: application/json
```

### Request body

\[03\]

urlstringreq

The public URL to fetch and clean. Must be http:// or https://. Private or internal hosts return INVALID\_URL.

freshboolean

If true, skip the cache lookup and fetch again. The result is still written to the cache. Default: false.

**PDFs and text formats work too.** Point this endpoint at a PDF and Onto extracts its text — no OCR, no vision. JSON, CSV, XML and plain text are read as well. Image-only PDFs (no embedded text layer) return `IMAGE_PDF` (422). Failed requests are refunded.

### Response

\[04\]

**Success (200):** JSON by default. Add `Accept: text/markdown` to get raw Markdown back instead. The `url` in the response is canonicalised — for example, GitHub blob URLs become `raw.githubusercontent.com` URLs. A `warnings` array appears only when something is worth flagging, such as a page that looks like a JavaScript shell with little server-rendered content.


```
{
  "status": "success",
  "url": "https://stripe.com",
  "markdown": "# Stripe — Online payments…\n\n…",
  "metadata": {
    "title": "Stripe | Financial Infrastructure...",
    "description": "Stripe powers online and...",
    "language": "en"
  },
  "stats": {
    "raw_html_size_kb": 605.0,
    "markdown_size_kb": 14.8,
    "reduction_percent": 97.6,
    "extraction_time_ms": 412
  },
  "cache": { "hit": false, "ttl_seconds": 3600 }
}
```

**Errors:** see [error codes](/api/errors). This endpoint can return `INVALID_URL` (400, also private or internal hosts and non-http schemes), `UNAUTHORIZED` (401), `PAYMENT_REQUIRED` (402), `ROBOTS_BLOCKED` (403), `WAF_BLOCKED` (403), `URL_NOT_FOUND` (404, also DNS failure), `TOO_LARGE` (413, over 10 MB), `UNSUPPORTED_TYPE` (415), `IMAGE_PDF` (422), `RATE_LIMITED` (429), `CONCURRENT_LIMIT` (429), `EXTRACTION_FAILED` (500), `TLS_ERROR` (502), `TIMEOUT` (504, 15 s page fetch).

### Response headers

\[05\]

X-Onto-Cache'HIT' | 'MISS'

Whether this came from the one-hour cache. A HIT means no outbound fetch happened. It still costs 1 credit.

X-RateLimit-Remainingint

Monthly plan credits left. Stays 0 while you are billed from bought credits. Not a short rate-limit window.

X-RateLimit-ResetISO date

When monthly plan credits reset (end of the UTC month).

X-Onto-Billed'plan' | 'credit'

Whether this request used monthly plan credits or bought credits.

X-Credits-Remainingint

Bought-credit balance after this request. Sent only when the request was billed to credits.

X-Concurrent-Remainingint

How many more in-flight requests you have headroom for at this moment.

### Examples

\[06\]

cURL — JSON response:


```
curl -X POST https://api.buildonto.dev/v1/read \
  -H "Authorization: Bearer $ONTO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://stripe.com"}'
```

cURL — Markdown response (no JSON wrapper):


```
curl -X POST https://api.buildonto.dev/v1/read \
  -H "Authorization: Bearer $ONTO_API_KEY" \
  -H "Accept: text/markdown" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://stripe.com"}'
```

Node (fetch):


```
const res = await fetch('https://api.buildonto.dev/v1/read', {
  method: 'POST',
  headers: {
    'Authorization': `Bearer ${process.env.ONTO_API_KEY}`,
    'Content-Type': 'application/json',
  },
  body: JSON.stringify({ url: 'https://stripe.com' }),
});
const data = await res.json();
console.log(data.markdown);
```

Python (httpx):


```
import os, httpx

r = httpx.post(
    "https://api.buildonto.dev/v1/read",
    headers={"Authorization": f"Bearer {os.environ['ONTO_API_KEY']}"},
    json={"url": "https://stripe.com"},
)
print(r.json()["markdown"])
```

Copy as Markdown[](/api/read.md "Open the raw Markdown")

---
## Structured Data (JSON-LD)
```json
{
  "@context": "https://schema.org",
  "@type": "WebSite",
  "name": "Onto Docs",
  "url": "https://docs.buildonto.dev",
  "description": "How to use Onto: serve AI agents Markdown from your Next.js site, call the Read API, connect over MCP, and read the AIO score.",
  "inLanguage": "en",
  "publisher": {
    "@type": "Organization",
    "name": "Onto",
    "url": "https://buildonto.dev"
  }
}
```