# POST /v1/extract

> Return the structured data a page already declares — JSON-LD, OpenGraph, and meta tags.

Source: https://docs.buildonto.dev/api/extract
Section: Read API

---

POST /v1/extract\[01\]

## POST /v1/extract

Return the structured data a page already declares — JSON-LD, OpenGraph/Twitter cards, and meta tags. 1 credit per call.

**Deterministic, no AI.** Onto parses the structured data the page itself publishes — it never invents fields. For free-form questions over a page, use [/v1/read](/api/read) and let your own model reason over the clean Markdown.

### Endpoint

\[02\]


```
POST https://api.buildonto.dev/v1/extract
Authorization: Bearer onto_sk_live_YOUR_KEY
Content-Type: application/json
```

### Request body

\[03\]

urlstringreq

The public URL to extract structured data from.

freshboolean

If true, skip the cache lookup and fetch again. The result is still written to the cache. A cache hit still costs 1 credit. Default: false.

**PDFs and text URLs are accepted, but this endpoint never returns page text.** For the text of a PDF use [/v1/read](/api/read) (or [/v1/batch](/api/batch) with `mode: "read"`). PDFs declare no HTML-level structured data, so `structured` comes back empty, the `counts` are all zero, and `aio_score`, `grade` and `hallucination_risk` are all `null` — a PDF effectively yields the title only. Text URLs (JSON, CSV, XML, plain text) behave the same way. Image-only PDFs return `IMAGE_PDF` (422). Failed requests are refunded.

### Response

\[04\]

**Success (200):** the parsed `structured` data (arrays/maps; empty when the page declares none), `counts` per type, and the AIO trust score for the page. `structured.openGraph` is a single merged map of both `og:*` and `twitter:*` card properties, and `counts.open_graph` counts both. To extract across a whole site in one call, use [/v1/batch](/api/batch) with `mode: "extract"`.


```
{
  "status": "success",
  "url": "https://vercel.com",
  "title": "Vercel: Build and deploy…",
  "aio_score": 90,
  "grade": "Excellent",
  "hallucination_risk": "low",
  "structured": {
    "jsonLd": [
      { "@context": "https://schema.org", "@type": "SoftwareApplication", "name": "Vercel" }
    ],
    "openGraph": {
      "og:title": "Vercel: Build and deploy the best web experiences…",
      "og:type": "website",
      "twitter:card": "summary_large_image"
    },
    "meta": {
      "description": "Vercel provides the developer tools…",
      "author": "Vercel"
    }
  },
  "counts": { "json_ld": 1, "open_graph": 13, "meta": 10 },
  "cache": { "hit": false, "ttl_seconds": 3600 }
}
```

**Errors:** common ones for this endpoint: `INVALID_URL` (400), `UNAUTHORIZED` (401), `ROBOTS_BLOCKED` (403), `WAF_BLOCKED` (403), `URL_NOT_FOUND` (404), `TOO_LARGE` (413), `UNSUPPORTED_TYPE` (415), `IMAGE_PDF` (422), `RATE_LIMITED` (429), `CONCURRENT_LIMIT` (429), `PAYMENT_REQUIRED` (402), `EXTRACTION_FAILED` (500), `TLS_ERROR` (502), `TIMEOUT` (504). See [error codes](/api/errors).

### Examples

\[05\]

cURL:


```
curl -X POST https://api.buildonto.dev/v1/extract \
  -H "Authorization: Bearer $ONTO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "url": "https://vercel.com" }'
```

Node (fetch):


```
const res = await fetch('https://api.buildonto.dev/v1/extract', {
  method: 'POST',
  headers: {
    'Authorization': `Bearer ${process.env.ONTO_API_KEY}`,
    'Content-Type': 'application/json',
  },
  body: JSON.stringify({ url: 'https://vercel.com' }),
});
const { structured } = await res.json();
console.log(structured.jsonLd, structured.openGraph['og:title']);
```

Copy as Markdown[](/api/extract.md "Open the raw Markdown")

---
## Structured Data (JSON-LD)
```json
{
  "@context": "https://schema.org",
  "@type": "WebSite",
  "name": "Onto Docs",
  "url": "https://docs.buildonto.dev",
  "description": "How to use Onto: serve AI agents Markdown from your Next.js site, call the Read API, connect over MCP, and read the AIO score.",
  "inLanguage": "en",
  "publisher": {
    "@type": "Organization",
    "name": "Onto",
    "url": "https://buildonto.dev"
  }
}
```