IMAGE_PDF
The URL is a scanned PDF with no text layer, so there's no text to extract.
What this means
[02]IMAGE_PDF fires when a URL resolves to a PDF whose pages are only images, with no embedded text. Onto doesn't run OCR, so there's nothing to return. It can happen on /v1/read, /v1/read-and-score, /v1/score, /v1/extract, and on single items in /v1/batch. The request is refunded.
When you'll see it
[03]HTTP 422. Body always includes code: "IMAGE_PDF". Branch on code, never on the human-readable message — wording can change without notice; the code is the stable contract.
Example response
[04]{
"status": "error",
"code": "IMAGE_PDF",
"message": "This PDF has no extractable text (likely scanned or image-only). Onto does not OCR."
}How to handle
[05]Run the PDF through OCR first so it has a text layer, or read the HTML version if there is one. If you make the PDFs, export them with selectable text instead of flattened images.
Suggested handling in a Node client:
if (data.code === 'IMAGE_PDF') {
// Image-only PDF — nothing to extract (and you were refunded). OCR upstream if needed.
return null;
}The URL returned something Onto can't read as a document, like an image, video or archive.
The fetch or the parse failed: an origin error, an unreachable host, a redirect problem, or a page or PDF Onto couldn't parse.
See the full error index for the complete catalog with the handling switch statement covering every code at once.