# Documents to text (OCR)


# Documents to text (OCR)

`POST /v1/ocr` reads a PDF or an image and gives back its text, page by page,
    in every EU language: å, ä, ö, ß, ł, ő and Greek or Bulgarian letters included.
    Pages of a digital PDF come straight from the text in the file; scanned pages and photos are read.
    Price: €0.003 / page, every page of the document.
    On a Managed plan a page counts as 1,000 of the plan's tokens.
    The document is read in memory and not stored.

## Send a document

Upload the file as a form field. Add `lang` with the document's language for the best result on scans.

```
curl -sS https://api.axforge.ai/v1/ocr \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -F file=@page-1.jpg \
  -F lang=en
```

```
# curl.exe (not PowerShell's curl alias) sends the file the same in 5.1 and 7
curl.exe -sS https://api.axforge.ai/v1/ocr `
  -H "Authorization: Bearer $env:AXFORGE_API_KEY" `
  -F "file=@page-1.jpg" `
  -F "lang=en"
```

```
curl.exe -sS https://api.axforge.ai/v1/ocr ^
  -H "Authorization: Bearer %AXFORGE_API_KEY%" ^
  -F "file=@page-1.jpg" ^
  -F "lang=en"
```

```
import os
import requests

with open("page-1.jpg", "rb") as f:
    r = requests.post(
        "https://api.axforge.ai/v1/ocr",
        headers={"Authorization": f"Bearer {os.environ['AXFORGE_API_KEY']}"},
        files={"file": ("page-1.jpg", f)},
        data={"lang": "en"},
        timeout=330,
    )
r.raise_for_status()
print(r.json()["text"])
```

```
import { readFile } from "node:fs/promises";

const form = new FormData();
form.append("file", new Blob([await readFile("page-1.jpg")]), "page-1.jpg");
form.append("lang", "en");
const r = await fetch("https://api.axforge.ai/v1/ocr", {
  method: "POST",
  headers: { Authorization: `Bearer ${process.env.AXFORGE_API_KEY}` },
  body: form,
});
console.log((await r.json()).text);
```

Or send JSON with the file in base64: `{"base64": "...", "filename": "contract.pdf", "lang": "sv"}`.
    Links to files are not fetched; send the file itself. Give your client a timeout of at least 6 minutes for long scans.

## The answer

```
{
  "text": "INVOICE 2041\nBrodbod AB, Stockholm\nDate 2026-09-01\nBread flour 25 kg EUR 480.00",
  "pages": [
    {"page": 1, "text": "INVOICE 2041\nBrodbod AB, Stockholm\nDate 2026-09-01\nBread flour 25 kg EUR 480.00",
     "source": "ocr", "lang": "en", "agreement": 1.0}
  ],
  "check": {"agreement": 1.0, "reviewCount": 0, "review": []},
  "totalPages": 1,
  "ocrPages": 1,
  "chars": 79,
  "object": "ocr.result",
  "model": "ocr-eu",
  "usage": {"pages": 1}
}
```

      - `text` is the whole document; `pages` has each page on its own.

      - `source` is `text-layer` for a digital PDF page (its text as it is in the file) and `ocr` for a page that was read.

      - `check` lists the places worth a second look on scanned pages (`review`: the page, the words before it, what was read). It is `null` when no page was scanned.

      - `usage.pages` is what was billed.

## Languages

`lang` takes one of the 24 EU languages: bg, cs, da, de, el, en, es, et, fi, fr, ga, hr, hu, it,
    lt, lv, mt, nl, pl, pt, ro, sk, sl, sv. Without it the language is found page by page; a page that is
    mostly numbers is then read as English, so pass `lang` for tables and financial statements.

## Limits

        | Files | One PDF or one image per call (PNG, JPEG, TIFF, WebP, BMP or GIF). Save Word, Excel and PowerPoint files as PDF first. |  |

        | Size | 20 MB per file |  |

        | Pages | Up to 80 scanned pages and 2,000 pages in all per call; split larger documents |  |

        | Time | About 3 seconds per scanned page; digital pages are near instant. A call waits up to 300 seconds. |  |

Printed text is read; handwriting is not, and tables come back as lines of text.
    Pages must be upright. A multi-page TIFF is read from its first page only; send a PDF instead.
    A page that mixes digital text with a scanned part is read from its text only.

## Errors

        | Status | When |  |

        | 400 | No file, base64 that does not decode, or a `lang` outside the 24 |  |

        | 413 | Over 20 MB, or more pages than one call reads (the answer says how many) |  |

        | 415 | Not a PDF or an image |  |

        | 422 | The PDF is password protected or damaged |  |

        | 503 | Reading is busy or starting up; try again after the `Retry-After` seconds. Nothing is charged. |  |



Source: https://dev.axforge.ai/docs/ocr/
