Convert any PDF into clean text and markdown using advanced OCR. Supports scanned PDFs, images, tables, and complex layouts with high accuracy across 30+ languages.
PDF to Text (OCR) is a powerful API that extracts text and structured markdown from any type of PDF using advanced OCR and document parsing — scanned, image-based, handwritten, multi-column, or digitally generated.
High-accuracy text recognition for scanned and image-based PDFs, even handwritten and low-quality scans.
English, Hindi, Arabic, Chinese, Japanese, Spanish, French, and more — including multilingual combinations.
Get clean, structured markdown with headings, lists, and tables preserved from the original layout.
Submit PDFs and get an instant job ID. Poll for status or receive a webhook when processing completes.
Upload PDFs directly, send a public URL, Google Drive share link, or pass base64-encoded data.
Extract all pages, a range, specific pages, or a single page — only the pages you process count toward your quota.
Files are processed in isolated workers and removed automatically. Output URLs are CDN-hosted with controlled access.
Simple REST API, flexible input formats, clear error messages, and consistent JSON responses across every endpoint.
Start free and scale as you grow. Every plan includes the full OCR engine — higher tiers unlock more pages, larger files, and premium features.
Everything you need to integrate the PDF to Text (OCR) API
All API requests require authentication using your API key. Send it via the x-api-key header with every request.
x-api-key: your_api_key_hereThe API exposes two endpoints — one to submit a PDF and one to check job status.
The API is fully asynchronous. Large or scanned PDFs can take from a few seconds to several minutes to process, so the API never blocks — it returns a job_id instantly and processes the PDF in a background worker queue.
Send a POST request to /pdf-to-text/v1 with the PDF and your options. The API validates input, reads page count, and queues the job. You get back a job_id and a status_url within seconds.
The worker picks up the job, runs OCR page-by-page, and stores the extracted text and markdown on the CDN. Most jobs finish in 5–60 seconds depending on page count and language complexity.
Either poll GET /pdf-to-text/status/{job_id} every 3–5 seconds, or supply a webhook_url at submission time to be notified the moment the job finishes.
When status is completed, the response includes output.txt and (if requested) output.md CDN URLs. Download or cache them — files stay on the CDN for 24 hours.
The API accepts the PDF in any of these formats. Use whichever fits your environment.
Upload the PDF as a multipart form field. The field name can be pdf, file, data, document, upload, or attachment — all are accepted.
curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
-H "x-api-key: your_api_key" \
-F "[email protected]"pdf_urlPass any public URL that returns a PDF. Use this when the PDF already lives on your CDN, S3, or any public host.
curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
-H "x-api-key: your_api_key" \
-d '{"pdf_url":"https://cdn.example.com/contract.pdf"}'pdf_urlPaste a Google Drive share link directly. The API automatically detects Drive links and fetches the file. Make sure the file is shared with "Anyone with the link can view."
curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
-H "x-api-key: your_api_key" \
-d '{"pdf_url":"https://drive.google.com/uc?export=download&id=1--EMd70IQHnLuVJXqXH......"}'pdf_base64Send the PDF inline as base64. Useful when your client can't perform multipart uploads (some no-code platforms).
PDF_B64=$(base64 -w0 document.pdf)
curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
-H "x-api-key: your_api_key" \
-d "{\"pdf_base64\":\"$PDF_B64\"}"Content-Type header — form-data, JSON, and query strings all work. The PDF is auto-detected from whatever field you send it in.| Header | Type | Description |
|---|---|---|
| x-api-keyRequired | String | Your API key. Send in request header. |
| Field | Type | Description |
|---|---|---|
| pdf / file / dataMultipart | File | PDF file uploaded as multipart form data. Any of these field names work. |
| pdf_urlJSON/Form | String | Public URL of a PDF, including direct CDN links and Google Drive share URLs. |
| pdf_base64JSON/Form | String | Base64-encoded PDF data. Supports raw base64 or data:application/pdf;base64,... prefix. |
| Field | Type | Description |
|---|---|---|
| langOptional | String | OCR language code or combination. Default: eng. Combine with + (e.g. eng+hin). See language list. |
| output_formatOptional | String | text (default), md, or both. Markdown requires Starter plan or higher. |
| extract_pagesOptional | String | Which pages to process. See page selection below. Default: all. |
| webhook_urlOptional | String | Public URL to POST the completed job to. Requires Starter plan or higher. |
extract_pagesControl exactly which pages get OCR'd. Only the pages you request count against your max_page_number plan limit.
| Value | Result |
|---|---|
all or omit | All pages in the document |
5 | Only page 5 (1 page) |
2to9 or 2-9 | Pages 2 through 9 (8 pages) |
2,7,9 | Pages 2, 7, and 9 only (3 pages) |
10to30 | Pages 10 through 30 (21 pages) |
2,7,9, only 3 pages are checked against your plan's page limit — not 100.curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
-H "x-api-key: your_api_key" \
-F "[email protected]"const form = new FormData();
form.append('pdf', fileInput.files[0]);
const res = await fetch('https://api.corenexis.com/pdf-to-text/v1', {
method: 'POST',
headers: { 'x-api-key': 'your_api_key' },
body: form
});
const data = await res.json();
console.log('Job ID:', data.data.job_id);import requests
with open("document.pdf", "rb") as f:
res = requests.post(
"https://api.corenexis.com/pdf-to-text/v1",
headers={"x-api-key": "your_api_key"},
files={"pdf": f}
)
print(res.json()["data"]["job_id"])$ch = curl_init("https://api.corenexis.com/pdf-to-text/v1");
curl_setopt_array($ch, [
CURLOPT_POST => true,
CURLOPT_RETURNTRANSFER => true,
CURLOPT_HTTPHEADER => ["x-api-key: your_api_key"],
CURLOPT_POSTFIELDS => ["pdf" => new CURLFile("document.pdf")],
]);
$resp = json_decode(curl_exec($ch), true);
echo "Job ID: " . $resp["data"]["job_id"];curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
-H "x-api-key: your_api_key" \
-F "[email protected]" \
-F "lang=eng+hin" \
-F "output_format=both"import requests
with open("bilingual.pdf", "rb") as f:
res = requests.post(
"https://api.corenexis.com/pdf-to-text/v1",
headers={"x-api-key": "your_api_key"},
files={"pdf": f},
data={"lang": "eng+hin", "output_format": "both"}
)
print(res.json())# Pages 4 through 9 — billed as 6 pages
curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
-H "x-api-key: your_api_key" \
-F "[email protected]" \
-F "extract_pages=4to9"# Pages 2, 7, and 9 only — billed as 3 pages
curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
-H "x-api-key: your_api_key" \
-F "[email protected]" \
-F "extract_pages=2,7,9"# Just page 5
curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
-H "x-api-key: your_api_key" \
-F "[email protected]" \
-F "extract_pages=5"curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
-H "x-api-key: your_api_key" \
-d '{"pdf_url":"https://cdn.example.com/contract.pdf","lang":"eng","output_format":"md"}'curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
-H "x-api-key: your_api_key" \
-d '{"pdf_url":"https://drive.google.com/uc?export=download&id=1--EMd70IQHnLuVJXqXHNoiO...L","lang":"eng+hin"}'curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
-H "x-api-key: your_api_key" \
-F "pdf_url=https://cdn.example.com/contract.pdf" \
-F "extract_pages=1to5"# 1. Submit
RESPONSE=$(curl -s -X POST "https://api.corenexis.com/pdf-to-text/v1" \
-H "x-api-key: your_api_key" \
-F "[email protected]")
JOB_ID=$(echo "$RESPONSE" | jq -r '.data.job_id')
echo "Job ID: $JOB_ID"
# 2. Poll every 3 seconds
while true; do
RESULT=$(curl -s "https://api.corenexis.com/pdf-to-text/status/$JOB_ID" -H "x-api-key: your_api_key")
STATUS=$(echo "$RESULT" | jq -r '.data.status')
echo "Status: $STATUS"
[ "$STATUS" != "processing" ] && break
sleep 3
done
# 3. Download text
TXT_URL=$(echo "$RESULT" | jq -r '.data.output.txt')
curl -o extracted.txt "$TXT_URL"
echo "Saved to extracted.txt"import requests, time
API_KEY = "your_api_key"
# 1. Submit
with open("document.pdf", "rb") as f:
sub = requests.post(
"https://api.corenexis.com/pdf-to-text/v1",
headers={"x-api-key": API_KEY},
files={"pdf": f},
data={"output_format": "both"}
).json()
job_id = sub["data"]["job_id"]
print("Job ID:", job_id)
# 2. Poll
while True:
r = requests.get(f"https://api.corenexis.com/pdf-to-text/status/{job_id}",
headers={"x-api-key": API_KEY}).json()
status = r["data"]["status"]
print("Status:", status)
if status != "processing":
break
time.sleep(3)
# 3. Download
txt_url = r["data"]["output"]["txt"]
text = requests.get(txt_url).text
print(text[:500])const API_KEY = 'your_api_key';
// 1. Submit
const form = new FormData();
form.append('pdf', fileInput.files[0]);
form.append('output_format', 'both');
const sub = await fetch('https://api.corenexis.com/pdf-to-text/v1', {
method: 'POST',
headers: { 'x-api-key': API_KEY },
body: form
}).then(r => r.json());
const jobId = sub.data.job_id;
console.log('Job ID:', jobId);
// 2. Poll
let result;
while (true) {
result = await fetch(`https://api.corenexis.com/pdf-to-text/status/${jobId}`, {
headers: { 'x-api-key': API_KEY }
}).then(r => r.json());
if (result.data.status !== 'processing') break;
await new Promise(r => setTimeout(r, 3000));
}
// 3. Download
const text = await fetch(result.data.output.txt).then(r => r.text());
console.log(text);curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
-H "x-api-key: your_api_key" \
-F "[email protected]" \
-F "output_format=both" \
-F "webhook_url=https://your-server.com/ocr-callback"After submission, use the status_url from the response (or build it manually) to check job progress.
| Status | Meaning |
|---|---|
| processing | Job is queued or actively being processed. Keep polling. |
| completed | All requested pages processed. output URLs are ready. |
| partial | 5-minute timeout reached. Some pages were processed and saved; remainder skipped. |
| error | Job failed (corrupt PDF, OCR engine failure). See error field in response. |
curl "https://api.corenexis.com/pdf-to-text/status/a3f9c1d2e4b567890123456789abcdef" \
-H "x-api-key: your_api_key"| Interval | Best For |
|---|---|
| Every 3 seconds | Interactive UI / waiting user — fastest feedback |
| Every 10 seconds | Background jobs, small PDFs |
| Every 30 seconds | Large scanned PDFs (50+ pages) |
Instead of polling, provide a webhook_url at submission time. We'll POST the complete job result to your URL as soon as processing finishes (completed, partial, or errored).
POST requests with Content-Type: application/json{
"job_id": "a3f9c1d2e4b567890123456789abcdef",
"status": "completed",
"lang": "eng+hin",
"submitted_at": 1748000000000,
"completed_at": 1748000060000,
"total_pages": 10,
"extraction_pages": "all",
"pages_processed": 10,
"output_format": "both",
"webhook_url": "https://your-server.com/ocr-callback",
"output": {
"txt": "https://cdn.corenexis.com/user_content/documents/a3f9c1d2.txt",
"md": "https://cdn.corenexis.com/user_content/documents/a3f9c1d2.md"
}
}app.post('/ocr-callback', express.json(), async (req, res) => {
const { job_id, status, output } = req.body;
res.sendStatus(200); // ack immediately
if (status === 'completed' && output?.txt) {
const text = await fetch(output.txt).then(r => r.text());
// ... save text, trigger next step, notify user ...
}
});webhook_url field and you're done.{
"success": true,
"plan": "starter",
"data": {
"status": "processing",
"job_id": "a3f9c1d2e4b567890123456789abcdef",
"status_url": "https://api.corenexis.com/pdf-to-text/status/a3f9c1d2e4b567890123456789abcdef",
"total_pages": 20,
"pages_to_process": 20,
"file_size": "2 MB"
},
"usage": {
"remaining": 98,
"rate_limit": 10,
"monthly_limit": 100
}
}{
"success": true,
"plan": "starter",
"data": {
"job_id": "a3f9c1d2e4b567890123456789abcdef",
"status": "processing",
"lang": "eng",
"submitted_at": 1748000000000,
"webhook_url": null
},
"usage": { "remaining": 98, "rate_limit": 10, "monthly_limit": 100 }
}{
"success": true,
"plan": "starter",
"data": {
"job_id": "a3f9c1d2e4b567890123456789abcdef",
"status": "completed",
"lang": "eng+hin",
"submitted_at": 1748000000000,
"completed_at": 1748000060000,
"file_size": "500 KB",
"total_pages": 10,
"extraction_pages": "all",
"pages_processed": 10,
"output_format": "both",
"webhook_url": null,
"output": {
"txt": "https://cdn.corenexis.com/user_content/documents/a3f9c1d2.txt",
"md": "https://cdn.corenexis.com/user_content/documents/a3f9c1d2.md"
}
},
"usage": { "remaining": 98, "rate_limit": 10, "monthly_limit": 100 }
}{
"success": true,
"plan": "starter",
"data": {
"job_id": "a3f9c1d2e4b567890123456789abcdef",
"status": "partial",
"warning": "Processing timed out after 5 minutes. Only 8 of 50 pages processed.",
"lang": "eng",
"submitted_at": 1748000000000,
"completed_at": 1748000300000,
"file_size": "5 MB",
"total_pages": 50,
"extraction_pages": "all",
"pages_processed": 8,
"output_format": "text",
"output": {
"txt": "https://cdn.corenexis.com/user_content/documents/a3f9c1d2.txt"
}
},
"usage": { "remaining": 98, "rate_limit": 10, "monthly_limit": 100 }
}{
"success": true,
"plan": "starter",
"data": {
"job_id": "a3f9c1d2e4b567890123456789abcdef",
"status": "error",
"error": "PDF to image conversion failed.",
"submitted_at": 1748000000000,
"completed_at": 1748000005000
},
"usage": { "remaining": 98, "rate_limit": 10, "monthly_limit": 100 }
}| Field | Type | Description |
|---|---|---|
| success | Boolean | true when the API call was accepted (job created or status retrieved) |
| plan | String | Your current plan slug (free, starter, pro) |
| data.job_id | String | 32-character unique job identifier |
| data.status_url | String | Full URL to poll for status |
| data.status | String | processing / completed / partial / error |
| data.total_pages | Integer | Total pages in the original PDF |
| data.pages_to_process | Integer | Pages that will be OCR'd (per extract_pages) |
| data.pages_processed | Integer | Pages actually completed (set after processing) |
| data.file_size | String | Human-readable file size (e.g. 500 KB, 2 MB) |
| data.output.txt | String | CDN URL to plain text output (24-hour expiry) |
| data.output.md | String | CDN URL to markdown output (24-hour expiry) |
| usage.remaining | Integer | Extractions remaining this billing period |
| usage.rate_limit | Integer | Max requests per minute on your plan |
| usage.monthly_limit | Integer | Total monthly extractions on your plan |
Combine multiple language codes with + for multilingual PDFs (e.g. eng+hin, chi_sim+eng). Multilingual and non-English language codes require Starter plan or higher.
| Code | Language | Code | Language |
|---|---|---|---|
eng | English | hin | Hindi |
ara | Arabic | fra | French |
deu | German | spa | Spanish |
por | Portuguese | ita | Italian |
rus | Russian | chi_sim | Chinese (Simplified) |
chi_tra | Chinese (Traditional) | jpn | Japanese |
kor | Korean | ben | Bengali |
urd | Urdu | tam | Tamil |
tel | Telugu | mar | Marathi |
guj | Gujarati | kan | Kannada |
mal | Malayalam | pan | Punjabi |
nld | Dutch | pol | Polish |
tur | Turkish | vie | Vietnamese |
tha | Thai | ind | Indonesian |
fas | Persian (Farsi) | heb | Hebrew |
Every error response is consistent JSON: { success: false, code: "...", message: "..." }. Many errors include extra fields (param, provided, max_allowed) to help you self-correct without contacting support.
| HTTP | Code | Description |
|---|---|---|
| 400 | INVALID_INPUT | Bad input — missing PDF, invalid extract_pages, unsupported language code, malformed file. |
| 400 | PARAM_LIMIT_EXCEEDED | Request exceeds your plan's limit (file size, pages, feature). Includes param + max_allowed fields. |
| 400 | INVALID_BODY | Request body is not valid JSON. |
| 401 | MISSING_KEY | x-api-key header not present. |
| 401 | INVALID_KEY | API key is invalid, expired, or not found. |
| 403 | KEY_DISABLED | API key has been disabled. Generate a new one. |
| 403 | ACCOUNT_SUSPENDED | Account suspended. Contact support. |
| 403 | ACCOUNT_INACTIVE | Account not activated. Verify your email first. |
| 403 | EMAIL_NOT_VERIFIED | Email address has not been verified. |
| 402 | NO_SUBSCRIPTION | No active subscription on the PDF to Text API. Subscribe first. |
| 402 | SUBSCRIPTION_EXPIRED | Subscription expired. Renew to continue. |
| 402 | SUBSCRIPTION_CANCELLED | Subscription cancelled. |
| 402 | SUBSCRIPTION_INACTIVE | Subscription is inactive. |
| 402 | BILLING_REQUIRES_ACTION | Billing requires action. Update payment method. |
| 404 | JOB_NOT_FOUND | No job exists for the provided job ID (status endpoint only). |
| 404 | API_NOT_FOUND | API slug not configured on the platform. |
| 429 | RATE_LIMIT_EXCEEDED | Too many requests per minute. Wait 60 seconds. |
| 429 | QUOTA_EXCEEDED | Monthly quota reached. Upgrade your plan or wait for renewal. |
| 502 | PROCESSING_FAILED | Internal OCR service rejected or failed the job. Quota was not used. |
| 503 | API_UNAVAILABLE | Service temporarily unavailable. Retry shortly. |
| 503 | AUTH_SERVICE_UNAVAILABLE | Authentication backend is temporarily unreachable. Retry. |
| 503 | UPSTREAM_UNAVAILABLE | OCR worker is unavailable. Retry. |
{
"success": false,
"code": "PARAM_LIMIT_EXCEEDED",
"message": "Invalid value for 'max_page_number'. Allowed for your plan: 12 or less. Your request contains: 50. Please change 'max_page_number' or upgrade your plan.",
"param": "max_page_number",
"provided": 50,
"op": "lte",
"max_allowed": 12
}{
"success": false,
"code": "PARAM_LIMIT_EXCEEDED",
"message": "Invalid value for 'max_file_size_bytes'. Allowed for your plan: 2097152 or less. Your request contains: 5242880. Please change 'max_file_size_bytes' or upgrade your plan.",
"param": "max_file_size_bytes",
"provided": 5242880,
"op": "lte",
"max_allowed": 2097152
}{
"success": false,
"code": "PARAM_LIMIT_EXCEEDED",
"message": "Feature 'markdown' is not allowed on your current plan. Allowed for your plan: false. Your request contains: true. Please disable 'markdown' or upgrade your plan.",
"param": "markdown",
"provided": true,
"allowed": false
}{
"success": false,
"code": "INVALID_INPUT",
"message": "Unsupported language code 'klingon'. See /pdf-ocr/languages for the supported list."
}{
"success": false,
"code": "QUOTA_EXCEEDED",
"message": "Monthly quota reached. Upgrade your plan or wait for the next billing cycle."
}{
"success": false,
"code": "RATE_LIMIT_EXCEEDED",
"message": "Too many requests. Please wait before making another request."
}Create your free account and start converting PDFs to text and markdown in seconds. No credit card required.