New User Offer Get $5 free credit instantly — no credit card required. Use code WELCOME Claim $5 Free
API V1 — Production Ready

PDF to Text & Markdown OCR API

Convert any PDF into clean text and markdown using advanced OCR. Supports scanned PDFs, images, tables, and complex layouts with high accuracy across 30+ languages.

30+
Languages
50MB
Max PDF Size
Async
Queue Processing
99.9%
Uptime SLA

Built for Every PDF Workflow

PDF to Text (OCR) is a powerful API that extracts text and structured markdown from any type of PDF using advanced OCR and document parsing — scanned, image-based, handwritten, multi-column, or digitally generated.

Advanced OCR

High-accuracy text recognition for scanned and image-based PDFs, even handwritten and low-quality scans.

30+ Languages

English, Hindi, Arabic, Chinese, Japanese, Spanish, French, and more — including multilingual combinations.

Markdown Output

Get clean, structured markdown with headings, lists, and tables preserved from the original layout.

Async Queue

Submit PDFs and get an instant job ID. Poll for status or receive a webhook when processing completes.

Multiple Input Modes

Upload PDFs directly, send a public URL, Google Drive share link, or pass base64-encoded data.

Page Selection

Extract all pages, a range, specific pages, or a single page — only the pages you process count toward your quota.

Secure Processing

Files are processed in isolated workers and removed automatically. Output URLs are CDN-hosted with controlled access.

Developer Friendly

Simple REST API, flexible input formats, clear error messages, and consistent JSON responses across every endpoint.

Simple, Transparent Pricing

Start free and scale as you grow. Every plan includes the full OCR engine — higher tiers unlock more pages, larger files, and premium features.

Free
Free
  • 10 extractions/month
  • 5 requests/minute
  • Up to 12 pages per PDF
  • Up to 2 MB file size
  • English OCR only (eng)
  • Plain text output
Start Free
Starter
$4.99/month
  • 100 extractions/month
  • 10 requests/minute
  • Up to 24 pages per PDF
  • Up to 12 MB file size
  • All 30+ languages (incl. multilingual)
  • Markdown output
  • Webhook callbacks
  • Email support
Get Started
Pro
$14.99/month
  • 500 extractions/month
  • 30 requests/minute
  • Up to 49 pages per PDF
  • Up to 15 MB file size
  • All 30+ languages (incl. multilingual)
  • Markdown output
  • Webhook callbacks
  • Priority support
Get Started
Max
$29.99/month
  • 100 extractions/month
  • 60 requests/minute
  • Up to 50 pages per PDF
  • Up to 20 MB file size
  • All 30+ languages (incl. multilingual)
  • Markdown output
  • Webhook callbacks
  • Dedicated support
  • Accelerated processing
Get Started
Enterprise

Enterprise & Agency Plan

High-volume document workflows, custom retention, dedicated workers, and SLAs. We'll build a plan around your exact requirements.

Unlimited extractions
100+ MB files
1000+ pages
Dedicated queue
24/7 Support

API Reference

Everything you need to integrate the PDF to Text (OCR) API

Authentication

All API requests require authentication using your API key. Send it via the x-api-key header with every request.

Header
x-api-key: your_api_key_here
Get Your API KeySign up at dash.corenexis.com to get your API key instantly.

API Endpoints

The API exposes two endpoints — one to submit a PDF and one to check job status.

Submit a PDF for extraction

POSThttps://api.corenexis.com/pdf-to-text/v1

Check job status

GEThttps://api.corenexis.com/pdf-to-text/status/{job_id}
Free Status ChecksStatus polling does not consume monthly quota — only successful extraction submissions do.

How Queue Processing Works

The API is fully asynchronous. Large or scanned PDFs can take from a few seconds to several minutes to process, so the API never blocks — it returns a job_id instantly and processes the PDF in a background worker queue.

Submit the PDF

Send a POST request to /pdf-to-text/v1 with the PDF and your options. The API validates input, reads page count, and queues the job. You get back a job_id and a status_url within seconds.

Wait for processing

The worker picks up the job, runs OCR page-by-page, and stores the extracted text and markdown on the CDN. Most jobs finish in 5–60 seconds depending on page count and language complexity.

Get the result — poll OR webhook

Either poll GET /pdf-to-text/status/{job_id} every 3–5 seconds, or supply a webhook_url at submission time to be notified the moment the job finishes.

Download the output

When status is completed, the response includes output.txt and (if requested) output.md CDN URLs. Download or cache them — files stay on the CDN for 24 hours.

When to poll vs use a webhookFor interactive UIs, poll every 3–5 seconds. For background jobs and integrations (n8n, Zapier, custom servers), webhooks are simpler — no polling overhead, your endpoint is called once when the job finishes.

Input Modes — Four Ways to Send a PDF

The API accepts the PDF in any of these formats. Use whichever fits your environment.

1. Direct file upload (multipart) — recommended

Upload the PDF as a multipart form field. The field name can be pdf, file, data, document, upload, or attachment — all are accepted.

cURL
curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
  -H "x-api-key: your_api_key" \
  -F "[email protected]"

2. Public URL — pdf_url

Pass any public URL that returns a PDF. Use this when the PDF already lives on your CDN, S3, or any public host.

cURL
curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
  -H "x-api-key: your_api_key" \
  -d '{"pdf_url":"https://cdn.example.com/contract.pdf"}'

3. Google Drive share link — pdf_url

Paste a Google Drive share link directly. The API automatically detects Drive links and fetches the file. Make sure the file is shared with "Anyone with the link can view."

cURL
curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
  -H "x-api-key: your_api_key" \
  -d '{"pdf_url":"https://drive.google.com/uc?export=download&id=1--EMd70IQHnLuVJXqXH......"}'

4. Base64-encoded PDF — pdf_base64

Send the PDF inline as base64. Useful when your client can't perform multipart uploads (some no-code platforms).

cURL
PDF_B64=$(base64 -w0 document.pdf)
curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
  -H "x-api-key: your_api_key" \
  -d "{\"pdf_base64\":\"$PDF_B64\"}"
Smart Input DetectionThe API does not require a specific Content-Type header — form-data, JSON, and query strings all work. The PDF is auto-detected from whatever field you send it in.

Request Parameters

Headers

HeaderTypeDescription
x-api-keyRequiredStringYour API key. Send in request header.

PDF Source (one of)

FieldTypeDescription
pdf / file / dataMultipartFilePDF file uploaded as multipart form data. Any of these field names work.
pdf_urlJSON/FormStringPublic URL of a PDF, including direct CDN links and Google Drive share URLs.
pdf_base64JSON/FormStringBase64-encoded PDF data. Supports raw base64 or data:application/pdf;base64,... prefix.

Processing Options

FieldTypeDescription
langOptionalStringOCR language code or combination. Default: eng. Combine with + (e.g. eng+hin). See language list.
output_formatOptionalStringtext (default), md, or both. Markdown requires Starter plan or higher.
extract_pagesOptionalStringWhich pages to process. See page selection below. Default: all.
webhook_urlOptionalStringPublic URL to POST the completed job to. Requires Starter plan or higher.

Page Selection — extract_pages

Control exactly which pages get OCR'd. Only the pages you request count against your max_page_number plan limit.

ValueResult
all or omitAll pages in the document
5Only page 5 (1 page)
2to9 or 2-9Pages 2 through 9 (8 pages)
2,7,9Pages 2, 7, and 9 only (3 pages)
10to30Pages 10 through 30 (21 pages)
Pages billed = pages processedIf your PDF is 100 pages but you request 2,7,9, only 3 pages are checked against your plan's page limit — not 100.

Code Examples

Simple Extraction

curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
  -H "x-api-key: your_api_key" \
  -F "[email protected]"
const form = new FormData();
form.append('pdf', fileInput.files[0]);

const res = await fetch('https://api.corenexis.com/pdf-to-text/v1', {
  method: 'POST',
  headers: { 'x-api-key': 'your_api_key' },
  body: form
});
const data = await res.json();
console.log('Job ID:', data.data.job_id);
import requests

with open("document.pdf", "rb") as f:
    res = requests.post(
        "https://api.corenexis.com/pdf-to-text/v1",
        headers={"x-api-key": "your_api_key"},
        files={"pdf": f}
    )
print(res.json()["data"]["job_id"])
$ch = curl_init("https://api.corenexis.com/pdf-to-text/v1");
curl_setopt_array($ch, [
    CURLOPT_POST => true,
    CURLOPT_RETURNTRANSFER => true,
    CURLOPT_HTTPHEADER => ["x-api-key: your_api_key"],
    CURLOPT_POSTFIELDS => ["pdf" => new CURLFile("document.pdf")],
]);
$resp = json_decode(curl_exec($ch), true);
echo "Job ID: " . $resp["data"]["job_id"];

Multilingual PDF (English + Hindi) with Markdown

curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
  -H "x-api-key: your_api_key" \
  -F "[email protected]" \
  -F "lang=eng+hin" \
  -F "output_format=both"
import requests
with open("bilingual.pdf", "rb") as f:
    res = requests.post(
        "https://api.corenexis.com/pdf-to-text/v1",
        headers={"x-api-key": "your_api_key"},
        files={"pdf": f},
        data={"lang": "eng+hin", "output_format": "both"}
    )
print(res.json())

Extract Specific Pages

# Pages 4 through 9 — billed as 6 pages
curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
  -H "x-api-key: your_api_key" \
  -F "[email protected]" \
  -F "extract_pages=4to9"
# Pages 2, 7, and 9 only — billed as 3 pages
curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
  -H "x-api-key: your_api_key" \
  -F "[email protected]" \
  -F "extract_pages=2,7,9"
# Just page 5
curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
  -H "x-api-key: your_api_key" \
  -F "[email protected]" \
  -F "extract_pages=5"

PDF from URL / Google Drive

curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
  -H "x-api-key: your_api_key" \
  -d '{"pdf_url":"https://cdn.example.com/contract.pdf","lang":"eng","output_format":"md"}'
curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
  -H "x-api-key: your_api_key" \
  -d '{"pdf_url":"https://drive.google.com/uc?export=download&id=1--EMd70IQHnLuVJXqXHNoiO...L","lang":"eng+hin"}'
curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
  -H "x-api-key: your_api_key" \
  -F "pdf_url=https://cdn.example.com/contract.pdf" \
  -F "extract_pages=1to5"

Full Submit + Poll + Download Pipeline

# 1. Submit
RESPONSE=$(curl -s -X POST "https://api.corenexis.com/pdf-to-text/v1" \
  -H "x-api-key: your_api_key" \
  -F "[email protected]")

JOB_ID=$(echo "$RESPONSE" | jq -r '.data.job_id')
echo "Job ID: $JOB_ID"

# 2. Poll every 3 seconds
while true; do
  RESULT=$(curl -s "https://api.corenexis.com/pdf-to-text/status/$JOB_ID" -H "x-api-key: your_api_key")
  STATUS=$(echo "$RESULT" | jq -r '.data.status')
  echo "Status: $STATUS"
  [ "$STATUS" != "processing" ] && break
  sleep 3
done

# 3. Download text
TXT_URL=$(echo "$RESULT" | jq -r '.data.output.txt')
curl -o extracted.txt "$TXT_URL"
echo "Saved to extracted.txt"
import requests, time

API_KEY = "your_api_key"

# 1. Submit
with open("document.pdf", "rb") as f:
    sub = requests.post(
        "https://api.corenexis.com/pdf-to-text/v1",
        headers={"x-api-key": API_KEY},
        files={"pdf": f},
        data={"output_format": "both"}
    ).json()

job_id = sub["data"]["job_id"]
print("Job ID:", job_id)

# 2. Poll
while True:
    r = requests.get(f"https://api.corenexis.com/pdf-to-text/status/{job_id}",
                     headers={"x-api-key": API_KEY}).json()
    status = r["data"]["status"]
    print("Status:", status)
    if status != "processing":
        break
    time.sleep(3)

# 3. Download
txt_url = r["data"]["output"]["txt"]
text = requests.get(txt_url).text
print(text[:500])
const API_KEY = 'your_api_key';

// 1. Submit
const form = new FormData();
form.append('pdf', fileInput.files[0]);
form.append('output_format', 'both');

const sub = await fetch('https://api.corenexis.com/pdf-to-text/v1', {
  method: 'POST',
  headers: { 'x-api-key': API_KEY },
  body: form
}).then(r => r.json());

const jobId = sub.data.job_id;
console.log('Job ID:', jobId);

// 2. Poll
let result;
while (true) {
  result = await fetch(`https://api.corenexis.com/pdf-to-text/status/${jobId}`, {
    headers: { 'x-api-key': API_KEY }
  }).then(r => r.json());
  if (result.data.status !== 'processing') break;
  await new Promise(r => setTimeout(r, 3000));
}

// 3. Download
const text = await fetch(result.data.output.txt).then(r => r.text());
console.log(text);

With Webhook (no polling needed)

cURL
curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
  -H "x-api-key: your_api_key" \
  -F "[email protected]" \
  -F "output_format=both" \
  -F "webhook_url=https://your-server.com/ocr-callback"

Status Endpoint & Polling

After submission, use the status_url from the response (or build it manually) to check job progress.

GEThttps://api.corenexis.com/pdf-to-text/status/{job_id}

Possible Status Values

StatusMeaning
processingJob is queued or actively being processed. Keep polling.
completedAll requested pages processed. output URLs are ready.
partial5-minute timeout reached. Some pages were processed and saved; remainder skipped.
errorJob failed (corrupt PDF, OCR engine failure). See error field in response.

Status Check Example

cURL
curl "https://api.corenexis.com/pdf-to-text/status/a3f9c1d2e4b567890123456789abcdef" \
  -H "x-api-key: your_api_key"

Polling Recommendations

IntervalBest For
Every 3 secondsInteractive UI / waiting user — fastest feedback
Every 10 secondsBackground jobs, small PDFs
Every 30 secondsLarge scanned PDFs (50+ pages)
Status checks rate-limitStatus calls count toward your per-minute rate limit but not your monthly quota. Avoid polling more than once every 2 seconds.

Webhooks

Instead of polling, provide a webhook_url at submission time. We'll POST the complete job result to your URL as soon as processing finishes (completed, partial, or errored).

Webhook Requirements

  • URL must be publicly accessible (no localhost / private IPs)
  • Must accept POST requests with Content-Type: application/json
  • Should respond with HTTP 2xx within 10 seconds
  • Available on Starter plan or higher

Webhook Payload (same shape as completed status)

POST → your_webhook_url
{
  "job_id": "a3f9c1d2e4b567890123456789abcdef",
  "status": "completed",
  "lang": "eng+hin",
  "submitted_at": 1748000000000,
  "completed_at": 1748000060000,
  "total_pages": 10,
  "extraction_pages": "all",
  "pages_processed": 10,
  "output_format": "both",
  "webhook_url": "https://your-server.com/ocr-callback",
  "output": {
    "txt": "https://cdn.corenexis.com/user_content/documents/a3f9c1d2.txt",
    "md":  "https://cdn.corenexis.com/user_content/documents/a3f9c1d2.md"
  }
}

Example Webhook Handler (Node.js)

Express
app.post('/ocr-callback', express.json(), async (req, res) => {
  const { job_id, status, output } = req.body;
  res.sendStatus(200); // ack immediately

  if (status === 'completed' && output?.txt) {
    const text = await fetch(output.txt).then(r => r.text());
    // ... save text, trigger next step, notify user ...
  }
});
n8n / Zapier / MakeWebhook URL is the easiest way to integrate with no-code automation tools. Drop your workflow's webhook URL into the webhook_url field and you're done.

Response Format

Submit Response — Job Queued

POST /pdf-to-text/v1
{
  "success": true,
  "plan": "starter",
  "data": {
    "status": "processing",
    "job_id": "a3f9c1d2e4b567890123456789abcdef",
    "status_url": "https://api.corenexis.com/pdf-to-text/status/a3f9c1d2e4b567890123456789abcdef",
    "total_pages": 20,
    "pages_to_process": 20,
    "file_size": "2 MB"
  },
  "usage": {
    "remaining": 98,
    "rate_limit": 10,
    "monthly_limit": 100
  }
}

Status Response — Variants

{
  "success": true,
  "plan": "starter",
  "data": {
    "job_id": "a3f9c1d2e4b567890123456789abcdef",
    "status": "processing",
    "lang": "eng",
    "submitted_at": 1748000000000,
    "webhook_url": null
  },
  "usage": { "remaining": 98, "rate_limit": 10, "monthly_limit": 100 }
}
{
  "success": true,
  "plan": "starter",
  "data": {
    "job_id": "a3f9c1d2e4b567890123456789abcdef",
    "status": "completed",
    "lang": "eng+hin",
    "submitted_at": 1748000000000,
    "completed_at": 1748000060000,
    "file_size": "500 KB",
    "total_pages": 10,
    "extraction_pages": "all",
    "pages_processed": 10,
    "output_format": "both",
    "webhook_url": null,
    "output": {
      "txt": "https://cdn.corenexis.com/user_content/documents/a3f9c1d2.txt",
      "md":  "https://cdn.corenexis.com/user_content/documents/a3f9c1d2.md"
    }
  },
  "usage": { "remaining": 98, "rate_limit": 10, "monthly_limit": 100 }
}
{
  "success": true,
  "plan": "starter",
  "data": {
    "job_id": "a3f9c1d2e4b567890123456789abcdef",
    "status": "partial",
    "warning": "Processing timed out after 5 minutes. Only 8 of 50 pages processed.",
    "lang": "eng",
    "submitted_at": 1748000000000,
    "completed_at": 1748000300000,
    "file_size": "5 MB",
    "total_pages": 50,
    "extraction_pages": "all",
    "pages_processed": 8,
    "output_format": "text",
    "output": {
      "txt": "https://cdn.corenexis.com/user_content/documents/a3f9c1d2.txt"
    }
  },
  "usage": { "remaining": 98, "rate_limit": 10, "monthly_limit": 100 }
}
{
  "success": true,
  "plan": "starter",
  "data": {
    "job_id": "a3f9c1d2e4b567890123456789abcdef",
    "status": "error",
    "error": "PDF to image conversion failed.",
    "submitted_at": 1748000000000,
    "completed_at": 1748000005000
  },
  "usage": { "remaining": 98, "rate_limit": 10, "monthly_limit": 100 }
}

Response Fields

FieldTypeDescription
successBooleantrue when the API call was accepted (job created or status retrieved)
planStringYour current plan slug (free, starter, pro)
data.job_idString32-character unique job identifier
data.status_urlStringFull URL to poll for status
data.statusStringprocessing / completed / partial / error
data.total_pagesIntegerTotal pages in the original PDF
data.pages_to_processIntegerPages that will be OCR'd (per extract_pages)
data.pages_processedIntegerPages actually completed (set after processing)
data.file_sizeStringHuman-readable file size (e.g. 500 KB, 2 MB)
data.output.txtStringCDN URL to plain text output (24-hour expiry)
data.output.mdStringCDN URL to markdown output (24-hour expiry)
usage.remainingIntegerExtractions remaining this billing period
usage.rate_limitIntegerMax requests per minute on your plan
usage.monthly_limitIntegerTotal monthly extractions on your plan
Output URL ExpiryExtracted text and markdown CDN URLs are available for 24 hours after job completion. Download or cache them within that window.

Supported Languages

Combine multiple language codes with + for multilingual PDFs (e.g. eng+hin, chi_sim+eng). Multilingual and non-English language codes require Starter plan or higher.

CodeLanguageCodeLanguage
engEnglishhinHindi
araArabicfraFrench
deuGermanspaSpanish
porPortugueseitaItalian
rusRussianchi_simChinese (Simplified)
chi_traChinese (Traditional)jpnJapanese
korKoreanbenBengali
urdUrdutamTamil
telTelugumarMarathi
gujGujaratikanKannada
malMalayalampanPunjabi
nldDutchpolPolish
turTurkishvieVietnamese
thaThaiindIndonesian
fasPersian (Farsi)hebHebrew

Error Codes

Every error response is consistent JSON: { success: false, code: "...", message: "..." }. Many errors include extra fields (param, provided, max_allowed) to help you self-correct without contacting support.

HTTP Status Reference

HTTPCodeDescription
400INVALID_INPUTBad input — missing PDF, invalid extract_pages, unsupported language code, malformed file.
400PARAM_LIMIT_EXCEEDEDRequest exceeds your plan's limit (file size, pages, feature). Includes param + max_allowed fields.
400INVALID_BODYRequest body is not valid JSON.
401MISSING_KEYx-api-key header not present.
401INVALID_KEYAPI key is invalid, expired, or not found.
403KEY_DISABLEDAPI key has been disabled. Generate a new one.
403ACCOUNT_SUSPENDEDAccount suspended. Contact support.
403ACCOUNT_INACTIVEAccount not activated. Verify your email first.
403EMAIL_NOT_VERIFIEDEmail address has not been verified.
402NO_SUBSCRIPTIONNo active subscription on the PDF to Text API. Subscribe first.
402SUBSCRIPTION_EXPIREDSubscription expired. Renew to continue.
402SUBSCRIPTION_CANCELLEDSubscription cancelled.
402SUBSCRIPTION_INACTIVESubscription is inactive.
402BILLING_REQUIRES_ACTIONBilling requires action. Update payment method.
404JOB_NOT_FOUNDNo job exists for the provided job ID (status endpoint only).
404API_NOT_FOUNDAPI slug not configured on the platform.
429RATE_LIMIT_EXCEEDEDToo many requests per minute. Wait 60 seconds.
429QUOTA_EXCEEDEDMonthly quota reached. Upgrade your plan or wait for renewal.
502PROCESSING_FAILEDInternal OCR service rejected or failed the job. Quota was not used.
503API_UNAVAILABLEService temporarily unavailable. Retry shortly.
503AUTH_SERVICE_UNAVAILABLEAuthentication backend is temporarily unreachable. Retry.
503UPSTREAM_UNAVAILABLEOCR worker is unavailable. Retry.

Example Error Responses

{
  "success": false,
  "code": "PARAM_LIMIT_EXCEEDED",
  "message": "Invalid value for 'max_page_number'. Allowed for your plan: 12 or less. Your request contains: 50. Please change 'max_page_number' or upgrade your plan.",
  "param": "max_page_number",
  "provided": 50,
  "op": "lte",
  "max_allowed": 12
}
{
  "success": false,
  "code": "PARAM_LIMIT_EXCEEDED",
  "message": "Invalid value for 'max_file_size_bytes'. Allowed for your plan: 2097152 or less. Your request contains: 5242880. Please change 'max_file_size_bytes' or upgrade your plan.",
  "param": "max_file_size_bytes",
  "provided": 5242880,
  "op": "lte",
  "max_allowed": 2097152
}
{
  "success": false,
  "code": "PARAM_LIMIT_EXCEEDED",
  "message": "Feature 'markdown' is not allowed on your current plan. Allowed for your plan: false. Your request contains: true. Please disable 'markdown' or upgrade your plan.",
  "param": "markdown",
  "provided": true,
  "allowed": false
}
{
  "success": false,
  "code": "INVALID_INPUT",
  "message": "Unsupported language code 'klingon'. See /pdf-ocr/languages for the supported list."
}
{
  "success": false,
  "code": "QUOTA_EXCEEDED",
  "message": "Monthly quota reached. Upgrade your plan or wait for the next billing cycle."
}
{
  "success": false,
  "code": "RATE_LIMIT_EXCEEDED",
  "message": "Too many requests. Please wait before making another request."
}
Quota-Safe ErrorsWhen a request is rejected because of plan limits (page count, file size, feature flags), language whitelist, or upstream failure — your monthly quota is not consumed. Only fully accepted submissions count.
Rate LimitingFree 5/min · Starter 10/min · Pro 30/min. Status checks count toward the per-minute limit but not the monthly quota.

Ready to Extract Smarter?

Create your free account and start converting PDFs to text and markdown in seconds. No credit card required.

Stay in the Loop

Get the latest updates delivered straight to your inbox