Snip Snipping Tool Chrome Extension OCR API Files API Private Cloud OCR Secure Conversion Service
Make Documents Accessible Process Chemical Documents Collaborate on Documents OCR API for Developers Train Language Models Support Academic Research Artificial Intelligence Fintech Edtech Pharma & Chemical Universities & Schools
Handwriting Recognition Digital Ink On-prem PDF Cloud Mathpix Markdown All Supported Languages Image Conversion PDF Conversion Markdown Conversion Table OCR Mathpix CLI PDF Search PDF Reader PDF Data Extraction Chrome Extension View Conversion Gallery
Snip APIs Private Cloud OCR SCS
Mobile Desktop Web Chrome Extension
Mathpix Snip Apps Mathpix OCR API Mathpix Markdown Python SDK
Blog
About Careers Contact
Contact Get Started
← Back to Blog

Introducing Mathpix Private Cloud OCR

2026-09-30 · api, updates
Today we’re excited to introduce Mathpix Private Cloud OCR (PCO): the full Mathpix document API, deployed entirely inside your own infrastructure.
PCO gives teams the accuracy and document understanding of Mathpix without documents ever leaving their network. It runs as a single container in your Kubernetes cluster or on a single GPU host, across AWS, GCP, Azure, or your own data center.

What’s new

PCO exposes the same endpoints as api.mathpix.com, with the same request and response shapes, so code you wrote against our hosted API works unchanged once you point it at your deployment (the short list of options a deployment refuses is in the overview):
  • POST /v3/pdf, GET /v3/pdf/{pdf_id}, GET /v3/pdf/{pdf_id}.{ext} - document OCR: PDF, Word, PowerPoint, Excel, EPUB, TIFF and images in; Mathpix Markdown, lines JSON, Markdown, HTML, LaTeX, DOCX, PPTX, XLSX and zips out
  • POST /v3/text - one image, synchronously
  • POST /pco/v1/jobs - batch conversion over a folder in your own bucket: the deployment lists the folder, converts every document, writes the outputs next to the inputs, and leaves a manifest and a report in the same bucket. Per-file status with error ids, retry only what failed, cancel, and webhooks with the same signed contract as our hosted API
  • GET /health, GET /metrics, GET /pco/v1/status, GET /pco/v1/usage, GET /pco/v1/license - health, Prometheus metrics, versions, the pages you processed, and the license state, all on your side
And a command-line tool, pco, that converts a file, a folder on your disk or a folder in your bucket against your deployment, with progress, resume and a report. It will be released as open source.

Examples

POST /v3/pdf, exactly as on the hosted API, minus the API keys:
$ curl -sF file=@paper.pdf -F 'options_json={"conversion_formats":{"docx":true}}' http://pco.internal:8080/v3/pdf
{"pdf_id": "d9c1e1f0c2d94b1e8f3a5b7c9d0e1f2a", "status": "processing"}

$ curl -s http://pco.internal:8080/v3/pdf/d9c1e1f0c2d94b1e8f3a5b7c9d0e1f2a
{"pdf_id": "d9c1e1f0...", "status": "completed", "num_pages": 12, "num_pages_completed": 12, "percent_done": 100.0,
 "conversion_status": {"docx": {"status": "completed"}}}
POST /pco/v1/jobs, a whole folder:
// Request
{
  "input": {"folder": "s3://example-docs/arxiv/2024-q3/"},
  "options": {"conversion_formats": {"md": true, "docx": true}},
  "callback_url": "https://hooks.example.internal/mathpix", "callback_events": ["job.completed", "file.error"]
}

// Response
{
  "job_id": "j_7f3a9c1e4b", "status": "enumerating",
  "files": {"discovered": 1820, "queued": 1790, "processing": 24, "completed": 6, "failed": 0, "skipped": 0},
  "output": {"report_uri": "s3://example-docs/arxiv/2024-q3/_mathpix/j_7f3a9c1e4b/report.jsonl", ...}
}
The Mathpix CLI:
$ mpx pco convert ./scans/ --formats md,docx
  ✓ contracts/2024-03.pdf  14 pages  2024-03.mmd 2024-03.lines.json 2024-03.md 2024-03.docx  (9.8s)
  ✓ contracts/2024-04.pdf  22 pages  2024-04.mmd 2024-04.lines.json 2024-04.md 2024-04.docx  (14.1s)
  ...
120 files: 120 completed, 0 failed, 0 skipped, 0 still processing

Why Private Cloud OCR?

  • Your data stays yours. No file or file metadata ever reaches Mathpix. A licensed deployment sends a one-time code every 60-90 seconds to prove its contract is active, and a metered one sends daily page counts. An airgapped deployment sends nothing at all.
  • The same models as our cloud API. New releases arrive as a new image tag that you roll out when you choose. An upgrade replaces containers one at a time; documents in flight resume from the page they reached.
  • Built for your infrastructure. One container per GPU, Redis and an S3-compatible object store beside it, both included or your own managed services. helm install on a cluster or docker compose up on one host.
  • Pay for what you use. An annual minimum with included pages plus simple overage pricing. The pages you are billed for are the pages your own deployment reports at GET /pco/v1/usage, so the invoice is auditable on both sides.
Standard Enterprise
Annual minimum $25,000/year Custom
Included usage 25M pages/year Custom
Additional usage $1 / 1,000 pages Volume pricing
Deployment Customer-controlled Kubernetes Customer-controlled Kubernetes
AWS / GCP / Azure ✓ ✓
Customer data center ✓ ✓
Mathpix software updates ✓ ✓
Standard support ✓ ✓
SLA / priority support - ✓
Air-gapped deployment - Optional
Customer infrastructure cost Customer Customer

When should you use it?

Choose Private Cloud OCR when regulatory, contractual or latency requirements rule out sending documents to a cloud API:
  • Documents that cannot leave a jurisdiction, a network, or a machine
  • Contracts that forbid third-party processing
  • Pipelines that already live next to the documents, in your VPC or your data center

What it doesn’t replace

If your constraint is data residency alone, our EU endpoint or the Files API writing results into your own bucket may be the lighter option. You get the same models without having to run anything yourself. Private Cloud OCR is best suited for when the documents themselves must not travel.

Getting started

A Mathpix engineer deploys PCO with you in a single session: a customer-specific image, the Helm chart or the single-host Compose stack, and a smoke test on your documents. Reply to this post’s announcement email, or contact sales@mathpix.com, to schedule it.

Data lifecycle

Inputs, page images and outputs stay in your bucket until your retention rule removes them; the deployment keeps document status for 7 days and its usage counters indefinitely. Mathpix retains no document content, because none is sent. What Mathpix does keep is the deployment metadata it needs to run the contract: a licensed deployment’s check-ins (a one-time code, its version, its deployment id and the time), and for a metered deployment the per-day page counters, the same numbers you can read yourself at GET /pco/v1/usage. An airgapped and unmetered deployment sends nothing at all.

Migrating from SCS Classic?

Your input folders, output layout, conversion options and OCR options carry over. The batch jobs guide has the field-by-field mapping.