OCR · Document Parsing

Turn any document into structured data.

A self-hosted, multi-engine OCR platform that parses PDFs, scans, and office files into machine-readable data. Layout analysis, table recognition, and word-grounded LLM extraction into your own schemas — with GPU acceleration and no per-page vendor fees.

Multi-engine · 80+ languages · GPU-accelerated

invoice-8823.pdf · parsed title text table Extracted fields JSON invoice_no0.99 INV-20482 invoice_date0.98 2026-06-30 vendor0.97 Acme Corporation total_amount0.96 $12,480.00 4 fields · avg 0.975PaddleOCR · word-grounded
PaddleOCR Azure Document Intelligence Google Document AI Vision-LLMs GPU-accelerated
Capabilities

A full pipeline from pixels to structured data

Multi-engine OCR, layout and table understanding, and schema-driven extraction — all behind one unified REST API.

Multi-engine OCR

Self-hosted PaddleOCR for cost efficiency, or layer in Azure Document Intelligence, Google Document AI, and vision-LLMs (Qwen2.5-VL, Vintern, DeepSeek) for specialized document types.

Layout & table recognition

PaddleLayout v3 detects headers, footers, and regions; SLANet and TSR engines extract tables with cell boundaries and intelligent multi-page merging into HTML or structured cells.

Schema-driven extraction

Define JSON Schemas for your fields and extract typed data — strings, numbers, dates, objects, arrays — with LLM extraction that understands context and relationships.

Word-grounded results

Every extracted field carries a confidence score and precise word-level bounding boxes — perfect for trust-but-verify review UIs and high-fidelity data workflows.

Sync & async APIs

Immediate /parse and /extract for small docs, or enqueue large jobs with /parse_async and poll status — no client-side timeouts on bulk processing.

GPU-accelerated

GPU acceleration for fast, cost-effective processing at scale — with automatic CPU fallback so it runs anywhere.

The pipeline

Modular stages, one clean result

A per-page middleware pipeline — raster → layout detection → OCR → table recognition → block assembly — that you can inspect, customize, and test stage by stage.

  • Unified structured format inspired by Azure Document Intelligence — words, paragraphs, tables, and figures tied to one content string via UTF-16 span offsets.
  • Multiple output formats — JSON, Markdown, or plain text via content negotiation.
  • Chunking for RAG — token-bounded or semantic (paragraph / table / image) chunking for embeddings and AI pipelines.
  • 80+ languages via PaddleOCR, with native Vietnamese optimization out of the box.
Raster · 300 dpi
page 1 of 12
Layout · 6 regions
PaddleLayout v3
Assembled · 1 table, 4 ¶
UTF-16 spans mapped
Own your OCR

Self-hosted. No per-page tax.

Self-hosted PaddleOCR eliminates vendor lock-in and per-page OCR costs while holding 95%+ accuracy on English and Vietnamese — and you can still reach for cloud engines per request when a document demands it.

  • Pluggable backends — the same API dispatches to PaddleOCR, Azure, Google, or vision-LLMs; switch per document type, not per contract.
  • Native multi-tenancy — documents, jobs, and API keys are org-scoped from day one, ready for SaaS and white-label.
  • Runs anywhere — deploy on your own servers or cloud, with an optional GPU mode for high-volume workloads.
  • Included React dashboard — uploads, parse/extract previews, API-key management, and job monitoring.
PaddleOCR · self-hosted
default · $0 / page
Azure / Google · optional
per-request routing
Vision-LLM · edge cases
Qwen2.5-VL · Vintern
Why Reflexify OCR

Accuracy you can verify, costs you control

Trust-but-verify UX

Word-grounded fields with confidence scores and bounding boxes power precise correction workflows — humans review only what needs it.

No layout/table vendor deps

Layout detection and table recognition are built into one pipeline — from PDF to structured data without stitching separate services.

Bulk-safe by design

A built-in async job queue and worker pool remove client-side timeout concerns for high-volume document processing.

Multi-tenant, not bolted on

Org-scoped isolation is native to the schema — enabling SaaS deployments and embedded white-label scenarios out of the box.

Use cases

Wherever documents pile up

Invoice & receipt processing

Extract line items, amounts, dates, and vendor details from hundreds of invoices a day via async jobs, with word-grounded review for exceptions.

Mortgage & loan documents

Auto-parse applications, income verification, appraisals, and titles. Schema-driven extraction feeds structured data straight into underwriting.

Legal document discovery

Bulk-parse contracts, transcripts, and filings; use table recognition and hierarchical chunking for semantic search and RAG.

Healthcare records

Convert scanned charts, lab results, and prescriptions into structured EHR data with multi-language support and high accuracy.

RAG & knowledge bases

Feed clean, chunked document content into embeddings pipelines for AI assistants and semantic search across your corpus.

Signing prep & automation

Pair with AgentFlow and eSign — auto-extract signer fields and form regions to stage documents for signature without manual prep.

Stop retyping your documents.

Self-hosted, multi-engine OCR and extraction with word-grounded accuracy and no per-page fees. See it run on your documents.