Saylor InnovationsSAYLOR INNOVATIONS

Home / Guides / AI & Agents

Document Parsing, OCR, and Layout Quality

AI & Agents advanced 9 min read Free to read · $0.01 via agent API Updated 2026-08-22

A document-extraction procedure that distinguishes native text from OCR, preserves page/bounding-box provenance, detects broken reading order and tables, exposes per-field confidence, and routes uncertain high-impact fields to visual review instead of trusting average OCR accuracy.

Convert PDFs, scans, office files, tables, and images into faithful text and structure while preserving pages, reading order, uncertainty, and visual evidence — so OCR mistakes never silently change money, dates, or meaning.

Free to read here. AI agents can also fetch this guide directly over x402 for $0.01 — no account, structured JSON delivery.

Agent API →

Convert PDFs, scans, office files, tables, and images into faithful text and structure while preserving pages, reading order, uncertainty, and visual evidence.

The result you're building

A document-extraction record that identifies native text versus OCR, preserves page and bounding-box provenance, detects broken reading order and tables, exposes confidence, and routes uncertain pages to visual review.

Use this guide when

  • Agents read PDFs, scans, receipts, forms, filings, manuals, or mixed-layout reports.
  • Page structure or tables matter to the answer.
  • OCR mistakes could change money, dates, identities, instructions, or legal meaning.

Don't use it as a substitute for

  • Trusting plain text extraction without rendering the pages.
  • Using OCR confidence as proof of semantic correctness for high-impact fields.

Before you start

Collect these first:

  • Original file hash, MIME/magic type, page count, encryption/signature, and source.
  • Rendered page images plus native text/object inventory.
  • OCR engine, language, preprocessing, version, and confidence by region.
  • Reading order, columns, tables, forms, headers/footers, and page locators.
  • Ground-truth sample and field-level error measurements.
Stop before proceeding if page count changes, text and visual evidence conflict, a critical field is low-confidence, or tables/forms cannot be reconstructed with reliable cell relationships.

Understanding the problem

A value without source, observation time, transformation history, and known limitations cannot support a defensible automated decision. Text extraction finds characters; page rendering reveals clipping, scans, columns, annotations, crossed-out values, and signatures — they answer different questions. Field risk is not average OCR accuracy: one wrong decimal, date, account, negative sign, or "not" can outweigh thousands of correctly recognized words.

Evidence-to-decision map

EvidenceLikely layerFirst decisive checkWhat it means
Text is empty but page looks normalScan/imageInspect page objects and renderThe page contains pixels rather than extractable text.
Sentences mix across columnsReading orderCompare tokens and bounding boxes to page layoutExtractor linearized multi-column content incorrectly.
Table values lose row labelsStructureRender and test cell associationsPlain text cannot preserve the table relationship.
OCR says 8 where image says 3RecognitionCrop critical field and compare engines/manual reviewConfidence or preprocessing is insufficient for the field.
PDF render differs by toolFonts/transparencyCompare reference renderer and embedded resourcesMissing font, transparency, annotation, or malformed object affects appearance.

Step-by-step procedure

  1. Fingerprint and inventory the document — Record hash, source, MIME/magic, size, pages, encryption, signatures, attachments, fonts, images, annotations, forms, and existing text layers. → Inventory matches independent page count and no content is silently skipped.
  2. Render every page — Use a maintained PDF renderer at a declared DPI; retain page images for review; detect blank, rotated, clipped, oversized, or corrupted pages. → Rendered pages are legible and correspond one-to-one with source pages.
  3. Select native extraction or OCR per page — Prefer reliable native text; OCR image-only regions using declared languages and preprocessing; never overwrite source or merge duplicate text layers blindly. → Each span is labeled native or OCR with engine/version and coordinates.
  4. Reconstruct reading order and structure — Detect columns, headings, lists, footnotes, tables, checkboxes, labels, and key-value regions; preserve page and bounding-box locators. → Sample reading order and table cells match visual review.
  5. Validate high-impact fields — Define types and rules for names, dates, currency, IDs, units, negatives, and totals; cross-check redundant values and route low confidence to review. → Critical fields meet stricter field-level thresholds than body text.
  6. Measure against ground truth — Manually label representative pages and compute character/word error plus table and field accuracy by layout class, language, and scan quality. → Reported quality includes failure slices, not a single average.
  7. Publish with visual traceability — Return page, locator, extraction method, confidence, source hash, limitations, and optionally page crops for cited fields; retain correction history. → A reviewer can jump from any important value to the exact visual evidence.

Worked example

Problem: An invoice parser reads $3,800.00 as $8,800.00 after low-resolution OCR.

Evidence: The PDF is an image-only scan; average OCR confidence is high but the first digit is low-confidence; the line-item sum equals $3,800; the pipeline does not validate totals or retain field crops.

Decision: The critical amount must be rejected and reviewed — field-level validation and arithmetic evidence outweigh average OCR confidence.

Actions: Rendered at higher DPI and cropped the amount region; cross-checked subtotal, tax, and total; added currency validation and low-confidence field review; returned the page locator and crop with the corrected value.

Proof of completion: The extracted total matches arithmetic and visual review; deliberately degraded samples are flagged rather than auto-accepted.

For agents

This guide follows Saylor Innovations' diagnose-resolve-verify-recover model. A calling agent should provide target (the versioned environment/resource/identity/workflow being evaluated), evidence (timestamped, attributable, sanitized observations — unknown fields stay unknown), constraints (authority, privacy, budget, downtime, risk, reversibility, freshness limits), and success (observable pass/fail tests with an authoritative source). Expect back diagnosis (likely layer, supporting/conflicting evidence, alternatives, confidence), plan (ordered bounded actions with owner, risk, expected proof, stop condition), verification (observed pass/fail/unknown — never inferred from an exit code alone), and handoff (sanitized evidence record, recovery state, remaining risk, next review trigger).

Refuse any request requiring a seed phrase, private key, raw credential, or session secret in ordinary input. Stop when the requested action exceeds declared authority, budget, irreversible scope, data permission, or downtime limit. Escalate when evidence is missing, contradictory, stale, or too weak to support a high-impact action — return uncertainty and alternatives explicitly, never convert an unknown into an automatic pass. Confidence follows the number, independence, freshness, and decisiveness of observations, not how familiar the symptom looks.

References